Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care (doi.org)

Introduction: Large language model (LLM)-as-a-judge systems offer scalable evaluation of artificial intelligence (AI)-generated clinical outputs, yet their susceptibility to prompt variability raises concerns regarding reproducibility and alignment with expert judgement. This study examined...

The authors included RRIDs in their in Anaesthesiology Intensive Therapy paper! We value the author's support of reproducibility.
#openresearch#openscience inferred — the author tagged nothing

0 comments — live from bluesky

No comments yet.