LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
Hacker News
Read full postResearchers evaluated large language model (LLM) judges' ability to detect omissions in AI-generated clinical notes by comparing flawed notes with transcripts. They found standard LLM judges struggle to reliably identify missing information but developed new methods that improve omission detection with acceptable false alarm rates. These methods were validated by physicians and released as benchmarks and tools for further research.



