The LLM Judge That Kept Agreeing With Itself
Towards Data Science
Read full postA system using a language model as a judge to approve SQL queries repeatedly approved incorrect queries due to structural bias, prompting developers to treat the judge as a component needing its own testing and calibration against human reviewers.



