Dev6 min reading time
Separating signal from noise in coding evaluations
Open AI News
Read full postOpenAI audited the SWE-Bench Pro coding benchmark, finding over a third of tasks flawed due to strict tests, underspecified prompts, low test coverage, or misleading instructions, impacting accurate model evaluation.




