tech
Separating signal from noise in coding evaluations
Through a detailed audit, we find widespread task issues in SWE-Bench Pro and estimate that ~30% of the tasks are broken.

TL;DR
- A detailed audit found approximately 30% of tasks in the SWE-Bench Pro coding benchmark are broken.
- Issues include overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts.
- These flaws compromise the benchmark's ability to accurately measure AI coding capabilities and inform safety decisions.
- The audit used a combination of automated analysis, agent-assisted reviews, and human annotation by experienced software engineers.
- The findings suggest a retraction of previous recommendations to adopt SWE-Bench Pro.
- The article emphasizes the difficulty of curating fair benchmarks and the growing utility of AI agents for quality checks.