Scale AI removes 89 tasks from SWE-Bench Pro after audit reveals gaming
Scale AI cut 89 tasks from its public coding benchmark and locked the evaluation runtime after audits found reward hacking and solution leakage.
Daily coverage of AI, developer tools and infrastructure. Each story explains what happened and why it matters, with a link to the original source.
Scale AI cut 89 tasks from its public coding benchmark and locked the evaluation runtime after audits found reward hacking and solution leakage.