Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead
In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details...