Tests are partial oracles
Suites can reject legitimate solutions or miss subtle regressions. Interpret results alongside inspection, requirements and independent checks.
A passing benchmark is evidence about a test, not proof of safe production behavior.
Benchmarks, correctness, security and the limits of passing tests. Our purpose is to establish the evidence and questions needed for rigorous coverage, not to imply that a proposed future is inevitable.
Suites can reject legitimate solutions or miss subtle regressions. Interpret results alongside inspection, requirements and independent checks.
OpenAI reported in February 2026 that SWE-bench Verified had become unreliable for frontier-model evaluation because of issues including contaminated examples and test design.
Threat modeling, reproducibility, code review, CI, dependency hygiene and incident response remain important even as generation improves.
Preserve held-out tests, measure false passes, and document benchmark exposure and version.
These are starting references for future reporting, not proof that every open question above is resolved.
Edition 1.0 · 9 October 2026. Initial scope and source register created. This brief is a research framework, not a completed investigation; material future corrections should be described rather than silently overwritten.
How corrections are documented ↗