OBSERVATORY/FIELD 06 OF 10RESEARCH BRIEF / EDITION 1.0
Benchmarks, correctness, security and the limits of passing tests

Testing & Verification.

A passing benchmark is evidence about a test, not proof of safe production behavior.

EVIDENCE FRAMEWORKREVISION-READYOPEN QUESTIONS
WHAT THIS FIELD EXAMINES

Benchmarks, correctness, security and the limits of passing tests. Our purpose is to establish the evidence and questions needed for rigorous coverage, not to imply that a proposed future is inevitable.

01

Tests are partial oracles

Suites can reject legitimate solutions or miss subtle regressions. Interpret results alongside inspection, requirements and independent checks.

02

Benchmark validity can decay

OpenAI reported in February 2026 that SWE-bench Verified had become unreliable for frontier-model evaluation because of issues including contaminated examples and test design.

03

Verification as a workflow

Threat modeling, reproducibility, code review, CI, dependency hygiene and incident response remain important even as generation improves.

RESEARCH PROTOCOL

How we would test the claim.

Preserve held-out tests, measure false passes, and document benchmark exposure and version.

STARTING SOURCE TRAIL

Documents to consult

These are starting references for future reporting, not proof that every open question above is resolved.

  1. We Are Changing Our Developer Productivity Experiment Design
  2. SP 800-218 — Secure Software Development Framework, Version 1.1
PUBLIC EDITION RECORD

Revision record

Edition 1.0 · 9 October 2026. Initial scope and source register created. This brief is a research framework, not a completed investigation; material future corrections should be described rather than silently overwritten.

How corrections are documented ↗