Measure the whole task
Time spent prompting, integrating, validating and debugging belongs in the denominator. A faster initial draft can still produce a slower finished task.
The important outcome is not lines produced: it is work completed correctly with appropriate human judgment.
New divisions of labor between developers, reviewers and models. Our purpose is to establish the evidence and questions needed for rigorous coverage, not to imply that a proposed future is inevitable.
Time spent prompting, integrating, validating and debugging belongs in the denominator. A faster initial draft can still produce a slower finished task.
A 2025 METR randomized study found longer completion time in one experienced-developer setting. Later follow-up work acknowledged selection effects and uncertainty; neither result can stand in for all developers or tools.
Review responsibility, permissions, context and attribution should remain explicit when AI assists development.
Compare matched work with and without AI, including reviewing and fixing time.
These are starting references for future reporting, not proof that every open question above is resolved.
Edition 1.0 · 9 October 2026. Initial scope and source register created. This brief is a research framework, not a completed investigation; material future corrections should be described rather than silently overwritten.
How corrections are documented ↗