ExecCritic: Learn to Test, Test to Improve for Coding Agents
Leitian Tao et al. โ UW-Madison / Microsoft Research / Georgia Tech
Key Insight
Independent Test/Repair agents in a fail-closed harness reach 72.6% on SWE-bench Verified, +11.4 points over a no-test baseline.
What's Actually Supported
Independent, separately-trained Test and Repair agents materially improve resolution rate over a single self-testing agent, per SWE-bench Verified results.
Caveat / Limitations
Depends on isolated sandboxes and substantial compute. A weak Test agent alone can reduce resolution โ role separation, not just testing, is doing the work.