Preprint Agentic AI ยท Evaluation

ExecCritic: Learn to Test, Test to Improve for Coding Agents

Leitian Tao et al. โ€” UW-Madison / Microsoft Research / Georgia Tech

Key Insight

Independent Test/Repair agents in a fail-closed harness reach 72.6% on SWE-bench Verified, +11.4 points over a no-test baseline.

What's Actually Supported

Independent, separately-trained Test and Repair agents materially improve resolution rate over a single self-testing agent, per SWE-bench Verified results.

Caveat / Limitations

Depends on isolated sandboxes and substantial compute. A weak Test agent alone can reduce resolution โ€” role separation, not just testing, is doing the work.

Part of