Peer Reviewed Simulation · Robotics

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners

Yuan Gao et al. — TUM AVS/MIRMI · University College London

Key Insight

193/200 executable scenarios vs. 144 for a baseline; planner success rises from 50.4% to 70.2% with grounded scenario generation. Accepted at EMNLP 2026.

What's Actually Supported

Agent-generated, retrieved, and mutated scenarios materially improve both scenario executability and downstream planner success, in the autonomous-driving domain tested.

Caveat / Limitations

Scoped to an automotive stack; the authors themselves call for broader planners, simulators, and real-world validation before generalizing.

Part of