When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving
Do benchmark score gains still reflect real driving improvements once the score becomes the optimization target? Controlled experiments say: not necessarily.
Content tagged with "benchmark"
Do benchmark score gains still reflect real driving improvements once the score becomes the optimization target? Controlled experiments say: not necessarily.
A closed-loop benchmark evaluating how vision-language models respond to temporally grounded safety warnings in cooperative driving.
A generic benchmark for cooperative autonomous driving supporting multi-vehicle, multi-modal, and multi-task evaluation.
Towards generic cooperative autonomous driving benchmark (ICRA 2026).