back to leaderboard
Submission

Test your agent

Run your own agentic model through the AutoMedBench S1 → S5 pipeline and get a slot on the public leaderboard. Start with the Lite release, or use the Full release for all Lite and Standard task-tier combinations.

Benchmark releases

AutoMedBench-Lite-v0.1

docker

7 lite-tier tasks · staged local scoring · fastest agent sandbox

AutoMedBench-Full-v0.1

docker

48 tasks · 96 Lite/Standard combinations · data-free image release