Agents' Last Exam
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue...
https://arxiv.org/abs/2606.05405

AI Agent Benchmark for Real-World Professional Workflows
Agents' Last Exam evaluates AI agents on long-horizon professional workflows with verifiable outcomes across industries such as finance, robotics, bioinformatics, media, and more.
https://agents-last-exam.org/

Seonglae Cho