M25.8 CONNECT THE MECHANISM
Evaluate the completed task and the trajectory that produced it
Wayfarer said "Booked!" nine times out of ten. The hotel's database says six. Learn to test an agent by what it changed, what it cost, and whether it succeeds every time.
LESSON OVERVIEW14 min lesson
Lesson overview
Wayfarer said "Booked!" nine times out of ten. The hotel's database says six. Learn to test an agent by what it changed, what it cost, and whether it succeeds every time.
What you’ll explore
- Agent evaluation needs reproducible environments, task-level success criteria, action traces, and resource accounting; success is judged by the environment's final state, not by the agent's claim that the work is done.
GO TO THE SOURCE
Original explanations, connected to the research.
AgentBenchSWE-bench: Can Language Models Resolve Real-World GitHub Issues? (Jimenez et al., 2023)WebArena: A Realistic Web Environment for Building Autonomous Agents (Zhou et al., 2023)τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (Yao et al., 2024)Evaluating Large Language Models Trained on Code (Chen et al., 2021), source of the unbiased pass@k estimatorSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.