Skip to main content
Anthropic’s Demystifying evals for AI agents is the clearest published account of how agent evals fail in practice. This page maps each recommendation to the thing you actually run. Every row names a command or API that exists in the release stated — no aspirational rows.

Vocabulary

The post’s vocabulary and multivon-eval’s API names, stated once. The API is not being renamed — a rename would break every existing suite for zero information gain. New warning and report copy uses the post’s terms (“40/40 trials passed”, “broken task or grader”, “graduate this suite”).

The mapping

The four 0.16.0 features protect the places where evals lie most confidently: the 0% floor (validate), the 100% ceiling (saturation monitor), the single-number middle (pass@k vs pass^k), and the coerced verdict (the judge-UNKNOWN row — hedged judge verdicts parse as UNKNOWN, never as a silent yes).