How to Evaluate AI Agents From Tool Calls to Task Completion
NVIDIA published a developer blog post describing how agent evaluation has moved from scoring individual function calls to scoring whether an entire task is completed. The post says the key question when shipping an agent is whether it can execute a chain of work across dozens of sequential tool calls against a live environment and recover when a step fails, and that scoring whether a model sounds right reveals little about whether the work finished.
The post states that full agentic evaluation requires an execution environment that runs each tool call, tracks state across steps, and reads the environment afterward to determine whether the work was done. Two scoring layers sit on top: step-level process scoring, which asks whether a call was valid, relevant, and useful given the state at that point, and end-to-end outcome scoring, which checks only the final state. The post describes both as readings of the trace, the ordered log of a single attempt.
The post says benchmark runs roll up through a fixed hierarchy of Benchmark, Trial, Task, Turn, and Step, with metrics on accuracy, verbosity, and cost axes, and that paired reporting matters, such as success rate with consistency ranges. It says benchmark comparability depends on task complexity, environment statefulness, and verification methodology, with executable verification preferred over reference-based evaluation and LLM-as-a-Judge, whose scores it describes as provisional until validated against human ratings on a sample.
As an example, the post says Nemotron 3.5 Lightning achieves 86% accuracy on PinchBench while finishing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy. It recommends that enterprise deployment decisions prioritize domain-specific evaluations built from real tickets and APIs, gated on environment state rather than isolated call accuracy, and points to reproducibility docs, build.nvidia.com, and a NIM guide.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and...