Your Agent Aced the Task. Will It Do It Again?
A new Hugging Face blog post describes a diagnostic and guideline system for measuring and reducing variability in LLM agent behavior, arguing that standard average accuracy benchmarks hide an unreliability problem.
The post reports that on AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions, but succeeded in all five runs for only 53.0% of tasks, a 24.4-point consistency gap. On hard tasks, the gap reaches 30 points. The author distinguishes Mean@k, the average pass rate commonly reported as accuracy, from Pass^k, the fraction of tasks where the agent succeeds on all k runs, noting Pass^k is always less than or equal to Mean@k.
The post attributes flipping behavior to the shape of the model's probability distribution at each decision: sharp distributions resist noise, while flat distributions with near-tied tokens can reorder under small perturbations. Because trajectories chain many decisions, the post says, small per-step flip chances compound. The agent ran at temperature 0.0, so the variance described is not ordinary sampling.
The described Consistency Analyzer resamples each decision point in a recorded trajectory using a single call requesting k completions (k=5 by default), producing a per-step consistency score without ground truth or full task replay. Flagged steps become consistency guidelines in the existing ALTK-Evolve storage and retrieval pipeline. The post reports the consistency gap falling from 24.4 points to 12.0 points, with Pass^5 rising from 53.0% to 69.0% and Mean@5 rising from 77.4% to 81.0% on AppWorld test_normal (168 tasks). Medium tiers gained 22.9 points and hard tiers 14.3 points.
On a related task in the same scenario, guidelines lifted Pass^5 by 13.0 points. With gpt-oss-120b, same-task Pass^5 rose from 10.1% to 16.1%, and similar-task generalization reached 8.7 points. The post says methodology and evaluations appear in a technical report on arXiv, and the ALTK-Evolve open-source repository now includes the Consistency Analyzer and consistency-guideline generation.
Based on reporting from the original publisher. Visit the source for full context and later updates.