AivexaNewsSearch
AI news for builders and product teamsChecked every hour

SOP-Bench: A new benchmark for evaluating AI agents on real business procedures

Collected Oct 1, 2026

Amazon has released SOP-Bench, an openly available benchmark that measures how well AI agents carry out real standard operating procedures (SOPs) written by domain experts. The company presented the benchmark at the 2026 Conference on Knowledge Discovery and Data Mining (KDD).

SOP-Bench covers 12 business areas, including healthcare intake, dangerous-goods classification, customer service, content moderation, financial compliance, and warehouse inspection, with more than 2,000 tasks. Each task includes the tool interfaces an agent needs and a correct outcome, allowing results to be checked against ground truth. The framework ships with two baseline agents and lets teams substitute their own or add procedures; each procedure consists of the SOP text, callable tools, tool specifications, and test cases with known answers.

Procedures were authored by experts from real industrial workflows, while an Anthropic Claude 3.5 Sonnet v2 model generated data schemas, mock APIs and tool specifications, tool code, and datasets mixing ordinary cases, edge cases, and outright failures. Experts reviewed and corrected every generated item and ran the code. Amazon said no proprietary or sensitive data was involved.

Amazon ran a function-calling agent and a reasoning-style agent across 11 frontier models. It reported that on the reasoning-style agent, the newer Claude 4.5 family scored lower than the older Claude 4 family, and that the same reversal held in comparisons of individual models on the same setup. In one video-annotation procedure, success nearly halved when the six required tools were joined by 20 extra plausible but useless tools.

No model-agent pairing led across the board. On the easiest procedures, such as triaging incoming e-mails by intent, agents were correct approximately nine out of ten times; on the hardest, such as annotating objects in a driving video, approximately one out of four. The reasoning-style agent came out slightly ahead on average in head-to-head comparison but won only eight of thirteen procedure runs and took about a third longer per task. Amazon also described an open question: a procedure with long stretches of reading and few decisions gave agents more trouble than one with many more decisions, though it said the two procedures differ in other ways, including tool count.

The full benchmark, including the 12 expert-authored procedures, generated tools and datasets, the two baseline agents, and evaluation code, is being released on GitHub and HuggingFace. Amazon said it plans to add harder variants, instructions with images and tables, and nested procedures that require context switching.

Read at Amazon Science

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Extendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.