AivexaNewsSearch
AI news for builders and product teamsChecked every hour

A new benchmark for evaluating patient-facing health AI agents

Collected Oct 1, 2026

Amazon Science introduced PatientAgentBench, a reproducible, clinician-vetted evaluation standard for patient-facing health AI agents in healthcare, per a company announcement. The framework generates a synthetic patient chart and health record, a clinical vignette derived from that record, and a patient agent that converses with the health AI system under evaluation. The evaluated system is itself an agent: a base model with a harness that reasons over patient context and uses the benchmark's stateful, simulated healthcare tools.

According to the announcement, an LLM-as-a-jury panel scores each conversation using more than 100 clinician-vetted criteria across six dimensions: clinical safety, triage quality, workflow accuracy, task completion, clinical helpfulness, and conversational quality. Licensed clinicians validated the automated evaluation by annotating a shared sample of conversations; their scores aligned strongly with the jury, on par with or exceeding human inter-annotator agreement. All patient profiles, clinical narratives, and conversations are fully synthetic, with no real patient health information used.

Amazon Science said it evaluated multiple families of frontier models on thousands of shared multiturn patient conversations. Out of the box, in a baseline agentic harness, even the most capable models fell short of the standard patient-facing care requires. Three findings were reported. Triage showed the widest divergence: routine administrative requests from clinically complex patients were hardest, producing what the announcement called a severity paradox in which agents score higher on clearly severe cases than on mild or routine ones. Safety failures concentrated in patterns: crisis resource omission, such as recognizing suicidal ideation but failing to provide hotline information, and clinical-information fabrication, including invented provider credentials, fake citations, and claimed tool executions that never ran. More-capable models narrowed clinical gaps but did not close them.

The framework is being released on GitHub, including the synthetic-scenario generation pipeline, the healthcare sandbox with stateful tools, the dual-agent-conversation runner, and the LLM-as-a-jury evaluation system with all six rubric prompts. It generates fresh scenarios on demand from a configurable seed distribution rather than including a fixed dataset. The announcement states these agents are designed to support, not replace, a patient's provider.

Read at Amazon Science

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.