AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Real AI Agents and Real Work

Collected Oct 1, 2026

OpenAI released a new test of AI ability that differs from benchmarks built around math or trivia, according to an account from One Useful Thing. Experts with an average of 14 years of experience across industries including finance, law, and retail designed realistic tasks that would take human experts an average of four to seven hours to complete.

Both AI models and other experts then performed the tasks. A third group of experts graded the results without knowing which answers came from the AI and which from humans, a process that took about an hour per question. Human experts won, but barely, and the margins varied dramatically by industry. More recent AI models scored much higher than older ones.

The major reason AI lost to humans was not hallucinations and errors, but a failure to format results well or follow instructions exactly, areas described as rapidly improving. If current patterns hold, the next generation of AI models should beat human experts on average on this test, the account states.

That does not mean AI is ready to replace human jobs soon, because what was measured was tasks, not jobs, according to the account. Jobs consist of many tasks; AI doing one or more tasks shifts what a person does rather than replacing the entire job, and AI remains jagged in its abilities and cannot substitute for all the complex work of human interaction.

The account also describes giving Claude Sonnet 4.5, to which the author had early access, the text of a sophisticated economics paper involving multiple experiments along with the archive of its replication data. The prompts were to replicate the findings from the dataset, doing the work itself, and to attempt a full replication or do what it could; because the work involved complex statistics, the author also asked it to replicate the full interactions as much as possible.

Read at One Useful Thing

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

The race between human-centered work and infinite PowerPoints