Giving your AI a Job Interview

A new essay argues that benchmarks, the primary way AI capability is measured, have significant shortcomings, and that users and organizations need other ways to judge which model fits their needs.
The problems described include public benchmarks and answer keys that some AIs may incorporate into training, uncertainty about what tests actually measure, uncalibrated scoring, and errors in test questions that can make top scores unachievable. Examples cited include MMLU-Pro questions such as the approximate mean cranial capacity of Homo erectus and the place named in the title of a 1979 live album by Cheap Trick.
Despite these flaws, the essay says benchmarks collectively trend upward and appear to measure some underlying ability factor, with ARC-AGI and METR Long Tasks showing the same trend. It cites AIME for math, GPQA for scientific and legal knowledge, MMLU for general knowledge, SWE-bench and LiveBench for coding, and Terminal-Bench for agentic ability.
Few robust individual benchmarks exist for writing, sociological analysis, business advice or empathy, according to the essay. It suggests informal "vibes" testing, such as asking models to draw a pelican on a bike, a reference attributed to Simon Willison, or an otter on a plane, as a way to sense a model's world model.
A writing exercise about someone with 47 words remaining is described as revealing differences among Claude 4.5 Sonnet, Gemini 2.5 Pro, GPT-5 Thinking and Kimi K2 Thinking. The essay says vibes testing is idiosyncratic and relies on feelings rather than real measures.
For organizations, the essay compares model selection to hiring and calls for a rigorous "job interview." It points to OpenAI's GDPval paper, which gathered experts with an average of 14 years of experience in industries from finance to law to retail to generate realistic projects taking human experts an average of four to seven hours.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
As AI advice becomes more important, we are going to need to get better at assessing it