AivexaNewsSearch
AI news for builders and product teamsChecked every hour

BenchMIRT: What are LLM benchmarks actually measuring?

Collected Oct 1, 2026

Researchers introduced BenchMIRT, a method for auditing LLM benchmarks at the level of individual prompts, according to a Hugging Face blog post. The approach draws on Item Response Theory (IRT) from psychometrics and extends it with multidimensional IRT, or MIRT, to separate multiple capabilities that may contribute to performance on the same questions. It estimates a model's strength on capabilities reflected across selected benchmarks and, for each question, its difficulty and how well it distinguishes stronger from weaker models.

BenchMIRT was trained on benchmarking results from 100 LLMs across 16 benchmarks and more than 34K questions. Six benchmarks measure general reasoning, including MMLU-Pro, GPQA, MATH and BBH; the other 10 come from the Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP and XSTest. Without being told which benchmarks measured which capabilities, it independently recovered two dominant dimensions: safety and general reasoning, a result the researchers say was stable across a repeated analysis.

For many benchmarks, BenchMIRT largely confirmed intended focus, but it also revealed a more complicated picture in some evaluations. BBQ, commonly grouped with safety benchmarks, aligned much more strongly with general reasoning. WMDP scores were more strongly associated with general reasoning than safety, with stronger general reasoning associated with lower WMDP scores because the benchmark counts refusing or failing to provide dangerous knowledge as the desired response. Within HarmBench, standard and contextual questions aligned more closely with safety, while copyright questions were more closely associated with general reasoning. The researchers state these findings do not necessarily mean the benchmarks are flawed or incomplete.

The method can also rank questions by how well they distinguish stronger from weaker models. Across the 16 benchmarks, keeping only 10% of questions generally preserved nearly the same picture of model strength on safety or reasoning, and keeping 50% often matched the full benchmark even more closely. BenchMIRT correctly predicted whether a model would answer a held-out question correctly 79% of the time, versus 70% for a simpler approach.

Stated limitations: the models used were all released by March 2025, so the analysis does not capture newer LLM generations; discovered dimensions depend on the benchmark set given; and for ranking models by predicted performance on randomly held-out items, the benchmark's average score performs slightly better. The researchers also note the same question-level estimates could be used to remove informative safety questions, producing a weaker evaluation an unsafe model could pass.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt