What AI gets wrong and what failure teaches us

Microsoft Research released a podcast episode featuring Jennifer Neville, a partner research manager at Microsoft Research and the Samuel D. Conte chair professor of computer science and statistics at Purdue, in conversation with principal applied scientist Chad Atalla.
According to the episode description, Neville explores the role evaluation plays in pushing the performance boundaries of today's AI systems to meet user needs, and the "surprising failures" that emerge when models are tested beyond traditional benchmarks. She also shares practical guidance for working with current AI systems and discusses why looking closely at data matters when results defy expectations.
Neville leads the AI Interaction and Learning team. She said the team focuses on pushing the frontier of AI system behavior in realistic work environments: studying where the boundary of current system performance lies, how users experience that in practice with real workflows, and how to improve and push model performance for complex tasks. She said standard machine learning and AI evaluation relies on benchmarks that are fairly simple relative to actual use, and that the team designs practical evaluations covering multiturn behavior, collaborative environments, and long-horizon tasks.
She described research in which public single-turn benchmark datasets were modified so that a fully specified complex instruction was instead provided as clarification over multiple turns by models simulating users. She said performance seen in single-turn scenarios degrades significantly over multiple turns, which she called a fairly surprising finding, and that users resonated with the result. A practical suggestion in the paper: if a model becomes confused after multi-turn specification, stop the chat, erase everything, and give a fully specified single turn.
Neville said the team also works with Microsoft product groups analyzing consumer logs at scale to identify main patterns of user failures, which motivates much of the work. She noted the team is working on reinforcement learning methods to help models learn better in these environments.
Biographical details from the episode: Neville joined Microsoft in 2021 from academia, has been teaching and advising for 20 years, has published more than 130 papers with over 10,000 citations, and has received a National Science Foundation CAREER Award, a spot on IEEE's "10 to Watch" in AI list, and best paper awards from the International Conference on Data Mining and the International Conference on Learning Representations.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research .