AI breakthroughs in robotics won’t change your life any time soon

Tesla's Optimus humanoid robot has become a familiar sight on video feeds, dancing, passing popcorn and pressing microwave buttons, and occasionally falling over while handing out water bottles. Elon Musk, Tesla's CEO, told shareholders in July that the robot will eventually have human and then superhuman dexterity, and argued that Optimus could automate almost all human labor, from hauling sheet metal to folding laundry, for as little as $20,000 each. Speaking at the World Economic Forum in Davos, Switzerland, in January, he predicted the robots could go on sale to the public by the end of 2027. Marc Andreessen has called robotics potentially the biggest industry in the history of the planet, Nvidia CEO Jensen Huang said in January that humanoid robots would match human-level ability this year, and Morgan Stanley projects that robots resembling and acting like humans could number nearly 1 billion by 2050, creating a market worth over $5 trillion.
Many robotics researchers push back on those timelines. Yann LeCun, often described as one of the godfathers of AI, said at a Davos event in January that none of the companies building humanoid robots has any idea how to make them smart enough to be useful. Researchers also argue that humanoid form and generalist capability are being conflated. Agility Robotics cofounder and chief robot officer Jonathan Hurst, a robotics professor at Oregon State University, put it plainly: it is very easy to make a robot that looks like a person, but dramatically more difficult to make a machine that moves or behaves dynamically or physically like one.
To understand what is actually working, look at ALOHA 2, short for A Low-cost Open-source Hardware System for Bimanual Teleoperation, the bench-top rig Google DeepMind uses to test its most advanced robotics AI, Gemini Robotics. It consists of a pair of arms, grippers and cameras, and under Gemini Robotics control it can pack a lunchbox: placing bread into a Ziploc bag, sealing it, putting grapes in a Tupperware container, securing the lid and zipping everything into the lunchbox. The key change behind such demonstrations is the robot policy, the logic that decides how a machine reads its surroundings, plans movement and executes a task. Policies used to be hard-coded rules, thousands of lines of software specifying each millimeter of motion across hundreds of tasks.
Over the past few years those policies have been handed to AI systems. Vision-language models, trained on images as well as words, brought contextual understanding: shown a coffee spill, a VLM can identify a cloth. Vision-language-action models added motion commands, and are trained on images or video of a task paired with data about how a robot arm moves to perform it, usually gathered through teleoperation, where a human drives the robot through an action. Place a VLA-powered robot at a desk and ask it to close a laptop or wrap a headphone wire and it will survey the scene and act, provided it has seen the task before. Gemini Robotics is a VLA, trained on many hours of human demonstrations, and can pick up snow peas with tongs, do origami or assemble a simple lunch. Outside its training set, it is highly likely to fail. Edward Johns of Imperial College London notes that a true generalist policy would handle everything along the spectrum of tasks, whereas today's models manage a few things here and a few things there.
The standard answer to that gap is more training data, and Google DeepMind wants to gather as much as possible, according to Pannag Sanketi, a former robotics tech lead there now working on his own AI robotics project. Unlike language models, which had oceans of text available, robotics has no comparable pool of high-quality physical demonstrations. The alternatives each have drawbacks: paying people to collect teleoperation data is costly and slow; training on videos of humans acting yields poor-quality data; and deploying robots in the real world to harvest experience is unsafe and unreliable outside labs. Sanketi favors a multi-prong approach drawing on all of those sources. Hurst calls the premise that data alone solves the problem fundamentally flawed, because real-world tasks explode in complexity. Making coffee, for instance, involves different kitchens, machines, cups and grips, and achieving generality through VLAs would require complete data coverage of everything a robot could ever do, which amounts to an almost infinite pool.
LeCun dismisses the whole approach, saying the AI methods that succeeded for language do not work for the high-dimensional, continuous, noisy data common in robotics and that something else is required. The leading candidate is the world model, trained less on text than on video, 3D scans and sensor data, and built to predict the outcomes of actions in the real world. A sufficiently faithful internal representation of how objects move, collide, fall and deform would let roboticists train in simulation that is faster, cheaper and safer than real-world testing, and would let robots anticipate consequences rather than merely react. Nvidia and Google are working on the technology, and investors are funding startups: World Labs, cofounded by Stanford's Fei-Fei Li, raised $1 billion in February and was acquired by AMD at the end of September for $8.2 billion, while AMI Labs, cofounded by LeCun, raised $1 billion in March. Li has described the field as nascent, with foundational approaches still being established, making world models a promising research direction rather than an immediate route to general-purpose robots.
San Francisco startup Physical Intelligence, or PI, offers one of the clearest glimpses of where the work stands. Its first generalist system, pi-zero, was published in 2024 as a VLA trained on a 10,000-hour proprietary collection of teleoperated human demonstrations plus open-source robot datasets. A spring 2025 version, pi-0.5, added labeled web images for versatility, and a fall 2025 update, pi-0.6, added reinforcement learning, with each iteration improving from slowly folding laundry to putting things away in new environments and folding boxes more reliably. The April 2026 release, pi-0.7, added a lightweight world model that generates images of the steps needed for a task and feeds them to the robot as it works. The company claims the first signs of compositional generalization, meaning the ability to perform skills never seen in training by recombining learned ones. In one test the model was asked to load a sweet potato into an air fryer, a task it had not encountered; a demonstration video shows false starts and a reasonable effort that does not quite finish. PI cofounder Sergey Levine, a professor at UC Berkeley, called it the first convincing instance of that kind of generalization. Investigation later found snippets of relevant labeled teleoperation data in the training material, including two examples of a human controller pushing an air fryer basket in.
A tension runs through much of what the public sees. Many flashy demonstrations involve humans controlling the robot or carefully scripting its actions; the robot that appeared onstage with Jensen Huang in March 2025 was remote-controlled, according to its makers, by a puppeteer behind the scenes. Fully autonomous motion planning in new, chaotic environments such as a construction site or an unfamiliar home remains largely unsolved, and getting robots to handle larger, ambiguous jobs like make dinner, which require deciding what to do and in what order, is harder still.
Why it matters: product and engineering teams weighing robotics investments should treat the 2027 consumer-sale prediction as an aspiration rather than a schedule, since the published capability frontier is narrow, teleoperation-heavy and fragile outside training conditions. The nearer-term opportunity lies in constrained, repetitive industrial and lab tasks where data can be collected systematically and failure is cheap, and in the tooling around data collection, simulation and evaluation. If world models mature, the training bottleneck shifts from gathering demonstrations to building simulation fidelity, which would favor teams with simulation and sensor expertise. The main uncertainty is whether scaling data and VLAs is enough or whether world models are required first, and that question is unresolved.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The story is a collaboration between MIT Technology Review and Aventine, a non-profit research foundation that creates and supports content about how technology and science are changing the way we live. A robot shaped like a human—white with a black head and torso—has been popping up on video feeds. Perhaps you’ve seen it dance or…