AGI Is Not Multimodal

An essay published by The Gradient argues that the multimodal approach to artificial general intelligence is "sure to fail in the near term" and will not produce human-level AGI capable of sensorimotor reasoning, motion planning, and social coordination. The author contends that these models scaled effectively on existing hardware rather than emerging as thoughtful solutions to intelligence.
The piece defines AGI as necessarily general across all domains, including problems originating in physical reality such as repairing a car, untying a knot, or preparing food. It argues such problems require intelligence situated in something like a physical world model and cannot be fully represented by symbols and solved through symbol manipulation. A book, Designing an Intelligence, edited by George Konidaris, is cited as forthcoming from MIT Press.
The author disputes the claim that large language models learn world models through next-token prediction, suggesting instead that they learn "bags of heuristics" and a model of syntax. Evidence cited includes the Othello paper, in which researchers predicted a game board from hidden states of a transformer trained on legal moves; the essay argues the results do not generalize because Othello resides in symbols, and notes a blog post stating OthelloGPT learned sequence prediction rules that do not hold for all possible games. Melanie Mitchell's recent piece and a paper are also cited as evidence that generative models can score well on sequence prediction while failing to learn the worlds that created the data.
The essay revisits Rich Sutton's "The Bitter Lesson," a recent Turing Award recipient with Andy Barto, arguing it has been misinterpreted as ruling out any structural assumptions. It cites Convolutional Neural Networks, the Transformer attention mechanism, and 3D Gaussian Splatting as human-intuition-driven architectural advances. It also argues multimodal models contradict the Bitter Lesson by assuming how modalities should be processed and joined, that modality-specific encoders and decoders leave "meaning" decentralized and potentially inconsistent, and that today's modalities may not appropriately partition observation and action spaces for an embodied agent.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of generative AI models have convinced some that AGI is imminent. While these models appear to capture the essence of human