AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Car-GPT: Could LLMs finally make self-driving cars happen?

Collected Oct 1, 2026

Jérémy Cohen's article "Car-GPT: Could LLMs finally make self-driving cars happen?", published in The Gradient in 2024, examines whether large language models could provide an unexpected answer to autonomous driving, comparing the possibility to Alexander Fleming's accidental 1928 discovery of penicillin.

The article describes the modular approach that most self-driving cars used in the 2010s, splitting autonomous software into Perception, Localization, Planning and Control modules, and notes that a decade later companies began taking End-To-End learning seriously, an approach it says introduces a black box problem because it replaces every module with a single neural network predicting steering and acceleration.

It identifies four active research areas in 2023 for LLMs in self-driving: Perception, Planning, Generation and Question & Answers. In perception, it cites GPT-4 Vision, HiLM-D, MTD-GPT and PromptTrack, describing PromptTrack as combining the DETR object detector with large language models, sending multi-view images to an encoder-decoder network, combining predicted annotations with a prompt such as "find the vehicles that are turning right", then finding 3D bounding box localization and assigning IDs using bipartite graph matching such as the Hungarian Algorithm.

For planning, the article cites Talk2BEV, which it says works with LLaVA and ChatGPT4 and accepts Bird Eye View input, and DriveGPT, described as sending perception output to Chat-GPT and fine-tuning it to output a driving trajectory directly. For generation, it cites Wayve's GAIA-1, which takes images, actions and text prompts as input and uses a world model to produce video, along with MagicDrive, Driving Into the Future and Driving Diffusion.

On trust, the article raises the possibility of hallucinations and says it is very early in the research, adding that it is not sure anyone has used these models live in a car on the streets, as opposed to in a headquarters for training or image generation. It concludes that it is too early to tell. The author is described as a self-driving car engineer and founder of Think Autonomous.

Read at The Gradient

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Exploring the utility of large language models in autonomous driving: Can they be trusted for self-driving cars, and what are the key challenges?