AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Academia is for Ambition — Alex Zhang, MIT

Collected Oct 2, 2026

Latent Space published a podcast episode featuring Alex Zhang of MIT, whose work spans GPU kernels, KernelBench, Recursive Language Models (RLMs), and large multi-agent swarms. The episode's description states that an RLM-based harness was the first to approximately solve ARC-AGI-3, before OpenAI's Astra, and that it is influencing new research.

In the transcript, Zhang said he became involved with GPU Mode (formerly CUDA Mode), a Discord started around 2023 by Mark Saroufim, Andreas, and Jeremy Howard, originally dedicated to learning how to write GPU kernels. He said he joined while interning at Snapchat, where he explored writing specialized kernels for a Google paper called Infinite Attention. He described KernelBench as having spawned from the idea of getting LLMs to automate GPU kernel code.

On AI-written kernels, Zhang said that on the GPU Mode leaderboard almost all recent solutions are AI-generated, but that a member known as Gauners, a strong GPU kernel writer, produced in the top 10 basically the only kernel stable in end-to-end systems. He said GPU kernels have a verification problem and that reward hacking has been an issue since KernelBench was released, adding that there is still "a lot of alpha" in being good at writing GPU kernels. He also said a knowledgeable person can uncover something for a model that would erase a trillion-token spend, and that compute efficiency still matters.

Asked about speed-of-light limits, Zhang said for matrix multiplication a theoretical fastest-kernel estimate is easy to compute, but that it is not clear such numbers are reachable, and that kernels people write are often not close to them. He discussed kernel fusion or megakernel approaches, noting they generally apply when memory-bound. The episode also covered research taste, GEV and alternatives to standard autoregressive language models, harnesses as compositional generalizers, Prime Agent, OpenAI's 10,000-agent experiment with 130B output tokens, Kimi swarms, Sakana AI, speculative programmatic tool calling, capability overhang, and Neuralese.

Read at Latent Space

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.