Mamba Explained
Mamba is a State Space Model (SSM) presented as an alternative class of model to Transformers, according to an explainer published by The Gradient. The article states that Mamba promises performance and scaling laws similar to those of the Transformer while remaining feasible at long sequence lengths, such as 1 million tokens, and that it runs up to 5x faster than a Transformer.
The Gradient writes that Mamba removes the quadratic bottleneck in the attention mechanism. Attention lets every token look back at every previous token, which the article describes as O(n²) time complexity in training and O(n) time per autoregressively generated token, with an O(n) key-value cache. Sliding Window Attention and FlashAttention are cited as ways to mitigate that bottleneck on the margin, but the article argues a different approach is needed for very long context windows.
The Mamba authors, Gu and Dao, are quoted as saying that Mamba enjoys fast inference and linear scaling in sequence length, and that its performance improves on real data up to million-length sequences. They also state that Mamba achieves state-of-the-art performance across several modalities including language, audio and genomics, and that their Mamba-3B model outperforms same-size Transformers and matches Transformers twice its size in both pretraining and downstream evaluation.
The article describes the Mamba block as replacing attention, used for communication between tokens, with a Control Theory-inspired SSM, while retaining MLP-style projections for computation. It explains the SSM through continuous-time differential equations, discretisation via Zero-Order Hold, and interpretations of the A, B, C and D matrices and the step size Delta.
A central claim in the explainer is selectivity: in Mamba, the A, B, C and D matrices are functions of the input x, making them context dependent rather than static. The authors are quoted saying the efficiency versus effectiveness tradeoff of sequence models is characterised by how well they compress their state, and that a fundamental principle for building sequence models is selectivity. The article states Mamba has O(1) space and O(n) time requirements, but notes that the Selection Mechanism loses the convolutional form, with hardware optimisation work making Mamba run faster than comparably sized Transformers. The Gradient also says the piece covers advantages and disadvantages of Mamba versus Transformers, analogies, and what Mamba means for interpretability, AI safety and applications.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Is Attention all you need? Mamba, a novel AI model based on State Space Models (SSMs), emerges as a formidable alternative to the widely used Transformer models, addressing their inefficiency in processing long sequences.