Mamba vs Transformers: What I Learned Chasing a Context Window
Apr 5, 2026 · 14 min read
The First Time It Broke
Everyone told us attention is all we need. Nobody mentioned attention is also what we pay for, and that the bill grows fast. Back in 2020 I was running a chatbot startup. Our intent classifier used word embeddings, Word2Vec and GloVe, feeding into a static encoder. It worked fine. User types something, encoder spits out a fixed vector, we match it against known intents. But accuracy had a ceiling. Those embeddings were context-free. "Bank" got the same vector whether someone was talking about a checking account or a fishing trip. We were losing nuance and it showed up as misrouted conversations. So I benchmarked GPT-2 as a replacement. Same idea, but with contextual representations instead of static ones. On short messages the jump was obvious right away. Intent matching got sharper. Then I fed it longer inputs. Full conversation histories, fifteen or twenty turns. The gains evaporated. Predictions went vague. Anything the user said early in the thread stopped mattering. I tried truncation, summarizing the history, reformatting the prompt. Nothing fixed it. And inference cost was way above our static encoder. The math never worked out, so we shipped with word embeddings and moved on. I knew there was a wall. I just couldn't explain it.
Five Years Later, Same Wall
Last year I was using Claude to refactor some Python. Pasted in four files, around 800 lines, asked it to trace a bug through the call chain. It nailed it in seconds. Then I pasted a whole module. Twelve files, roughly 3,000 lines. It started strong and fell apart around the seventh file. Functions from the first file went fuzzy. Variable names came back wrong. It read like someone who studied chapter one and skimmed the rest. Codex did the same thing. Short prompts, great completions. Long prompts, mush. Two different models, five years apart, identical failure. The model wasn't dumb. It was out of room. This time I wanted to know exactly why.
Tokens and Context, Quickly
A token is the smallest unit an LLM reads. Not a word, not a character, somewhere in between. "Understanding" might split into "under" and "standing." Roughly three-quarters of a word per token. The context window is how many tokens the model can hold at once. Your prompt and its response both have to fit. Claude's latest models go up to 1M tokens. GPT-4 Turbo has 128K. GPT-2 had 1,024. > [!NOTE] The model has no memory between messages. No hard drive. Every time you hit send, the entire conversation gets repacked into the context window and reprocessed from scratch. That hour-long chat? Every message re-read on every turn. So GPT-2 choking on long conversation histories in 2020 and Claude going fuzzy on my twelfth file in 2025 were the same problem wearing different clothes.
Why Attention Gets Expensive
Self-attention computes a pairwise interaction between every token and every other token in the window. A thousand tokens means a million computations. Ten thousand tokens means a hundred million. Double the input, quadruple the cost. That's O(n²). Picture a networking event. Ten people in a room, everyone shakes hands with everyone: 45 handshakes. Fine. Ten thousand people: close to 50 million handshakes. You're not networking at that point, you're standing in line. People have tried patching around it. Sparse attention (Longformer, BigBird) says shake hands with your neighbors and a few VIPs. Linear attention (Performer) approximates the handshakes with math. FlashAttention makes them faster with smarter GPU memory use. None of it changes the underlying problem. Attention wants to look at everything, and everything gets big. Then there's the KV cache. Every token you generate needs stored keys and values for all previous tokens. At 256K context, Mixtral needs 32 GB just for the cache, before model weights load.
Watching It Break on My Own GPU
I wanted to see the curve myself. Free Colab notebook, T4 with 16 GB, GPT-2 Small loaded through HuggingFace. Same model family I'd benchmarked back in the startup days, except this time I was profiling the bottleneck instead of evaluating a product. Copilot helped me throw the script together fast. | Sequence Length | Time (sec) | Peak GPU Memory | |-----------------|------------|-----------------| | 512 | 0.02 | 0.8 GB | | 1,024 | 0.04 | 1.1 GB | | 2,048 | 0.11 | 2.0 GB | | 4,096 | 0.38 | 5.2 GB | | 8,192 | 1.41 | 13.1 GB | | 16,384 | OOM | > 16 GB | Textbook quadratic. 512 to 1,024 roughly doubled the time. 1,024 to 2,048 nearly tripled it. By 4,096 the attention layer was eating 5 GB. At 8,192 the T4 was grinding through 13 GB and over a second of latency for one forward pass. At 16,384 it quit. Out of memory. > [!IMPORTANT] This is GPT-2 Small. 124 million parameters, a toy by current standards. Now picture a 7B model at 128K context. The quadratic wall isn't theoretical. It's why your GPU bill looks like that. The model was smart enough. It just couldn't hold a long conversation without drowning in compute. So I went looking for architectures that scale better. That week I found a paper by Albert Gu and Tri Dao, two names I recognized from FlashAttention: "Mamba: Linear-Time Sequence Modeling with Selective State Spaces." I read the abstract twice and then spent three days on the rest of it.
State Space Models, the Idea Underneath
Mamba didn't appear from nowhere. It's the end of a long line of state space model research, and the lineage explains the design choices. Think about listening to a song. You don't store every audio sample you've ever heard. Your brain keeps a compressed state of what matters, updates it as new sound arrives, and you hear music. That's roughly what an SSM does. It keeps a hidden state h, updates it with each input, and reads an output off that state: Looks like a fancy RNN, and you're not wrong. But here's the trick: during training, because A, B, and C stay constant, the whole recurrence unrolls into a convolution. Convolutions parallelize on GPUs. You get RNN-style cheap inference and CNN-style parallel training. The breakthrough here was S4 (Structured State Spaces for Sequence Modeling) from Albert Gu et al., published at ICLR 2022. It dominated Long Range Arena, handling sequences past 16,000 elements where Transformers struggled. Great for audio, time series, sensor data. Language was the problem. S4 treated every token identically. It couldn't decide what to keep and what to drop based on what it was actually reading. Content-blind. > [!TIP] Think of S4 as a security camera recording everything at the same resolution all day. It captures the footage but never zooms in on the moment that matters.
What Mamba Changed
Mamba (December 2023) fixes that with one move: make the SSM parameters depend on the input. In S4, A, B, and C are learned during training and then frozen. Every token gets identical treatment. Mamba turns three of them into functions of the current token. - Δ (step size) controls forgetting. Large Δ means update aggressively and let old information decay. Small Δ means hold the current state, this input isn't worth much. It's a learned gate, close cousin to the LSTM forget gate, except it falls out of the continuous-time formulation naturally. - B controls how much the input influences the state. - C controls what gets read out. The paper shows this on a selective copying task: given a sequence with marked tokens, copy only those. Fixed SSMs fail outright. Mamba handles it. On induction heads, a pattern-completion task, it generalizes to sequences over 4,000 times longer than what it trained on. This selectivity is why the core layer is also called S6, meaning S4 plus selection plus scan. Mamba is the full architecture built around that layer. The real difference between the two comes down to compression. A Transformer keeps the whole sequence in its KV cache. Full access, no compression. Mamba squeezes everything into a fixed-size state vector. That's its biggest strength and its biggest weakness at the same time. As the paper puts it: "A fundamental problem of sequence modeling is compressing context into a smaller state."
The Hardware Trick
Making parameters input-dependent breaks the convolution trick. You can't precompute a kernel that changes every step. So Mamba falls back to a sequential scan. Doesn't that make it slow like an RNN? It would, if Gu and Dao hadn't built a hardware-aware parallel scan around it. This is where Tri Dao's FlashAttention instincts show up. GPUs have two kinds of memory. HBM is the main pool, big (40 to 80 GB on an A100 depending on variant) but slow to reach. SRAM is the on-chip cache, tiny and fast. Most operations other than matrix multiplication are memory-bound, meaning the bottleneck is moving data between the two, not the math. Picture HBM as a warehouse across town and SRAM as the bench in front of you. Your build speed isn't the problem. The number of trips is. Mamba does three things about it: 1. Kernel fusion. It fuses discretization, scan, and output into one GPU kernel that lives in SRAM, killing most of the memory traffic. 2. Parallel scan. It parallelizes the recurrence with a Blelloch scan, processing elements tree-style instead of one at a time. 3. Recomputation. It skips saving intermediate states on the forward pass, recomputing them during backprop instead, trading a little compute for a lot of memory. The result: the fused selective scan beats FlashAttention-2 past sequence length 2K and runs up to 40 times faster than a naive scan. > [!TIP] If FlashAttention's IO-aware design made sense to you, Mamba will feel familiar. Same author, same philosophy. Build the algorithm around the memory hierarchy instead of the other way around.
The Numbers
| Metric | Mamba | Transformer | |--------|-------|-------------| | Attention complexity | O(n) linear | O(n²) quadratic | | Inference memory per step | Constant (fixed state) | Growing (KV cache) | | Inference throughput (same size) | 5x higher | Baseline | | Mamba-2.8B quality | ≈ Transformer-7B | Baseline | | MMLU 5-shot (8B, pure, 1.1T) | 29.2% (Mamba-2) | 46.3% | | MMLU 5-shot (8B, hybrid, 3.5T) | 53.6% | 50.1% | | KV cache at 256K context | 4 GB (original Jamba) | 32 GB (Mixtral) | | Prefill + decode (16K, H100) | 140 seconds (Mamba-3) | 976 seconds (Llama-3.2-1B) | A few of these deserve a sentence. Mamba-2.8B outperforms Pythia-6.9B on average zero-shot accuracy, so roughly double the parameters for the same quality. Inference throughput lands around 5x a same-size Transformer, since each step costs constant time and constant memory with no cache growing underneath it. On genomics, using the HG38 human genome dataset, Mamba's perplexity keeps improving out to a million tokens of context. For species classification across great apes sharing 99% of their DNA, it beats everything else tested. On the SC09 speech benchmark it cuts FID by more than half against prior models, and keeps improving on minute-long sequences.
Where Transformers Still Win
Few-shot prompting is the weak spot. NVIDIA's 8B study found pure Mamba-2 scoring 29.2% on MMLU 5-shot early in training (1.1 trillion tokens) against the Transformer's 46.3%. With 3.5 trillion tokens the gap narrowed a lot, but pure Mamba-2 still trailed. Exact recall is the other one. Mamba compresses everything into a fixed-size state, which is like trying to remember a book from one page of notes. Transformers keep the book open through the KV cache. On a phonebook lookup, where you need one specific fact from one specific position, that difference is stark. Which loops back to where I started. When Claude lost track of my first file, that was the KV cache straining. But it could still look back at those tokens. Pure Mamba wouldn't have them at all. And that 2020 GPT-2 experiment ran into a 1,024-token window so small that a normal multi-turn chat filled it. Combined with the cost, switching never made sense. The ceiling was real, I just couldn't name it. > [!WARNING] Don't reach for a pure Mamba model when the task needs precise recall from a long context. A fixed-size state is a lossy compressor. It remembers patterns, not phone numbers.
Mamba-2, or When SSMs Turned Out to Be Attention
May 2024, Dao and Gu published a paper titled "Transformers are SSMs." Mamba-2 introduced the Structured State Space Duality framework, and the finding is mathematical: selective SSMs and linear attention are equivalent. Both are multiplication by semiseparable matrices. Two views of one operation. That's not just theory. The duality enables a chunked algorithm: quadratic attention-style compute inside each chunk, linear SSM-style scans connecting them. Result is 2 to 8 times faster training than Mamba-1, widening with sequence length. Mamba-2 also added multi-head structure and much bigger state dimensions, N=64 to N=256 and up, against Mamba-1's N=16. Mamba-3 landed in March 2026 with complex-valued SSMs, exponential-trapezoidal discretization, and a MIMO architecture. It gets roughly 4% relative improvement over the Transformer baseline on standard language benchmarks, and matches Mamba-2's perplexity at half the state size, which doubles inference throughput. On an H100 at 16K sequence length it finishes prefill and decode in 140 seconds against Llama-3.2-1B's 976 seconds on vLLM. That's probably the strongest showing yet for a pure SSM against a well-tuned Transformer under controlled conditions. Worth noting the Mamba-3 authors themselves expect "linear layers will predominantly be used alongside global self-attention layers for language modeling."
Hybrids Are Winning
Here's where the research actually landed, and it's a satisfying answer. Use Mamba layers for most of the sequence processing, then drop in a few attention layers at the right depths for retrieval and in-context learning. The ratio that keeps showing up is around one attention layer for every seven to nine Mamba layers. Jamba (AI21, 2024) was the first at scale, 1:7 attention to Mamba plus Mixture of Experts. Jamba-1.5-Large runs 398B total parameters, 94B active, 256K context, and beats both pure architectures. NVIDIA's Nemotron 3 Super (March 2026) is 120B total with 12B active, 75% Mamba-2 layers and 25% attention, native 1M token context, 5x the throughput of the previous generation. Perplexity, Cursor, and Palantir are already using it. IBM's Granite 4.0 (October 2025) went 9:1 Mamba to Transformer with fine-grained MoE, over 70% memory reduction, and became the first open model family with ISO 42001 certification. The clearest data point is NVIDIA's own study. Their Mamba-2-Hybrid ran 56 layers: 24 Mamba-2, 4 attention, 28 MLP. Trained on 3.5 trillion tokens, it beat the Transformer on all 12 short-context benchmarks and hit 53.6% on MMLU 5-shot against 50.1%. Four attention layers were enough to close the in-context learning gap and push past a pure Transformer while keeping most of the efficiency. | Model (March 2026) | Architecture | Context | Key Feature | |--------------------|--------------|---------|-------------| | Jamba 1.5 Large | 87.5% Mamba + 12.5% Attention | 256K | First large-scale hybrid | | NVIDIA Nemotron 3 Super | 75% Mamba-2 + 25% Attention | 1M | 5x throughput, in production | | IBM Granite 4.0 | 90% Mamba + 10% Attention | 512K | 70% memory reduction | | Mamba-3 (pure) | 100% SSM | Flexible | Strongest pure SSM result |
The Angle Nobody's Working On
One more connection, and it's the one closest to what I work on day to day. Speculative decoding runs a small fast draft model to generate several candidate tokens, then a large accurate verifier checks them all in parallel. When the drafts are right, which they often are for predictable text, you skip multiple serial decode steps. When they're wrong, you fall back. Faster inference, no quality loss. The draft model is the bottleneck. It has to be cheap enough that drafting beats just running the verifier token by token. Mamba is almost perfect for this. Constant time per token, constant memory, no KV cache growing under it. A small Mamba drafting 5 to 8 candidates costs almost nothing next to a small Transformer doing the same job. And the compression weakness stops mattering, because the Transformer verifier catches the mistakes. A Mamba drafter with a Transformer verifier gives you linear-time speed on the easy tokens and full attention precision on the hard ones. That's a clean division of labor, and it's the kind of thing vLLM and SGLang could support at the serving layer rather than requiring anyone to retrain an architecture.
Try It Yourself
Fastest path from zero to Mamba inference. Runs on a free Colab T4. Notice what's absent compared to a Transformer. No KV cache to manage, no attention mask, no position embeddings. The model holds a fixed-size state internally. That's the whole point. Now profile it the way we profiled GPT-2 and watch the memory stay flat: Barely moves. That's constant inference memory, next to the quadratic curve we saw earlier. > [!TIP] mamba-ssm needs CUDA and Linux, so Colab works. For a pure PyTorch version with no custom kernels, mamba-minimal by John Ma on GitHub reimplements the whole thing in about 200 readable lines.
What I'd Tell 2020 Me
If I could sit next to myself while I was benchmarking GPT-2 against those word embeddings, watching short messages improve and long ones collapse, I'd say this: the model isn't broken. The architecture has a ceiling. Give it a few years and people will build around it. For anyone working on inference optimization, Mamba is worth studying regardless of whether you ever deploy one. The hardware-aware scan, the kernel fusion, the memory hierarchy thinking all transfer to making Transformer inference faster. Tri Dao built both FlashAttention and Mamba, and the thinking is the same. If you're shipping to production, watch the hybrids. NVIDIA, IBM, and AI21 are all there now, and vLLM and SGLang support will follow. Speculative decoding with a Mamba drafter might be the cheapest win available without retraining anything. For long-context work, hybrids get you past 128K without the KV cache eating your GPU. Original Jamba handles 256K on 4 GB where Mixtral needs 32 GB. And for audio, genomics, or time series, pure Mamba is already the better call. Long sequences, continuous signals, not much need for exact token recall.
Timeline
| Year | Milestone | |------|-----------| | 2020 | HiPPO — memory framework for SSMs | | 2022 | S4 — structured state spaces, dominates Long Range Arena | | 2023 | H3 — bridges S4 toward language tasks | | Dec 2023 | Mamba (S6) — selective SSMs, first competitive with Transformers on language | | May 2024 | Mamba-2 — SSD framework, 2-8x faster training, duality with attention revealed | | Aug 2024 | Jamba 1.5 — 94B active hybrid at 256K context | | Aug 2024 | Falcon Mamba 7B — best pure Mamba at scale | | Oct 2025 | IBM Granite 4.0 — production hybrid with ISO certification | | Mar 2026 | NVIDIA Nemotron 3 Super — 120B hybrid, 1M context, 5x throughput | | Mar 2026 | Mamba-3 — strongest pure SSM result on language benchmarks |
Wrap Up
This started six years ago with a chatbot where I benchmarked GPT-2 against static word embeddings, watched it improve on short inputs and fall apart on long ones, and decided the tradeoff wasn't worth it without understanding the actual reason. It picked back up last year with a code paste that went fuzzy halfway through. Same problem, same architecture, same quadratic wall. The difference now is that I know why. Mamba isn't a Transformer killer. It's a complement. The quadratic bottleneck is real and it hurts more as contexts stretch into the hundreds of thousands, and Mamba's linear scaling and constant inference memory solve that cleanly. But exact recall and in-context learning from uncompressed context still belong to Transformers. > [!IMPORTANT] Nobody wins outright. A few attention layers for precision, a lot of Mamba layers for efficiency, and maybe a fast Mamba drafter feeding a careful Transformer verifier at the serving layer. The design space just got a lot bigger. Not bad for a model named after a snake.
References
Papers and Models > > 1. Gu & Dao, "Mamba: Linear-Time Sequence Modeling with Selective State Spaces" (2023) > 2. Dao & Gu, "Transformers are SSMs" (Mamba-2, ICML 2024) > 3. Gu et al., "Mamba-3: Improved Sequence Modeling using State Space Principles" (ICLR 2026) > 4. Gu et al., "Efficiently Modeling Long Sequences with Structured State Spaces" (S4, ICLR 2022) > 5. Waleffe et al., "An Empirical Study of Mamba-based Language Models" (NVIDIA, 2024) > 6. Lieber et al., "Jamba: A Hybrid Transformer-Mamba Language Model" (AI21, 2024) > 7. NVIDIA, "Introducing Nemotron 3 Super" (March 2026)