Why does a model reread the whole conversation instead of just remembering it?
the paper →Mamba: Linear-Time Sequence Modeling with Selective State SpacesGu & Dao, 2023 ↗The selection mechanism is the contribution; the hardware-aware scan in section 3.3 is what made it trainable. Read both or neither.A fixed memory can only keep what it chose to keep.
Every reply a transformer writes begins by reading the whole conversation again. You do not do that. You keep a summary in your head, update it as each new thing is said, and never rewind. Mamba is a model built that way: a fixed amount of memory, revised once per word, that never grows. It makes the notebook unnecessary — and it pays for that with everything the notebook was keeping.
On the machineThis lands on HBM, station 3 of 10 on the path the constraint took through the hardware. Removes the occupant that grew with every token and leaves one that never does. See it on the drawing →Read as far as you want. Each level assumes the one above it and nothing more.
A state space model carries a fixed-size state from one token to the next and updates it once per token, the way a recurrent network does. What makes Mamba work where earlier recurrent models did not is that the update depends on the input: the model decides, per token, what to write into the state and what to let decay. That choice costs the trick that made older state space models fast to train, so it is bought back with a parallel scan that keeps the state in SRAM. The result is a model whose memory per sequence is a constant, at any length.