AI Insight
This paper introduces a "design-model framework" for building efficient recurrent neural network layers by explicitly modeling memory through Bayesian filtering principles. The authors demonstrate that their Bayesian Layer, which propagates both mean and covariance to track uncertainty in stored information, unifies and improves upon several existing sub-quadratic sequence models including linear attention, GLA, Mamba-2, and DeltaNet. Experiments show that incorporating covariance tracking improves associative recall, long-context retrieval, and robustness in out-of-distribution scenarios, though with modest perplexity costs in some settings.
Why it matters
This work provides a principled theoretical foundation for understanding and improving recurrent sequence models, which are crucial for efficient processing of long sequences in language models and other applications. The framework could enable better memory management in AI systems while reducing computational costs compared to standard attention mechanisms.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: We introduce the emph{design-model framework}: a way to derive efficient recurrent sequence maps from explicit assumptions about memory. A design model writes evidence into memory by exact Bayesian filtering; a query- dependent readout produces a predictive distribution whose mean is the layer output. In our linear-Gaussian instantiation, the emph{Bayesian Layer} propagates both a mean and a covariance: the covariance tracks uncertainty over stored associations, steering writes toward uncertain directions, attenuating gains as evidence accumulates, and preserving confident memories. The same framework unifies several sub-quadratic recurrences: linear attention, GLA, and Mamba-2/SSD are exact filters under a latent-input design model, whereas DeltaNet and related Delta-rule models are covariance-reset reductions of the Bayesian Layer’s design model. Restoring covariance propagation yields closed-form predictions for retrieval dynamics, which we verify empirically, and improves robustness beyond the training regime in controlled collision studies, learned associative recall, and the Zoology MQAR benchmark. Training from scratch on WikiText-103 under matched state budgets lowers perplexity on associative-recall hits. Distilling Bayesian Layers into a pretrained 340M Gated DeltaNet improves RULER long-context retrieval over a matched-compute control, at a 2.5–2.7% held-out perplexity cost.