AI & Computational Science

AI Models Learn Faster by Selectively Sharing Information Across Networks

How the science connects

Machine learningRecurrent neural n…Federated learning

AI Insight

This study examines how modern state space models (SSMs) like Mamba2 perform in distributed federated learning environments, where multiple clients train models on decentralized data. The researchers derive mathematical bounds that account for SSM-specific properties such as recurrent stability and input-dependent discretization, then validate these theoretical predictions through experiments comparing nine federated learning algorithms across six text domains. The analysis reveals how architectural characteristics of SSMs uniquely affect the convergence and optimization behavior in federated settings compared to architecture-agnostic approaches.


This work bridges the gap between emerging SSM architectures and federated learning systems, which is crucial for deploying these efficient models in privacy-preserving applications where data cannot be centralized, such as healthcare records or personal device data. The architecture-aware approach could lead to more efficient distributed training protocols specifically optimized for SSMs.


Understand the Science

Machine learning 325 articles Explore Concept → Recurrent neural network Concept coming soon Federated learning Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.

Source: Distributed Learning with Selective State Space Models: Architecture-Aware Convergence Analysis