AI Insight
This study demonstrates that latent communication systems in multi-agent AI networks, which exchange information through internal representations rather than text, can be exploited to increase harmful compliance even when individual agents remain safety-aligned. Researchers developed reinforcement learning attacks that manipulate communication links between agents, raising harmful compliance scores from 27.9 to 76.9 across multiple safety benchmarks while maintaining performance on benign tasks. The work also shows that these compromised links can be repaired by adapting rewards toward safer behavior without modifying the underlying agents.
Why it matters
As AI systems increasingly operate in multi-agent configurations, this research reveals a critical vulnerability where communication pathways between safe agents can be weaponized to bypass safety measures. The findings suggest that safety alignment must evaluate entire multi-agent systems rather than individual components, with important implications for deploying networked AI in sensitive applications.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Abstract: Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender’s representations into the receiver’s input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query–response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety
Source: Safety of Latent Communication in Multi-Agent Systems