AI & Computational Science

AI Chatbots’ Safety Training May Be More Fragile Than Expected

How the science connects

Machine learningLarge language modelAI safety

AI Insight

Researchers developed SKIN-DEEP, a geometric diagnostic tool that analyzes the internal representations of large language models to predict how stable their safety behaviors will be after updates or fine-tuning. By comparing aligned and base model versions across twenty-one instruction-tuned models, they found that harmful requests create distinct low-rank separation patterns in the model's activations. Models with lower Geometric Fragility Scores before fine-tuning maintained better safety performance after being trained on benign examples, suggesting that representation geometry can predict susceptibility to safety degradation.


This work provides a proactive method to assess AI safety robustness before deployment or updates, rather than discovering vulnerabilities only after interventions occur. The diagnostic could help developers identify models at risk of losing safety alignment during routine fine-tuning, potentially preventing the deployment of systems that appear safe initially but degrade easily.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an intervention. We introduce SKIN-DEEP, a geometric diagnostic that examines the unmodified model’s residual-stream activations. It compares aligned and base checkpoints to identify safety-separating directions, tests their behavioral relevance through ablation, and summarizes the layer-wise pattern in the Geometric Fragility Score (GFS). Across twenty-one instruction-tuned models, harmful requests and benign instructions exhibit a recurring low-rank separation pattern. Selected direction ablations weaken refusal, with the effective direction varying across models. In benign low-rank fine-tuning experiments, the initially safe model with the lowest score before fine-tuning has the lowest harmful-compliance rate when trained on the largest tested set of harmless examples. These findings connect representation geometry to subsequent behavioral susceptibility and support activation-based diagnostics as a complement to refusal tests. Our code is available at https://github.com/js-lee-AI/skin-deep.

Source: Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations