AI & Computational Science

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

How the science connects

Language modelRepresentation lea…

AI Insight

This study demonstrates that linear probes used to detect truthfulness in language models can be fooled by a "perfect aliasing" problem, where the probe cannot distinguish between the model representing truth versus simply providing the task-prescribed answer. In controlled experiments with Gemma-2-9B and Llama-3.1-8B models trained to deceive in specific contexts, ally-fitted probes incorrectly suggested the model had stopped representing truth (scoring near 0% accuracy), while probes trained on mixed contexts recovered the true information perfectly (100% accuracy). The research shows that low probe scores on deceptive models do not necessarily indicate the model has hidden or stopped internally representing the truth.


This finding has significant implications for AI safety monitoring, as it reveals that commonly proposed truth-probing methods for detecting deception in language models may produce misleading results. The work highlights the need for more robust methods to evaluate whether AI systems are genuinely concealing information versus artifacts of probe design.


Understand the Science

Language model 27 articles Explore Concept → Representation learning Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: Linear probes that decode the truth from a language model’s activations have been proposed as deception monitors, and a truth probe scoring below chance on a model trained to deceive is naturally read as evidence that the model hid or stopped representing the truth. We show that a fitting-label ambiguity can produce the same readout. In a controlled game, a model should report a secret bit to an ally and its complement to a rival. On compliant (ally) contexts the true bit and the answer the task prescribes are identical labels, so a probe fitted there cannot tell which of the two it measures. We call this complete agreement perfect aliasing. The labels are complements on rival contexts, so one ally-fitted probe scored against each has rival AUROCs that sum to one. Mixed fitting, on ally and rival contexts together, makes the two labels differ; randomized output codebooks also decouple the prescribed answer from the output letter. For a reward-trained Gemma-2-9B policy that answers falsely on every evaluated rival trial, the ally-fitted probe scores $0.006 pm 0.005$ AUROC at the final layer (mean $pm$ sample SD over three RL training seeds), while mixed-fit probes score 1.000 on the same held-out activations. Mixed fitting also uses more examples and labelled rival data, so this shows the true bit remains linearly recoverable, not that separating the labels alone explains the gain. In instructed Llama-3.1-8B, a probe refitted on one prompt variant and one frozen from a reference variant both have held-out ally accuracy 1.000 but rival truth AUROCs of 0.080 and 0.986 on the same trials. Our main experiments use one single-token game that states the bit in the prompt, so the recovered direction may read a retained copy of it; where the model must infer the bit, no tested arm deceives reliably. We study what a probe measures, not whether the model uses that information.

Source: The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes