Toy Model of Activation Obfuscation
Jesse Li
Abstract
I completed this work as part of the BlueDot Impact Technical AI Safety Project . This linkpost is a somewhat condensed version of the writeup on my blog. Training against probes is considered a forbidden technique , because the model might learn to obfuscate its activations instead of behaving better. Can we create a toy example of this? More specifically: under optimization pressure, will a toy model learn to encode a feature to be challenging to detect with linear probes? In this research, I give theoretical and empirical evidence that models can and will defeat adversarially-trained linear probes, at least in some configurations. I investigated this problem with a toy residual-stream MLP architecture: The model tries to learn mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width=