Back to papers
lesswrong8.0 / 10

Eliciting hidden knowledge from monitors with NLAs

David Africa

Abstract

Aleksandr Bowkis* and David Africa* TL;DR Chain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface. We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging an agent's trajectory. NLAs can be useful for monitoring in two ways: Monitor-side: Eliciting latent capabilities from weak monitors by surfacing unverbalised knowledge of reward hacking. Agent-side: An agent’s NLA readouts can be self incriminating. NLAs surface a monitor’s knowledge of reward-hacking better than direct verbalised judgements for some of our datasets. Monitor-side NLAs are less useful for recovering unverbalised knowledge of reward hacking than inspecting a monitor’s CoT. There is some decorrelation between monitor-side NLAs and CoT, suggesting a combined approach may be most effective at eliciting monitor capabilities. Introduction Chain of thought monitorability is very important to ensuring safe deployment of powerful AI. It is also fragile ; [1] both OpenAI and Anthropic have accidentally trained against chain of thought; various research grou

Research area

mechanistic interpretabilitymonitoringreward hacking
Published
15 Jul 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.