Back to papers
lesswrong7.0 / 10

J-space auditing might be unreliable

Kartikay Luthra

Abstract

Across these preliminary experiments, decoded J-space did not seem particularly informative about reward-hacking behaviour. The readouts remained substantially similar across checkpoints and monitoring conditions despite meaningful behavioural differences, and providing J-space to an LLM auditor produced little additional discrimination beyond the information already available from the task or transcript. These results are limited to one model family, one model organism, and one behavioural setting, so I treat them as motivation for further stress-testing. This was originally produced as part of application to Neel Nanda's MATS research stream; the scope and depth reflect that constraint, and Future Work outlines where I'd take it with more time . Introduction J-space or global workspace was introduced by Anthropic's recent work which showcases the intermediate tokens a model uses for it's computation, my original thought when i read this paper was what kind of information can the intermediates disclose about the underlying model's policies which guides its decision-making under some context, by policy I roughly mean will J-space tokens help me in inferring the intention behind a m

Research area

evaluationsmechanistic interpretabilityreward hacking
Published
16 Sept 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.