Back to papers
lesswrong8.0 / 10

2B scoring model flags out-of-domain misalignment, suggesting specialist judges have potential for audits

burnssa

Abstract

Some evidence that narrow ‘specialist’ models could be useful as part of deployed model misalignment audits, complementing larger frontier auditing agents and offering potential cost, discrimination and transparency benefits. A Gemma 2B specialist judge trained on Betley et al 2025b (‘Betley’) code examples is able to distinguish between responses provided by insecure-fine-tuned ‘misaligned’ and ‘secure-fine-tuned’ models responding to out-of-domain general safety prompts (‘ICEBERG’ - details below), where prompted Sonnet 4.5 judgments on the same prompts are not discriminative Similar tests on SecurityEval code prompts (out-of-distribution but in-domain) show the specialist judges track alignment drift and improve on Sonnet's paired discrimination (winrate of ~66% vs ~64%), but comparisons are partly confounded by pattern-matching Betley training data attributes The results tentatively indicate promise for extending narrow classifiers now often used in safety contexts (e.g. Anthropic classifiers for pretraining filtering and jailbreak protection ) to deployed model audits, but more work is required to prove specialist judges' value Replication guide here Scalable auditing of

Research area

alignmentevaluationmodel infrastructure
Published
14 May 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.