Back to papers
lesswrong8.0 / 10

TASTE: Can AI Models Judge AI Safety Research Proposals?

Hasan Baig

Abstract

tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confidence. We find models perform worse than human researchers on TASTE (Fable 5, 60%). 📝 Blog , 📄 Paper This work was done as part of the Anthropic Fellows Program . Background While some aspects of AI safety research are relatively straightforward to measure , progress on many questions in AI safety cannot be evaluated with verifiable rewards. For instance, research into mitigating risks from AI misalignment often involves forecasting risks posed by future AI systems. Another example is detecting when models are deceptive , which depends on the difficult task of accurately attributing beliefs and intentions to models. If we want to automate AI safety research — which might become necessary if

Research area

ai safetybenchmarksevaluations
Published
28 Aug 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.