Back to papers
lesswrong8.0 / 10

AI in 2025: Gestalt (Annual Safety Review)

Zvi Mowshowitz

Abstract

Comprehensive annual review of AI capabilities and safety progress in 2025. Covers evaluation awareness (Claude Sonnet 4.5 verbalizing eval awareness 58% of the time), reward hacking, deceptive alignment, lab safety plans, and the gap between safety and capability progress. Influential synthesis for community orientation.

Research area

alignmentdeceptionpost-training
Published
2025
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.