lesswrong8.0 / 10
AI in 2025: Gestalt (Annual Safety Review)
Zvi Mowshowitz
Abstract
Comprehensive annual review of AI capabilities and safety progress in 2025. Covers evaluation awareness (Claude Sonnet 4.5 verbalizing eval awareness 58% of the time), reward hacking, deceptive alignment, lab safety plans, and the gap between safety and capability progress. Influential synthesis for community orientation.
Research area
alignmentdeceptionpost-training
Published
2025
Source
lesswrong
Org
Alignment Forum
Sign in to read and join the discussion.