lesswrong9.0 / 10
AI Safety at the Frontier: Paper Highlights (Monthly Series, 2025–2026)
AISafety.com editorial team
Abstract
Monthly curated digest of the most important AI safety papers, with detailed summaries and analysis. Covers sandbagging, emergent misalignment, cyber capabilities, control evaluations, and more. Highly influential community resource for tracking the frontier of safety research.
Research area
agent foundationsmodel robustnessrobustness to domain shifts
Published
2025
Source
lesswrong
Org
Alignment Forum
Sign in to read and join the discussion.