Back to papers
lesswrong9.0 / 10

AI Safety at the Frontier: Paper Highlights (Monthly Series, 2025–2026)

AISafety.com editorial team

Abstract

Monthly curated digest of the most important AI safety papers, with detailed summaries and analysis. Covers sandbagging, emergent misalignment, cyber capabilities, control evaluations, and more. Highly influential community resource for tracking the frontier of safety research.

Research area

agent foundationsmodel robustnessrobustness to domain shifts
Published
2025
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.