Back to papers
arxiv8.0 / 10

An Approach to Technical AGI Safety and Security

Shane Legg, Jan Leike, Rohin Shah, Victoria Krakovna et al. (Google DeepMind)

Abstract

Comprehensive 145-page DeepMind report laying out their technical approach to AGI safety. Covers alignment, interpretability, evaluation, robustness, and security. One of the most significant org-level safety documents of 2025 alongside Anthropic's RSP.

Research area

agent foundationsalignmentrobustness to domain shifts
Published
2025
Source
arxiv
Org
Google DeepMind
View paper
Sign in to read and join the discussion.