Back to papers
lesswrong8.0 / 10

AI for AI Safety

Joe Carlsmith

Abstract

Fourth essay in Carlsmith's 'How do we solve the alignment problem?' series. Argues for using frontier AI labor to strengthen safety progress, risk evaluation, and capability restraint. Critically examines elicitation/evaluation failures, differential sabotage risks, and the viability of automated alignment research as a strategy.

Research area

alignmentevaluationgovernance
Published
2025
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.