lesswrong8.0 / 10
AI for AI Safety
Joe Carlsmith
Abstract
Fourth essay in Carlsmith's 'How do we solve the alignment problem?' series. Argues for using frontier AI labor to strengthen safety progress, risk evaluation, and capability restraint. Critically examines elicitation/evaluation failures, differential sabotage risks, and the viability of automated alignment research as a strategy.
Research area
alignmentevaluationgovernance
Published
2025
Source
lesswrong
Org
Alignment Forum
Sign in to read and join the discussion.