Back to papers
other8.0 / 10

A sketch of an AI control safety case

Abstract

Partnership with UK AISI to describe how developers can construct structured arguments that models cannot subvert control measures

Research area

agent foundationsagentic misalignmentcontrol theory
Published
Source
other
Org
Redwood Research
Sign in to read and join the discussion.