other8.0 / 10
A sketch of an AI control safety case
Abstract
Partnership with UK AISI to describe how developers can construct structured arguments that models cannot subvert control measures
Research area
agent foundationsagentic misalignmentcontrol theory
Published
—
Source
other
Org
Redwood Research
Sign in to read and join the discussion.