Back to papers

Can we steer AI models toward safer actions by making these instrumentally useful?

Francesca Gomez

Abstract

An empirical study adapting and testing insider risk mitigations for Agentic Misalignment

Research area

behavioural boundariesinference-time controlssteering controls for ai
Published
22 Oct 2025
Source
other
Org
Wiser Human
View paper
Sign in to read and join the discussion.