other8.0 / 10
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Abstract
Alignment research on suppressing unwanted traits in LLMs through inoculation prompting
Research area
alignment
Published
—
Source
other
Org
UK AI Safety Institute
Sign in to read and join the discussion.