other8.0 / 10
Unelicitable Backdoors in Language Models
Abstract
Novel class of unelicitable backdoors in transformer models that can evade detection by conventional cybersecurity monitoring systems.
Research area
agentic misalignment
Published
—
Source
other
Org
Palisade Research
Sign in to read and join the discussion.