Back to papers
other8.0 / 10

Unelicitable Backdoors in Language Models

Abstract

Novel class of unelicitable backdoors in transformer models that can evade detection by conventional cybersecurity monitoring systems.

Research area

agentic misalignment
Published
Source
other
Org
Palisade Research
Sign in to read and join the discussion.