Back to papers
other8.0 / 10

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Abstract

Testing effectiveness of defense-in-depth AI safety strategies using STACK attack method, achieving 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0%

Research area

deceptionpost-trainingrobustness to domain shifts
Published
Source
other
Org
FAR AI
Sign in to read and join the discussion.