Back to papers
arxiv8.0 / 10

Badllama 3: removing safety finetuning from Llama 3

Abstract

We show that extensive LLM safety fine-tuning is easily subverted when an attacker has access to model weights. We evaluate three state-of-the-art fine-tuning methods-QLoRA, ReFT, and Ortho-and show how algorithmic advances enable constant jailbreaking performance with cuts in FLOPs and optimisation power. We strip safety fine-tuning from Llama 3 8B in one minute and Llama 3 70B in 30 minutes on a single GPU, and sketch ways to reduce this further.

Research area

agentic misalignmentmodel robustnessrobustness to domain shifts
Published
Source
arxiv
Org
Palisade Research
View paper
Sign in to read and join the discussion.