Back to papers
arxiv8.0 / 10

Model tampering attacks enable more rigorous evaluations of LLM capabilities

Abstract

Research paper by MATS scholars. Authors: Zora Che

Research area

post-training
Published
Source
arxiv
Org
MATS
View paper
Sign in to read and join the discussion.