arxiv8.0 / 10
Model tampering attacks enable more rigorous evaluations of LLM capabilities
Abstract
Research paper by MATS scholars. Authors: Zora Che
Research area
post-training
Published
—
Source
arxiv
Org
MATS
Sign in to read and join the discussion.