other8.0 / 10
Exploiting Novel GPT-4 APIs
Abstract
Red-teaming GPT-4 fine-tuning, function calling and knowledge retrieval APIs, finding that fine-tuning on as few as 15 harmful examples can remove core safeguards
Research area
deceptionfine-tuningpost-training
Published
—
Source
other
Org
FAR AI
Sign in to read and join the discussion.