Back to papers
arxiv9.0 / 10

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov

Abstract

As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.

Research area

ai securityjailbreakingred-teaming
Published
18 Aug 2026
Source
arxiv
Org
Basic Research of Artificial Intelligence Laboratory
View paper
Sign in to read and join the discussion.