Back to papers
arxiv7.0 / 10

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez

Abstract

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.

Research area

machine unlearningreward hackingreward specification
Published
18 Aug 2026
Source
arxiv
Org
University of Valencia
View paper
Sign in to read and join the discussion.