Releasing Benign Trajectories for MonitoringBench
Can Küçükkurt
Abstract
Authors: Can Kucukkurt, Ethan Elasky Palaestra Research TL;DR: We are releasing MonitoringBenchHonest , a supplement to the original attack trajectory dataset MonitoringBench used by both Anthropic and OpenAI in monitoring settings. Our set is intended for use in more expensive tasks, rather than the cheaper bulk monitoring setting where there is a large acceptable number of false positives. We control for tasks, format, and model, and release 550 benign trajectories alongside 550 paired attack trajectories from the original dataset. What is MonitoringBench? MonitoringBench is a difficulty-graded benchmark of 2,644 successful attack trajectories for evaluating coding-agent monitors. The attacks are generated through the process of strategy generation, execution, and post-hoc refinement, and are substantially more difficult for models to detect than prompt elicitation only of frontier attackers. … MonitoringBench was used for monitor evaluations of internal Mythos deployment (see Anthropic's Mythos risk report) and OpenAI's Auto-Review . Why honest trajectories matter Despite a significant portion of status quo control literature focusing on step monitoring, and despite status quo c