Your AIs don't do what you want. This is really bad
Kaustubh Kislay
Abstract
Replit AI deletes entire database during code freeze, then lies about it — a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym , a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substantial compute looking for a way out of the isolated evaluation environment rather than solving the problem at hand found and exploited a previously unknown zero-day in third-party software OpenAI used as a proxy and cache for package registries used it to get unrestricted internet access chained stolen credentials and several vulnerabilities into a remote code execution path on Hugging Face’s servers reached the ExploitGym solutions sitting in Hugging Face’s production database This is one of the most egregious examples of reward hacking in the wild, if not the most. Without special care, this will come to be one of the least egregious. The internet is full of complaints about