Back to papers
lesswrong7.0 / 10

The temporal lockbox: a hardened observatory for AI misalignment

kmenou

Abstract

This is a linkpost for https://kmenou.github.io/aips_website/temporal_lockbox_v0.1.html Summary: Weather forecasts by AI agents can be scored against measurements that do not yet exist and cannot plausibly be influenced by the forecaster. That causal gap enables harder-to-game AI evaluations under sustained optimization pressure - and an observatory for how agents act in various misaligned ways. Discuss

Research area

agentic evaluationevaluationsmulti-agent safety
Published
31 Jul 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.