Back to papers
lesswrong8.0 / 10

Stickiness in AI Behavioral Design

James_T

Abstract

Current model specs aim to shape the behaviors of near-present models, rather than the behaviors of models arbitrarily far into the future. OpenAI writes that their model spec aims to apply “0-3 months ahead of the present.” Anthropic’s Constitution for Claude notes that the document “is likely to change in important ways in the future.” So these documents are presented as provisional guidelines, not as trying to set behavioral standards for the far future. But what if current model behaviors transfer into future models by default? My thesis is that the behavioral targets that spec authors set for present LLMs will have a large influence on the behavior of future, more powerful LLMs. As a result, future AIs may be governed by rules poorly suited to their greater capabilities and more pervasive roles. The extremely capable, long-running, and ubiquitous LLMs of the future might end up acting according to behavioral targets written for less capable, shorter-running, and rarer LLMs of the past. This could be quite bad, especially if such defaults become so entrenched that they are not only hard to undo, but hard even to notice as contingent features of reality. First, I’ll make the des

Research area

agentic misalignmentagent misalignmentalignment
Published
13 May 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.