Back to papers
arxiv9.0 / 10

Harmonizing AI Safety Thresholds

Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey

Abstract

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

Research area

ai governanceevaluationssystem-level risk assessment
Published
17 Jul 2026
Source
arxiv
Org
Brown University
View paper
Sign in to read and join the discussion.