Back to papers
arxiv6.0 / 10

Single-Query Black-Box Calibration Auditing via Logit Bias

Roman Plaud, Antoine Saillenfest, Matthieu Labeau, Thomas Bonald, Willem Waegeman

Abstract

Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.

Research area

evaluationsystemic governance & auditabilityuncertainty quantification
Published
4 Sept 2026
Source
arxiv
Org
Institut Polytechnique de Paris
View paper
Sign in to read and join the discussion.