Back to papers
arxiv6.0 / 10

Test-Time Training for Modality Order Consistency in Vision-Language Models

Aditi Gupta, Yossi Gandelsman

Abstract

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations diverge sharply between prompt orders. We find that the test-time training method repairs this misalignment across layers. Together, our results identify modality-order sensitivity as a circuit-level failure in VLMs and demonstrate that simple, asymmetric test-time adaptation can effectively mitigate it and even improve performance over the baseline.

Research area

activation patchingcircuit analysismodel robustness
Published
22 Jul 2026
Source
arxiv
Org
University of Chicago
View paper
Sign in to read and join the discussion.