← 피드로
[Submitted on 21 Jul 2026]
Abstract:Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their ability to recode visual representations when presented with goal-directed language remains poorly characterized. Indeed, prior work largely treats visual representations in VLMs as static repositories of visual information that are manipulated by language representations. In the present work, we provide evidence for two concrete instances of language-induced recoding of visual representations. First, we identify an abstract reference representation that denotes which objects are goal-relevant under a natural language prompt. We extract contrastive steering vectors corresponding to this reference representation and demonstrate that they are causally implicated in model predictions. These reference representations are abstract in that they generalize to different objects, different task contexts, and even from synthetic to naturalistic images. Second, we demonstrate language-induced attribute modulation: later layers selectively amplify goal-relevant attributes in visual representations of objects. We demonstrate this phenomenon across a range of different prompts. Finally, we provide a causal intervention that demonstrates that attribute modulation mediates a VLM’s response distribution. Together, our results support a more dynamic account of cross-modality processing in VLMs — rather than vision tokens serving as static repositories of information, they are modulated to support queries articulated in language.
Submission history
From: Michael Lepori Jr. [view email]
[v1]
Tue, 21 Jul 2026 06:06:20 UTC (8,721 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.00035
답글 남기기