Source-Modality Monitoring in Vision-Language Models
Etha Tianze Hua
Brown University
Tian Yun
Brown University
Ellie Pavlick
Brown University
Conference on Language Modeling (COLM), 2026
[Paper]   [Code]


Abstract

We define and investigate source-modality monitoring — the ability of multimodal models to track and communicate the input source from which pieces of information originate. We consider source-modality monitoring as an instance of the more general binding problem, and evaluate the extent to which models exploit syntactic versus semantic signals in order to bind words like image in a user-provided prompt to specific components of their input and context (i.e., actual images). Across experiments spanning 11 vision-language models (VLMs) performing target-modality information retrieval tasks, we find that both syntactic and semantic signals play an important role, but that the latter tend to outweigh the former in cases when modalities are highly distinct distributionally. We discuss the implications of these findings for model robustness, and in the context of increasingly multimodal agentic systems.


The Target-Modality Retrieval Task

We show a model an inconsistent image–caption pair — the caption describes a different scene than the image — and ask it to report the content of one specified source. Because the two inputs carry mutually exclusive information, answering correctly requires retrieving from the queried modality rather than simply preferring one modality overall.
An image of a three-wheeled motorcycle paired with the mismatched caption 'A picture of a person riding a horse'. Asking what the image shows should return the motorcycle; asking what the caption says should return the horse.
An example task instance. The same input pair yields two different correct answers depending on which source the question requests.


Most VLMs Track Source Modality Well

We measure selectivity: how often a model reports the queried modality rather than the other one, on a scale from −1 to 1. Across 11 VLMs and two captioning datasets, almost every model lands well above chance. Models whose input pipeline inserts explicit image wrapper tokens (★) tend to do better than those without.
Selectivity scores for 11 vision-language models, nearly all above zero, with InstructBLIP-7B negative. A heatmap shows valid-response rates near 1.0.
Left: aggregated source-modality selectivity across VLMs and datasets; error bars are the standard deviation across datasets. ★ marks VLMs with special image wrapper tokens. Right: the rate of valid responses under inconsistent pairs and under single-modality inputs.
VLMs can do this task. The interesting question is how — are they simply keying off the marker tokens?


Symbolic Binding Is Not the Whole Story

First, VLMs are genuinely capable of symbolic binding. If we replace the modality words with content-free labels like Dax and Wug, so that nothing about the label hints at what it refers to, models still retrieve the right span.
Bar charts showing high selectivity for InternVL3-14B, Qwen2.5-VL-32B and Gemma-3-12B under four arbitrary label assignments.
Performance on the purely symbolic retrieval task across four arbitrary label assignments, where image and caption content are tied to meaningless labels rather than modality names.
However, symbols are not the only cue available. Image and caption tokens differ markedly in their distributions, which offers an alternative route to modality identity. Already at the embedding layer, before any contextualization, tokens are far more similar to others of their own modality than across modalities, and a linear probe can separate the two perfectly.
Table reporting within-image cosine similarity 0.325, within-caption 0.101, average within-modality 0.213, cross-modality 0.020, linear probe accuracy 0.999 and a shuffled-label control accuracy of 0.491.
Distributional separation between visual and textual content-token embeddings, across 16 model–dataset pairs. The shuffled-label control sits at chance, confirming the probe is reading modality rather than probe capacity.
To separate the symbolic and distributional cues, we systematically perturb the symbolic marker information. If binding were purely symbolic, removing the markers should drop performance to chance and swapping them should reverse the model's answers. Neither happens: selectivity falls but stays clearly positive, and swapped markers do not flip predictions.
Table of the four input conditions: Unperturbed keeps the original image and caption markers; Arbitrary replaces them with Dax and Wug; Remove deletes both markers; Swap exchanges the image and caption markers.
The four marker conditions. Unperturbed keeps the original modality markers; Arbitrary replaces them with modality-irrelevant labels; Remove removes both markers; Swap exchanges the image and caption markers.
Bar charts comparing purely symbolic and purely distributional expectations against three models under unperturbed, marker-removal and marker-swap conditions. Observed behavior sits between the two extremes.
Selectivity when symbolic markers are removed or swapped, shown against the predictions of a purely symbolic and a purely distributional account. Real models sit between the two. The right-hand panels split by target modality, where image retrieval stays robust but caption retrieval collapses.
Source-modality monitoring is not reducible to symbolic markers. Models also read modality off the distributional character of the tokens themselves — image patches simply do not look like text.


The Word You Use to Refer Matters

Querying a piece of text with the label "document" makes the model rely heavily on marker information: it searches for content explicitly marked as a "document." In contrast, when queried for "text," the markers become largely irrelevant — the label is already semantically aligned with the textual content itself.
Selectivity for Qwen2.5-VL-32B across image-caption, image-text and image-document settings under unperturbed, marker-removal and marker-swap conditions.
Qwen2.5-VL-32B. Marker perturbations barely dent the image–text setting, but substantially reduce image–caption and image–document.
Semantic fit can substitute for explicit markers. Models readily recognize textual content as “text” even when marker cues are perturbed, but rely much more heavily on those cues to identify the same content as a “document.”


Markers Write Modality Into the Content Tokens

Where does the marker's contribution actually live? We collect hidden states at the content-token positions from a clean run with markers intact, then patch them into a run where the markers have been deleted. Selectivity partially recovers — so during contextualization the markers write modality identity into the surrounding content representations rather than keeping it local to themselves.
Diagram of the freeze-remove condition: hidden activations are collected at content-token positions from a normal run with markers intact, then patched into a second run in which the markers have been removed.
The freeze-remove intervention. Activations are collected at the content-token positions from a normal run with intact markers (bottom), then patched into a run where the marker tokens have been removed (top).
Bar charts showing that the freeze-remove condition recovers much of the selectivity lost by removing markers, especially for caption-target retrieval.
Selectivity under the freeze-remove intervention, compared with leaving the markers in place (unpert) and deleting them outright (RM).


Source Attribution Can Be Steered

Finally, we ask whether these representations can be manipulated. Keeping the model frozen, we learn two vectors added to the image span and the caption span, optimized to make the model report information from the wrong source. In early and middle layers this drives selectivity close to −1: the model reliably reports information from the modality it was not asked about. Marker positions are the more effective handle, and the effect fades in later layers.
Diagram of the learned-vector intervention: two trainable vectors are added either at the symbolic marker token positions or across the modality-specific content token positions of the image and caption spans.
The learned-vector intervention. Two trainable vectors, δ1 and δ2, are added either at the symbolic marker positions (marker deltas) or across the modality-specific content tokens (content deltas), applied respectively to the image span and the caption span.
Selectivity of Qwen2.5-VL-32B after learned interventions at different relative layer depths, dropping near -1 in early and middle layers.
Qwen2.5-VL-32B. Dashed and dotted lines mark the uninterventioned model and the token-level marker-swap condition.
The marker and content positions carry source-modality information that is causally accessible and adversarially exploitable.


Why It Matters

As models become more multimodal and agentic, not every source a user refers to will come with an explicit tag. People ask about “yesterday's meeting” or “the discussion about dinner plans” — micro-modalities that are distinct enough for humans to name, but not necessarily distinct enough for models to identify from distributional cues alone. When those semantic cues are weak, our results show that source binding becomes increasingly dependent on explicit markers, creating a potential point of fragility for reliable source attribution.


Paper and Bibtex

First page of the paper Source-Modality Monitoring in Vision-Language Models
Etha Tianze Hua, Tian Yun, Ellie Pavlick
Conference on Language Modeling (COLM), 2026

[arXiv]   [Code & Data]
Bibtex
@inproceedings{hua2026source,
  title     = {Source-Modality Monitoring in Vision-Language Models},
  author    = {Etha Tianze Hua and Tian Yun and Ellie Pavlick},
  booktitle = {Third Conference on Language Modeling},
  year      = {2026}
}


Acknowledgements

We are very grateful to the members of the Language Understanding and Representation (LUNAR) Lab at Brown University — especially Zhuonan Yang and Jennifer Meng Lu — as well as other members of the Brown Superlab for their valuable feedback on this paper. Etha Tianze Hua would like to thank Selina Liang for her support during this work. This project was in part supported by Schmidt Sciences Grant #GR5300958 and the NSF AI Research Institute on Interaction for AI Assistants Grant #GR5300593. Ellie Pavlick is a paid consultant for Google DeepMind. The content of this article does not necessarily reflect the views of the US Government or of Google, and no official endorsement of this work should be inferred.

This template was originally made by Phillip Isola and Richard Zhang for a colorful ECCV project; the code is available here.