Source-Modality Monitoring in Vision-Language Models
Etha Tianze Hua, Tian Yun, Ellie Pavlick
Department of Computer Science, Brown University
Third Annual Conference on Language Modeling
San Francisco · October 6, 2026
What is the user referring to?
USER
Image
Caption
“A dog fetching a stick.”
USER
What is in the image?
VLM
A cat laying on a shelf.
USER
What is in the caption?
VLM
A dog fetching a stick.
Most VLMs do well on this task, but we want to study how they do it.
2
A general problem
Source monitoring across domains
Language MODEL
Alice lives in Paris.
Bob lives in Tokyo.
Where does Bob live?
→ Tokyo
Feng & Steinhardt, 2024; Gur-Arieh et al., 2025
Agents
[search]
Flights from $420
[calendar]
Trip:
Nov 3–7
What did search find?
→ Flights from $420
Tool outputs, files, past turns
HUMAN witness
Saw:
him in NY
Hearsay:
“he was at COLM”
Where did you see him?
→ In New York
Loftus & Palmer, 1974; Johnson et al., 1993
3
Inside the model
A reference word must point to a span of tokens
<image>
IM₁
···
IMₙ
</image>
Caption:
A
dog
fetching
a
stick
.
“A dog fetching a stick.”
vision encoder
tokenizer
4
Inside the model
A reference word must point to a span of tokens
···
~ cat, shelf
~ dog, stick
What
is
in
the
image
?
VLM: A cat laying on a shelf.
query
What
is
in
the
caption
?
VLM: A dog fetching a stick.
How does the word find the right span?
5
a special case of the binding problem
A binding problem with two possible cues
<vision_start>
IM₁
IM₂
…
IMₙ
<vision_end>
Caption
:
a
dog
fetching
a
stick
Symbolic cue — the markers
Outlined tokens index the span, like entity–attribute or variable binding in LMs (Feng & Steinhardt, 2024; Wu et al., 2025).
Distributional cue — the content
Filled tokens carry their own signal: some inputs just are images, whether or not they are tagged.
Evidence: made-up labels (Dax, Wug) still bind. Gemma 1.00 · Qwen 0.98 · InternVL 0.68.
Evidence: a linear probe separates image and caption embeddings: 99.9% (control 49.1%).
Either cue would suffice. Which one do VLMs actually use?
6
a special case of the binding problem
A binding problem with two possible cues
<vision_start>
IM₁
IM₂
…
IMₙ
<vision_end>
Caption
:
a
dog
fetching
a
stick
Symbolic cue — the markers
Outlined tokens mark the span, like entity–attribute or variable binding in LMs (Feng & Steinhardt, 2024; Wu et al., 2025).
Evidence: VLMs can retrieve information based on arbitrary labels (see §4.1 in the paper).
Distributional cue — the content
Filled tokens carry their own signal: some inputs just are images, whether or not they are tagged.
Evidence: a linear probe separates image and caption embeddings: 99.9% (control 49.1%).
Either cue would suffice. Which one do VLMs actually use?
7
a special case of the binding problem
A binding problem with two possible cues
<vision_start>
IM₁
IM₂
…
IMₙ
<vision_end>
Caption
:
a
dog
fetching
a
stick
Symbolic cue — the markers
Outlined tokens mark the span, like entity–attribute or variable binding in LMs (Feng & Steinhardt, 2024; Wu et al., 2025).
Distributional cue — the content
Content tokens from different modalities carry their own signal: they have very different distributions.
Evidence: VLMs can retrieve information based on arbitrary labels (see §4.1 in the paper).
Evidence: a linear probe separates image and caption embeddings: 99.9% (control 49.1%) (see §4.2).
In principle, either cue would be enough. But which one do VLMs actually use?
8
Marker perturbation
Which cue do VLMs rely on? Two initial hypotheses
Hypothesis 1
Purely symbolic
Only the markers tell the model which span is the image and which is the caption.
<image>
IM₁
···
IMₙ
</image>
Caption:
a
dog
Uses: marker tokens only
Hypothesis 2
Purely distributional
Only the content tells the model which span is which; markers are ignored.
<image>
IM₁
···
IMₙ
</image>
Caption:
a
dog
Uses: content tokens only
9
Task Setup and evaluation
Task design and the selectivity metric
1
Sample pairs
“A dog fetching a stick.”
Flickr30k & MSCOCO, images paired with a mismatched caption.
2
Control order
Image
Caption
Caption
Image
We test on both orders, averaged, so VLMs can't rely on positions to solve the task.
3
Ask the VLM
What is in the image?
“A cat laying on a shelf.”
Same for “caption”. Free-form answers from 11 VLMs.
4
LLM judge
answer came from…
✓ image
caption
neither
We use GPT-5.4-mini as our LLM judge
5
Selectivity
S = P(target)
− P(other)
−1
0
+1
+1: always the requested source;
0: no preference;
−1: always the other source.
10
Marker perturbation (cont'd)
Two conditions to test the two initial hypotheses
Condition
Input format, token by token
Normal
<image>
IM₁
···
IMₙ
</image>
Caption:
a
dog
fetching
a
stick
.
Remove
IM₁
···
IMₙ
a
dog
fetching
a
stick
Swap
Caption:
IM₁
···
IMₙ
.
<image>
a
dog
fetching
a
stick
</image>
Outlined = marker tokens, filled = content tokens, dashed = removed marker.
If purely symbolic
Markers decide.
assumption
predicted
1
0
−1
Selectivity
Normal
≈ 0
Remove
If purely distributional
Content decides.
assumption
predicted
1
0
−1
Selectivity
Normal
≈ 1.0
Remove
11
Marker perturbation (cont'd)
Two conditions to test the two initial hypotheses
Condition
Input format, token by token
Normal
<image>
IM₁
···
IMₙ
</image>
Caption:
a
dog
fetching
a
stick
.
Remove
IM₁
···
IMₙ
a
dog
fetching
a
stick
Swap
Caption:
IM₁
···
IMₙ
.
<image>
a
dog
fetching
a
stick
</image>
Outlined = marker tokens, filled = content tokens, dashed = removed marker.
If purely symbolic
Markers decide.
assumption
predicted
1
0
−1
Selectivity
Normal
≈ 0
Remove
≈ −1*
Swap
If purely distributional
Content decides.
assumption
predicted
1
0
−1
Selectivity
Normal
≈ 1.0
Remove
≈ 1.0
Swap
12
Marker perturbation · Result
Neither initial hypotheses hold
Averaged over image and caption targets
1
0.5
0
−0.5
−1
Selectivity
0.86
I
0.99
G
0.98
Q
Normal
0.20
I
0.50
G
0.46
Q
Remove
distrib.
symbolic
0.30
I
0.65
G
0.52
Q
Swap
distrib.
symbolic
hypothesis predictions
I = InternVL3-14B
G = Gemma-3-12B
Q = Qwen2.5-VL-32B
13
Marker perturbation · Result by target modality
Images don’t need markers as much as captions do
1
0.5
0
−0.5
Selectivity
Asked about the image
I
G
Q
Normal
I
G
Q
Remove
I
G
Q
Swap
Asked about the caption
I
G
Q
Normal
I
G
Q
Remove
I
G
Q
Swap
I = InternVL3-14B
G = Gemma-3-12B
Q = Qwen2.5-VL-32B
Image target: robust
Holds still under removing or swapping markers.
Caption target: collapses
Falls to chance level.
14
Follow-up · Referential words for the textual span · Setup
Does the word you use to refer matter?
But “caption” is just one specific way we refer to a span of text. What if we use other words?
Input
<image>
IM₁
···
IMₙ
</image>
Caption:
a
dog
fetching
a
stick
.
→
“What is in the caption?”
Text:
a
dog
fetching
a
stick
.
→
“What is in the text?”
Document:
a
dog
fetching
a
stick
.
→
“What is in the document?”
15
Follow-up · Referential words for the textual span
Yes, the word you use to refer matters.
Qwen2.5-VL-32B · asked about the text source
1
0.5
0
−0.5
−1
Selectivity
“caption”
0.98
−0.02
0.07
“caption” | collapses to chance
Results re-plotted from marker perturbation.
“text”
0.90
0.72
0.81
“text” | stays highly selective
The referential relation survives the explicit markers.
“document”
0.88
−0.72
−0.68
“document” | flips to negative
When you asked for the document, it describes the image content.
Normal
Remove
Swap
Some words (i.e. "text") are more inherently related to spans of inputs than other words ("caption", "document").
16
The big question
As models take in information from more and more sources, can they keep track of where each piece came from, and use the one we asked for?
This work
→ VLMs use both distributional AND symbolic cues
→ Different sources have a different level of reliance on the presence of explicit marker tags
Future work
→ Models that take in more modalities (Omni-models)
→ Agentic systems with tool outputs, files, past turns under real-world use cases
17
Thank you & questions?
in the paper: check out the experiments not covered today
1
Behavioral evaluation across 11 VLMs
§3
Almost all VLMs can track source modality well.
2
Marker perturbation
§4.3
Covered today
Images need content alone; captions need markers.
3
Referential words
§4.3
Covered today
“text” refers without markers; “caption” and “document” don’t.
4
Interaction between marker and content tokens
§4.4
Marker information is written into the content tokens.
5
Steering VLMs for modality misattribution
§5
Marker positions are a stronger causal handle than content.
Project webpage
paper · code · data
Come see our poster:)
Main conference
Imperial Ballroom · Poster #2
Tue, Oct 6 · 11am–1pm
ActInterp workshop
Continental Ballroom 7–8
Fri, Oct 9 · 2:25–3:30pm
18