A diagnostic framework for multimodal visual reasoning

See2Think
Do Multimodal Models Really Use Intermediate Visual States?

See2Think disentangles whether an intermediate visual state is useful from whether a model’s reasoning behavior actually depends on it.

Siyu Yan1,3,†, Zhuoran Yan2,†, Haiying Xu3,4,†, Panhao Zhou2, Jingyu Chen2, Chenhao Ji3, Shuo Cao3,5, Yongheng Zhang2, Haoze Liu3, Siyu Zhang3,6, Xiwen Gu7, Yihao Liu3, Alex Jinpeng Wang2,§

1HKUST 2Central South University 3Shanghai AI Laboratory 4HKUST (Guangzhou) 5USTC 6Fudan University 7Wuhan University

Equal contribution  ·  § Corresponding author

Beyond final answers

Seeing a visual trace is not the same as knowing it was used.

Multimodal models increasingly draw auxiliary lines, annotate regions, invoke visual tools, and construct intermediate images. Yet final-answer accuracy cannot tell us whether the model chose a relevant visual action, whether that action was rendered faithfully, or whether the returned state influenced later reasoning.

See2Think provides a unified framework for answering all three questions. It combines a visually grounded benchmark with an observable and intervenable inference protocol, enabling matched outcome-level and process-level evaluation.

See2Think framework comparing traditional evaluation with process-level diagnosis

Figure 1. See2Think moves beyond final-answer evaluation to diagnose action relevance, render faithfulness, and feedback uptake.

See2ThinkBench

1.2K visually dependent problems across 12 diverse tasks.

Every category contains 100 samples. Caption-only solvability filtering reduces text-dominant shortcuts, while a unified answer interface supports consistent evaluation.

Representative examples from the twelve See2ThinkBench task categories

See2ThinkBench at a glance. Twelve task categories organized into three visual reasoning environments.

6 tasks

2D Structured

Geometry, spatial puzzles, physics, chemistry, science QA, and abstract patterns.

2 tasks

3D Scene

Object attributes and compositional reasoning over synthetic 3D environments.

4 tasks

Real-world

Robot manipulation, state change, commonsense scenes, and physical plausibility.

Visual Action-of-Thought

An observable, intervenable loop for visual reasoning.

VAoT records alternating textual thoughts, structured visual actions, rendered states, and subsequent reasoning. Because every visual action is externally executed and inspected, the trajectory supports controlled intervention and process-level diagnosis.

The four controlled See2Think inference settings

Four matched settings. CoT, VAoT-NoRender, VAoT, and VAoT-WrongRender isolate planning, execution, and dependence on visual feedback.

Plan

Action Relevance

Does the selected operation target evidence that matters for the task?

Render

Render Faithfulness

Does the returned image correctly realize the requested visual action?

Use

Feedback Uptake

Does subsequent reasoning incorporate information from the visual state?

What we found

Visual reasoning is conditional—not a universal accuracy booster.

1.2K samples per model and setting
4 representative multimodal models
>10 pp drop under corrupted feedback in 3D scenes
See2Think final-answer accuracy across models, tasks, and inference settings

No single setting dominates. The best visual-reasoning strategy changes across model families and visual environments.

Finding 1

Task regime matters.

Direct perception leads on explicit 2D diagrams, action planning slightly leads in 3D scenes, and real-world tasks show no clear aggregate winner.

Finding 2

Rendering is the bottleneck.

Models usually select a relevant visual action, but faithfully executing that action remains substantially harder.

Finding 3

Usefulness and dependence differ.

Correct rendering may add little net accuracy, while corrupted visual feedback can still cause measurable performance degradation.

Process-level See2Think diagnosis across visual environments

Process-level diagnosis. High action relevance contrasts with lower render faithfulness; high feedback uptake does not necessarily produce higher accuracy.

BibTeX

@article{yan2026see2think,
  title   = {See2Think: Do Multimodal Models Really Use
             Intermediate Visual States?},
  author  = {Yan, Siyu and Yan, Zhuoran and Xu, Haiying and
             Zhou, Panhao and Chen, Jingyu and Ji, Chenhao and
             Cao, Shuo and Zhang, Yongheng and Liu, Haoze and
             Zhang, Siyu and Gu, Xiwen and Liu, Yihao and
             Wang, Alex Jinpeng},
  year    = {2026}
}