Attention Maps Are Useless Predictors of VLM Correctness

A new mechanistic study on arXiv dismantles a core assumption in vision-language model evaluation: that sharp attention maps signal trustworthy answers. Testing LLaVA-1.5, PaliGemma, and Qwen2-VL across 3,090 samples, the researchers found attention structure predicts correctness at R=0.001, essentially random. Attention still matters for feature extraction, but it tells you nothing about whether the model is right.
Reliability lives later in the computation. A single hidden-state linear probe reaches AUROC above 0.95 on two of three families. Self-consistency at K=10 is the strongest behavioral predictor measured, at ten times inference cost. The study also surfaces a sharp architectural difference: LLaVA concentrates reliability in a fragile late bottleneck, losing 8.3 percentage points after ablating just five neurons. PaliGemma and Qwen2-VL distribute it widely, absorbing loss of roughly 50 percent of peak-layer hidden dimensions with under one point of degradation.
For anyone building monitors or deployment filters on top of VLMs, the source of ground truth just moved. Watch attention less. Probe hidden states more. Architecture choice now has direct implications for how brittle your reliability layer will be.