Anonymous ACL submission · supplementary site
A correct diagnosis is not a fix.
The paper's abstract
The question
Does the fix a diagnosis suggests beat plain fine-tuning?
Papers that look inside vision–language models often find a defect and then propose a mechanism to fix it. The mechanism is usually compared with the unchanged model. But a mechanism adds trainable parameters and is trained on task data, so the fair baseline is plain low-rank fine-tuning (LoRA) of the same model, on the same data, with a comparable number of trainable parameters. We call it the matched control, and we ran the whole cycle against it.
Q2 · Do the mechanisms help?
Nineteen mechanisms, each against its matched control
Each row is one mechanism minus its matched control, in accuracy points. The dot is the difference; the thick bar is one standard error and the thin bar two. A row in red is worse than its control by more than two standard errors. Click a row to see the model, benchmark, accuracies and the result file it comes from.
- Tip
- Click any row.
A tie, question by question
Same 1,000 questions, three seeds: the mechanism and its control trade answers
The two central mechanisms were re-run three times each, with one record per test question. That lets us look past the accuracy and ask on which questions the mechanism and its control differ. A mechanism that really helped would tend to win the same questions on every seed. Instead the wins go both ways, and retraining the control with a new random seed changes about as many answers as the mechanism does, or more.
Real questions where RIFR and LoRA-26 disagree
Answers are scored as in the paper: the first word of the answer must match GQA's answer (or one must start with the other). So a synonym such as “lady” for “woman” counts as wrong for whichever model says it. Photos: GQA test-dev (Visual Genome images, CC BY 4.0).
The oracle gap
Each lever works only when the answer chooses where to pull it
An oracle is a fix that is allowed to use the correct answer. For each lever we measured what it gains at every level of knowledge: with hindsight, when the correct answer picks the inputs, applied to every input, and learned as a trainable mechanism. Pick a lever.
Differences are in accuracy points against the unchanged model, except learned rows, which are against their fine-tuned controls. "≈" marks the hindsight value, which is extrapolated from 44 recoveries among 150 swept failures.
Q1 · Are the defects real?
Three defects that survive their controls, and one caution
We read the model's best guess for the answer at every one of its 28 layers (the logit lens). Correct answers first become the top guess at layer on average. In the failures, the bars show where the correct answer got its best rank: most peaks sit in the last six layers, which leaves almost no depth to keep the answer.
What happens in the failures
What an oracle recovers
Pick a question. The line shows where the correct answer stands in the model's ranking of all words at each layer (guess #1 is the model's top guess). In a “found, then lost” question the correct answer climbs into the top five late in the network and then drops out at the last layer. Press play to watch it layer by layer.
Qwen2-VL gives every image token three position ids: time, height and width. This explorer applies the model's rule for one image: image tokens start after the text, every token in a row gets the same height id, every token in a column the same width id, and the text after the image continues from the largest id plus one. Click an image token to see which tokens share its ids.
Where each coordinate lives in one attention head
The repair, and what it changes
A linear probe reads the object count from the image tokens at seven points along the pipeline. On rendered scenes it reads the count far better than the model answers. On real images (CountBenchQA) the gap almost disappears, so we treat this present-but-unused information as a caution, not a defect.
Rendered scenes
Real images
Q3 · If not, why not?
Three measured causes
1. The fix is chosen by the answer. The same patch that recovers failures when the correct answer picks the inputs costs accuracy when applied to every input (see the ladder above).
2. The signals a gate could use do not say where to intervene. We scored label-free signals on the least a gate needs: predicting a wrong answer. Output confidence does this well; the signals that follow the answer through the middle layers do not; and none of them says which layer to patch.
3. The model works around the defects. The position-code repair trains (its offsets reach a norm of 0.15–0.22) and still changes nothing. We read this as the decoder routing spatial information by content, so a cleaner position code goes unused.
Results
Every table of the paper
Parsed from the paper's LaTeX at build time, so the numbers are the paper's numbers. Each table comes with the paper's own note on how to read it.
Visualize
Every figure of the paper
Click a figure to enlarge it and read its caption and the paper's note on how to read it.
Code
All the code, as used
One script per experiment, with its arms as flags, and the scripts that render every table and figure of the paper. Files are byte-for-byte copies; their paths are the paths in the research repository.
Result files
Every number traces to one of these files
The list was captured by running each of the paper's table and figure scripts and recording every file it opened. Under each file: the scripts that read it.
Reproduce
Rebuild the paper's tables and figures in a minute
Download the archive, unzip it, and run from its top folder. The regenerated tables match the shipped ones byte for byte; we checked this on a clean copy before publishing.
pip install numpy matplotlib pillow python3 papers/arr2026_oracle_gap/make_tables.py # every generated table python3 papers/arr2026_oracle_gap/make_rerun_tables.py # Appendix F, from per-example records python3 papers/arr2026_oracle_gap/make_audit_figure.py # Figure 3 (other make_*.py: other figures)
Re-running an experiment needs one GPU with 20–48 GB and the public models and benchmarks, which the scripts download. For example, the three-seed re-runs of Appendix F:
python3 scripts/rifr.py --variant rifr --l1 16 --l2 26 --n-train 4000 --n-test 1000 --epochs 2 --bsz 6 --max-pixels 50176 --seed 0 python3 scripts/ang.py --variant ang --seed 0 python3 scripts/logitlens_probe.py --model Qwen/Qwen2-VL-2B-Instruct --n-eval 400
Every other re-run command is in runpod/arr_jobs/*.jobs. A mechanism and its control always share the image resolution of their experiment (224×224 for RIFR, deep supervision and the position-code repair; 448×448 for the neuron gates).
Checks
How this page was checked
Nothing on this page is typed by hand. A single build script reads the paper's LaTeX, its figures and the result files, and it refuses to publish if any of these checks fails:
Limits
What the evidence does not show
- Most cells have one seed. Deep supervision gave +1.8 points on its first seed and +0.6 over three. The two central D1 mechanisms were re-run with three seeds and still tie.
- Small differences are inside noise. Test sets have 100–1,000 items, so differences below 2–5 points cannot be told apart. The claim is not that the mechanisms hurt, but that none shows a gain distinguishable from its control.
- Parameter counts are comparable, not always identical. A control has the same data, schedule and image resolution, and equal (within 5%) or fewer trainable parameters; RIFR is compared with both a smaller and a larger control and loses to both.
- Coverage. Each axis rests mainly on one model family (Qwen2-VL-2B for reasoning, SmolVLM for perception), all training is LoRA on a frozen base, and D1 is weaker at 7B.
- One signal is left out. The error-detection signal "late change" included the output layer in its window, which makes it nearly constant; it is not used in the paper.
- Synthetic is not real. The present-but-unused gap (D4) is large on rendered scenes and at most +0.05 on real images.