ORACLE GAP Code, result files and interactive figures for an anonymous ACL submission anonymous · no tracking

Anonymous ACL submission · supplementary site

A correct diagnosis is not a fix.

The paper's abstract

The question

Does the fix a diagnosis suggests beat plain fine-tuning?

Papers that look inside vision–language models often find a defect and then propose a mechanism to fix it. The mechanism is usually compared with the unchanged model. But a mechanism adds trainable parameters and is trained on task data, so the fair baseline is plain low-rank fine-tuning (LoRA) of the same model, on the same data, with a comparable number of trainable parameters. We call it the matched control, and we ran the whole cycle against it.

Figure 2 of the paper where each mechanism attaches, beside the matched control
Qwen2-VL-2B with the attachment point of each audited mechanism family, and the matched control

Q2 · Do the mechanisms help?

Nineteen mechanisms, each against its matched control

Each row is one mechanism minus its matched control, in accuracy points. The dot is the difference; the thick bar is one standard error and the thin bar two. A row in red is worse than its control by more than two standard errors. Click a row to see the model, benchmark, accuracies and the result file it comes from.

The audit
Tip
Click any row.

A tie, question by question

Same 1,000 questions, three seeds: the mechanism and its control trade answers

The two central mechanisms were re-run three times each, with one record per test question. That lets us look past the accuracy and ask on which questions the mechanism and its control differ. A mechanism that really helped would tend to win the same questions on every seed. Instead the wins go both ways, and retraining the control with a new random seed changes about as many answers as the mechanism does, or more.

Question by question GQA two-hop test-dev, n = 1,000 · Appendix F re-runs

Real questions where RIFR and LoRA-26 disagree

Answers are scored as in the paper: the first word of the answer must match GQA's answer (or one must start with the other). So a synonym such as “lady” for “woman” counts as wrong for whichever model says it. Photos: GQA test-dev (Visual Genome images, CC BY 4.0).

The oracle gap

Each lever works only when the answer chooses where to pull it

An oracle is a fix that is allowed to use the correct answer. For each lever we measured what it gains at every level of knowledge: with hindsight, when the correct answer picks the inputs, applied to every input, and learned as a trainable mechanism. Pick a lever.

The oracle ladder Table of the paper

Differences are in accuracy points against the unchanged model, except learned rows, which are against their fine-tuned controls. "≈" marks the hindsight value, which is extrapolated from 44 recoveries among 150 swept failures.

Q1 · Are the defects real?

Three defects that survive their controls, and one caution

D1 the two-hop answer forms late and is then lost · Qwen2-VL-2B, GQA two-hop

We read the model's best guess for the answer at every one of its 28 layers (the logit lens). Correct answers first become the top guess at layer on average. In the failures, the bars show where the correct answer got its best rank: most peaks sit in the last six layers, which leaves almost no depth to keep the answer.

What happens in the failures

What an oracle recovers

Real examples watch the answer form and get lost · GQA two-hop questions, logit lens at every layer

Pick a question. The line shows where the correct answer stands in the model's ranking of all words at each layer (guess #1 is the model's top guess). In a “found, then lost” question the correct answer climbs into the top five late in the network and then drops out at the last layer. Press play to watch it layer by layer.

D3 image patches in the same row share a position coordinate · try Qwen2-VL's own index rule

Qwen2-VL gives every image token three position ids: time, height and width. This explorer applies the model's rule for one image: image tokens start after the text, every token in a row gets the same height id, every token in a column the same width id, and the text after the image continues from the largest id plus one. Click an image token to see which tokens share its ids.

The image, one cell per merged token (h, w)
The sequence the decoder sees

Where each coordinate lives in one attention head

The repair, and what it changes

D2 and D4 the count is present in the image tokens · SmolVLM-256M, frozen

A linear probe reads the object count from the image tokens at seven points along the pipeline. On rendered scenes it reads the count far better than the model answers. On real images (CountBenchQA) the gap almost disappears, so we treat this present-but-unused information as a caution, not a defect.

Rendered scenes

Real images

probe accuracythe model's own accuracychance

Q3 · If not, why not?

Three measured causes

1. The fix is chosen by the answer. The same patch that recovers failures when the correct answer picks the inputs costs accuracy when applied to every input (see the ladder above).

2. The signals a gate could use do not say where to intervene. We scored label-free signals on the least a gate needs: predicting a wrong answer. Output confidence does this well; the signals that follow the answer through the middle layers do not; and none of them says which layer to patch.

Label-free signals AUROC for predicting a wrong answer · Qwen2-VL-2B, n = per benchmark

3. The model works around the defects. The position-code repair trains (its offsets reach a norm of 0.15–0.22) and still changes nothing. We read this as the decoder routing spatial information by content, so a cleaner position code goes unused.

Results

Every table of the paper

Parsed from the paper's LaTeX at build time, so the numbers are the paper's numbers. Each table comes with the paper's own note on how to read it.

Visualize

Every figure of the paper

Click a figure to enlarge it and read its caption and the paper's note on how to read it.

Code

All the code, as used

One script per experiment, with its arms as flags, and the scripts that render every table and figure of the paper. Files are byte-for-byte copies; their paths are the paths in the research repository.

Result files

Every number traces to one of these files

The list was captured by running each of the paper's table and figure scripts and recording every file it opened. Under each file: the scripts that read it.

Reproduce

Rebuild the paper's tables and figures in a minute

Download the archive, unzip it, and run from its top folder. The regenerated tables match the shipped ones byte for byte; we checked this on a clean copy before publishing.

pip install numpy matplotlib pillow
python3 papers/arr2026_oracle_gap/make_tables.py         # every generated table
python3 papers/arr2026_oracle_gap/make_rerun_tables.py   # Appendix F, from per-example records
python3 papers/arr2026_oracle_gap/make_audit_figure.py   # Figure 3 (other make_*.py: other figures)

Re-running an experiment needs one GPU with 20–48 GB and the public models and benchmarks, which the scripts download. For example, the three-seed re-runs of Appendix F:

python3 scripts/rifr.py --variant rifr --l1 16 --l2 26 --n-train 4000 --n-test 1000 --epochs 2 --bsz 6 --max-pixels 50176 --seed 0
python3 scripts/ang.py --variant ang --seed 0
python3 scripts/logitlens_probe.py --model Qwen/Qwen2-VL-2B-Instruct --n-eval 400

Every other re-run command is in runpod/arr_jobs/*.jobs. A mechanism and its control always share the image resolution of their experiment (224×224 for RIFR, deep supervision and the position-code repair; 448×448 for the neuron gates).

Checks

How this page was checked

Nothing on this page is typed by hand. A single build script reads the paper's LaTeX, its figures and the result files, and it refuses to publish if any of these checks fails:

    Limits

    What the evidence does not show

    • Most cells have one seed. Deep supervision gave +1.8 points on its first seed and +0.6 over three. The two central D1 mechanisms were re-run with three seeds and still tie.
    • Small differences are inside noise. Test sets have 100–1,000 items, so differences below 2–5 points cannot be told apart. The claim is not that the mechanisms hurt, but that none shows a gain distinguishable from its control.
    • Parameter counts are comparable, not always identical. A control has the same data, schedule and image resolution, and equal (within 5%) or fewer trainable parameters; RIFR is compared with both a smaller and a larger control and loses to both.
    • Coverage. Each axis rests mainly on one model family (Qwen2-VL-2B for reasoning, SmolVLM for perception), all training is LoRA on a frozen base, and D1 is weaker at 7B.
    • One signal is left out. The error-detection signal "late change" included the output layer in its window, which makes it nearly constant; it is not used in the paper.
    • Synthetic is not real. The present-but-unused gap (D4) is large on rendered scenes and at most +0.05 on real images.