LLM·

LoopCD — Contrastive Decoding Across Recurrent Loops — What does it mean?

The news. On October 1, 2026, a team working at Apple posted Decoding Looped Transformers Better for (Almost) Free (arXiv 2610.02185). Across four looped model families — Ouro, Huginn, Parcae and Looped-Qwen3 — LoopCD raised Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33% and Huginn's HumanEval pass@1 from 22.56% to 31.71%, with no extra training. Run at half the loops, the guided models matched or beat their unguided full-depth versions on a seven-benchmark multiple-choice suite while using 22.5% to 48.2% fewer forward FLOPs. Read the paper →

Picture an editor holding two versions of the same paragraph. The first draft says one thing, the final draft says something a little different, and the editor's best guess at what the writer is reaching for is not the final draft itself but the final draft pushed a bit further in the direction the revisions were already going. LoopCD makes that move on every token: it reads the direction in which the loops have been changing the prediction, and continues it.

That only works because of how a looped Transformer is built. An ordinary model sends each token through a stack of different blocks, so an early layer's output lives in a different representation space from the last layer's. A looped model instead applies the same block R times, so the hidden state after loop 1 and the hidden state after loop R are the same kind of object, readable by the same output head, and both can be turned into logits for the same next token. The paper measures the gap between them: read alone, the loop-1 prediction trails the final one by 4.0 to 23.7 points on a seven-benchmark average across six models. Every token therefore arrives with a weak draft and a strong draft, and ordinary decoding throws the weak one away.

The update is one line: guided = final + ω × (final − first). LoopCD-Logits applies it to the two logit vectors, which costs one extra pass through the post-loop layers and the output head to score the reference. LoopCD-Hidden applies it to the two hidden states before those layers, so the output head still runs once and the added cost rounds to zero. Setting ω = 0 gives ordinary decoding back. Settings vary by model and task: the paper uses fixed strengths of ω = 0.2–0.3 for answer and code generation, because a nudged early token compounds through everything written after it, 0.4–0.5 for mathematical reasoning, and 0.5 for multiple-choice scoring, which tolerates a stronger push.

An editor's changes only matter where the writer hesitated, and the same holds here. The paper splits the contrast into two parts: one that scales every logit by the same factor, which acts like temperature and never changes the ranking, and one that re-ranks tokens. The re-ranking part carries the gain, and it can only flip a decision when its push is larger than the gap between the top two candidates. On ARC-Challenge, guidance added +6.4 to +13.3 points on the least-confident fifth of questions and at most +0.4 on the most-confident fifth, and only 5.8% to 14.9% of answers flipped at all. That is why the adaptive rule sets ω = ω_max × (1 − (p1 − p2)), where p1 and p2 are the two largest probabilities: a full push on a coin flip, almost none on a settled choice, much like greedy decoding being left alone whenever it is already sure. The paper uses this adaptive rule only with LoopCD-Logits; LoopCD-Hidden keeps a fixed strength, because computing p1 and p2 would cost the extra output pass it exists to avoid.

Weak referenceWhere it comes fromExtra work per tokenRelationship to the strong prediction
Smaller auxiliary model (Li et al., 2023)A second, separately trained LMA full forward pass of the second modelA different model; the two predictions are compared over the shared vocabulary
Perturbed context (Shi et al., 2023)The same model run on an altered inputA second forward pass of the same modelThe same model, given a different input context
Intermediate layer (Chuang et al., 2024)An earlier layer of the same stackAn extra readout through the output headAn intermediate layer of one stack, not a complete pass of the same block
Earlier loop (LoopCD)Loop 1 of the same shared blockOne extra coda + head pass (Logits) or none (Hidden)A complete earlier pass of the same shared block, for the same token

Worked example

Hold three things fixed: two candidate tokens, A and B; strength ω = 0.5; and these logits (illustrative). After loop 1, A scores 2.0 and B scores 1.2. After the last loop, A scores 2.0 and B scores 1.9, so A still wins by a hair. The edits between the drafts are 0 for A and +0.7 for B. Guided scores: A = 2.0 + 0.5 × 0 = 2.0, and B = 1.9 + 0.5 × 0.7 = 2.25. B now wins: the loops had been pulling B up and leaving A flat, and LoopCD finishes a trend the remaining loops had not completed. Had the last loop put A at 2.5 instead, B's push of 0.35 would be smaller than the 0.6 gap and A would keep the token — the close-call rule above, in numbers.

The same lever can buy back compute. In the paper's FLOP table for one 512-token prefill, Huginn at 32 loops costs 54.5 TFLOPs; at 16 loops it costs 28.26, and LoopCD-Logits multiplies that by 1.042 for the extra readout, giving about 29.4 TFLOPs, roughly 46% less (our arithmetic from the paper's per-configuration numbers). At that half depth, Huginn with LoopCD-Logits scored 1.02 points above its unguided 32-loop self on the seven-benchmark multiple-choice mean. In that half-depth setting the paper takes its reference from loop 5 rather than loop 1, to suit the shorter trajectory. These are arithmetic counts, not measured latency.

Where it does not reach

LoopCD needs a looped model. A standard Transformer has no earlier draft in the same space, which is why the older methods in the table borrowed one from a smaller model, a changed prompt, or an intermediate layer. The tested models are small — the largest, Looped-Qwen3, is built from Qwen3-4B — and the half-depth savings were measured on the multiple-choice suite only. The knob is also touchy: for generation, quality falls steeply past roughly ω ≈ 0.6, pushing toward the weak draft (ω below zero) lowers the multiple-choice mean for every evaluated model, and on GSM8K the gains were mixed, with a drop of up to 1.29 points on Huginn. For Huginn, whose loop starts from random noise, the hidden-state variant also needs a later reference (loop 6 or 7) instead of loop 1. It sits next to early exit and self-speculative decoding as another way to get more out of states a model already computes.

Goes deeper in: LLM Internals → Text Generation → From Logits to Probabilities

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based