HARISSA self-checks for local LLMs — Hidden-state correctness probes — What does it mean?
The news. On September 29, 2026, Kenan Alkiek, Moontae Lee, David Jurgens and V.G. Vinod Vydiswaran posted HARISSA (Hidden-Activation Reads for Inference-time Safety and Self-Assessment) to arXiv. They fine-tune Qwen 3.5 models so that two hidden states predict correctness, then test on MedQA (US medical-licensing questions), MedMCQA (medical entrance-exam questions) and BBH (general reasoning puzzles), all multiple choice. With one 4B model on a device, HARISSA lands within one accuracy point of chain-of-thought at 2.7 times lower latency, averaged over the three tasks. On a server holding the 2B, 4B, 9B and 27B sizes, it is more accurate than the FrugalGPT and Self-REF cascades at the same latency. Read the paper →
Picture a student in an exam. Before writing anything, they read the question and get a feeling: I know this one, or I don't. After writing a quick answer, they get a second feeling: that looks right, or something is off. A careful student uses both. A doubtful quick answer gets redone on scratch paper, and an answer that still feels wrong becomes a question for the proctor instead of a mark on the answer sheet. HARISSA gives a language model the same two gut checks, and reads them straight out of its activations.
The first gut check is the prefill state: the hidden vector at the last prompt token after the model's single forward pass over the prompt, the prefill phase, before any output token exists. The second is the answer state: the hidden vector at the last token of the generated answer. A linear probe (logistic regression on one layer, calibrated on held-out data so that a score of 0.8 means roughly 80% of such answers are right) turns each vector into a probability that the answer is correct. Each probe is one dot product, so the paper reports it adds no measurable latency, and no separate judge model is needed to score anything.
The probes work better when the model is trained to make them work. HARISSA runs every answering mode on every training query, grades each answer, and fine-tunes the model with LoRA so that linear heads on both states predict those grades. Fine-tuning sharpens the probes and barely changes the answers: on the 4B device model it lifts the answer probe's AUROC by 0.12 on MedQA, while each mode's accuracy moves by at most 1.7 points.
At inference the policy walks the modes from cheapest to most expensive. Before each mode except the last, it checks the prefill probability; below the skip threshold, that mode never generates. Otherwise the mode generates and the policy checks the answer probability; at or above the fall-through threshold it stops, and below it the answer is discarded and the next mode runs. If the policy reaches the last mode, that mode always generates, since nothing is left to try. If the answer it stops with is still under the deferral threshold (set equal to the fall-through threshold in the paper), the answer is withheld and the query goes to the person. That last move is a fail-safe default: when the system cannot vouch for an output, it stops instead of guessing.
The two gut checks are not equally sharp. The answer state predicts correctness better than the prefill state, because it has already seen the answer. On the device, the answer probe reaches 0.83 AUROC on MedQA against 0.63 for the prefill probe; on BBH the two are 0.92 and 0.81. A separate text classifier (ModernBERT) that reads only the question scores 0.53 on MedQA, below the prefill probe, so the prefill state carries something a difficulty classifier cannot: what this particular model knows.
Which gut check the policy relies on is set by prices, not by preference. On the device, thinking costs about eight times a direct answer, but the prefill probe is weak on the medical tasks, so the threshold sweep picks a skip threshold of 0 and lets the answer check do the work. On the server, skipping the 2B size costs only its prefill pass, so the policy skips it on most queries; without that prefill skip, it is 6.3 points less accurate on MedQA at similar latency. This is the same question as deciding when to spend more tokens, answered by a learned signal for one model on one task.
Here is where the speed-up comes from, using the paper's MedQA device numbers for Qwen 3.5 4B. A direct answer takes 9.8 s and is right 78.8% of the time; chain-of-thought takes 81.0 s and is right 86.2% of the time. HARISSA always writes the direct answer first, then thinks only when the answer probe doubts it, which happens on 29% of queries. The average query therefore costs about 9.8 s + 0.29 × 81.0 s ≈ 33 s, matching the reported 33.0 s at 85.7% accuracy: half a point below chain-of-thought for about 2.5 times less waiting on this task (the 2.7× headline is the average over three tasks). FrugalGPT has to think on 75% of queries to reach the same accuracy, because its external scorer misses more of the wrong direct answers. Deferral then adds the safety layer: handing 8.7% of MedQA queries to the person removes 32% of the wrong answers that would otherwise be delivered, nearly four times what deferring the same share at random removes.
| Policy (Qwen 3.5 4B, MedQA, device) | Accuracy | Mean latency | Source |
|---|---|---|---|
| Direct (thinking off) | 78.8% | 9.8 s | Table 1 |
| Chain-of-thought | 86.2% | 81.0 s | Table 1 |
| Self-consistency (3 samples) | 80.0% | 29.1 s | Table 1 |
| FrugalGPT cascade | 85.6% | 70.6 s | Table 1 |
| Self-REF cascade | 84.5% | 45.0 s | Table 1 |
| HARISSA | 85.7% | 33.0 s | Table 1 |
Latency in that table is computed from each query's token counts at rates measured on one NVIDIA A40, with every query answered and deferral off.
Two limits are worth knowing. Each probe is calibrated within one mode, so it can say whether an answer is likely right, but not which of two model sizes is better. A router that generates at every server size and keeps the answer with the highest probability gains only about one point over the 27B size alone, at roughly twice its latency, so HARISSA cascades instead of routing. The savings also depend on how expensive the slow mode is: on Gemma 4, thinking costs only 1.5 times a direct answer on MedQA, so there is little latency to save and the policy thinks on 69% of MedQA queries. Earlier work found that probes of this kind transfer poorly across datasets, which is why HARISSA fits them on each deployment's own training split.
Goes deeper in: Agent Engineering → Layered Guardrails → Fail-Safe vs Fail-Open
Related explainers
- Cluster, Route, Escalate: a cost-aware cascade: the cloud-escalation cascade HARISSA replaces when no bigger model is available
- When2Think: difficulty-aware think routing: a post-training answer to the same "think or not" decision