LLM·

HEAR protocol — Harness-engine coordination for workflow-aware KV caching — What does it mean?

The news. On October 5, 2026, a team from Harbin Institute of Technology (Shenzhen) posted Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving to arXiv. The paper defines HEAR (Harness–Engine pAiRing), a bidirectional protocol with four message categories, and tests it with two policies: cache-aware request ordering with KV retention, and per-role selection of inference modes. It reports a 1.61× batch speedup and 2.23× lower median time to first token on SCBench, and 1.23× and 2.45× end-to-end speedups on BrowseComp-Plus and DeepResearchBench, without observed loss in task quality. Read the paper →

Picture a busy restaurant where the waiter and the kitchen never talk. The kitchen sees only order slips: when the counter fills up, it clears whichever station has been idle longest. The waiter sees only the tables: table 3 is finishing its main course and will order dessert in a moment, but the kitchen does not know that, so it clears table 3's station and has to prep it again from scratch. The kitchen is missing the waiter's knowledge of who comes back next, and the waiter is missing the kitchen's knowledge of what is still prepped. Each side makes a choice that looks reasonable with the information it has, and the meal gets slower.

An agent run has the same split. The agent harness knows the workflow: which subagent is waiting on a tool call, which context will be resumed, which branch is finished. The inference engine knows the runtime: which contexts still have KV cache resident on the GPU, how long the queue is, how much memory is free. The paper names both failures. Without workflow intent, the engine may evict a context's KV state shortly before it is reused, which forces a costly re-prefill. Without residency feedback, the harness may send a cold request ahead of an equally ready one whose cache is still warm, so the scheduler wastes the warm cache too.

HEAR is the set of notes through the hatch. The harness sends two kinds of message (a description of the workflow, and requests for action), and the engine sends two kinds back (its current state and capabilities, and the outcome of each request). The paper is strict about what each note means, with four rules: saying a context will be reused is not the same as asking the engine to keep it; a preference may be ignored but a requirement must be met or explicitly rejected; a report that a prefix is on the GPU is an observation, not a reservation; and an engine that accepted a request to prepare a cache has not yet completed it. These rules matter because both layers change on their own, so a policy that treats a hint as a command or an old report as a promise will act on a state that no longer exists.

Walk the paper's SCBench run, holding three things fixed: the model (Qwen3-8B), the hardware (one RTX PRO 6000 GPU), and the workload (multi-turn conversations with random arrivals and think time between turns). Under plain FCFS, only 19.2% of follow-up prompt tokens reused cached KV and the rest were recomputed, and finishing every conversation took 481 seconds. When the harness uses the engine's residency reports to run warm requests first, reuse rises to 87.6%, median TTFT falls from 63.1 s to 4.6 s, and the batch finishes in 224 s (481 ÷ 224 ≈ 2.15×). The cost is the cold requests that keep getting skipped: the worst time to first token rises from 79.0 s to 93.9 s. Adding a 40-second guard, after which a waiting request can no longer be bypassed, gives a worst time to first token of 65.0 s and a batch time of 299 s (481 ÷ 299 ≈ 1.61×). The guard spends 75 seconds of batch time (299 − 224) to remove about 29 seconds from the worst time to first token (93.9 − 65.0), which is the trade a tail-latency target asks for. That is also why the guard is a requirement in HEAR's terms, not a preference: the scheduler may reorder for cache reuse, but never past a request that has already waited 40 seconds. The guard limits skipping; it does not cap the wait itself, since queueing and prefill still add time after the guard applies.

Policy (SCBench)Median TTFTMax TTFTBatch completionKV reuseSource
FCFS (default engine order)63.1 s79.0 s481 s19.2%Table 3
Cache-aware ordering4.6 s93.9 s224 s87.6%Table 3
Cache-aware + 40 s guard28.3 s65.0 s299 s66.1%Table 3
Cache-aware + 60 s guard4.5 s80.4 s230 s84.7%Table 3

The second test, on a production-derived Mooncake trace served by Qwen3-8B on one H100, shows that the harness's hint alone is not enough. Here the harness estimates how likely each session is to continue and sends that as a description, not a command; the engine combines it with recency to decide what to keep. At 50% load, where requests rarely queue and idle-time eviction is the main loss, this raised KV reuse from 19.5% to 31.9%. At full load, stacking session-aware retention, cache-aware ordering and the guard together gave almost no latency gain and 0.89× the FCFS throughput, while cache-aware ordering alone or session-aware retention alone each beat FCFS. No single policy won at every load, so the hint about future reuse is only useful when the engine can weigh it against its live queue and memory state. This is the paper's main argument for a two-way protocol instead of a one-way hint field.

The same messages also choose how each agent role runs, on a slower timescale. In a LangGraph harness with a Main agent and Reader subagents, the harness describes each role's workload and the engine reports which KV-compression modes it supports and how they perform. Running the Readers with H2O (a mode that keeps KV only for the most-attended tokens plus the most recent ones) cut their mean latency by 5.60× on BrowseComp-Plus and 7.22× on DeepResearchBench. The best mode for the Main agent depended on the workload, and the selected setups gave 1.23× and 2.45× end-to-end speedups with accuracy of 46.63% vs 46.15% and a RACE report-quality score of 41.1% vs 40.6% against the plain vLLM setup. Using one fixed mapping for both workloads would have cost 7.8% and 14.0% more time, which supports the paper's point that the choice should come from the workload description, not from a fixed setting.

Two limits are worth stating. HEAR is a contract, not an optimizer: the speedups come from the policies built on top of it, and the paper says its benefits depend on which information and controls a given serving stack exposes. The tests also ran a single engine on a single GPU per role, so behaviour in a large multi-node deployment is not measured. For production agent work, the durable idea is the boundary itself: the harness should tell the engine which contexts it will reuse and when (in prompt-cache terms, which prefix is worth keeping), and the engine should tell the harness what is actually cached, instead of each side guessing.

Goes deeper in: LLM Serving → Prefix Caching & RadixAttention → Eviction and Copy-on-Write

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based