Agent·

Not all GUI-agent experience belongs in the weights — Component routing: context vs weights — What does it mean?

The news. On October 1, 2026, researchers at South Dakota State University posted Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents (arXiv 2610.01787). Earlier studies disagreed on whether an agent should learn from its own past runs by fine-tuning or by retrieving notes into the prompt. This paper argues the question was asked at the wrong size: a trajectory is a bundle of different kinds of knowledge, and each kind has its own best destination. It tests this on two Android benchmarks (MobileGym and AndroidWorld) with three 7–8B vision-language backbones and three seeds. Read the paper →

Picture a new hire in a busy office. After a month, some things live in their hands: where the stapler is, which drawer holds the forms. Nobody re-reads a note for that. Other things live on sticky notes on the monitor: "this client's file needs the manager's signature before it goes out." That rule only matters when that exact client's file is on the desk, so it is better read at the moment than baked into a reflex. The paper's claim is that an agent's experience splits the same way, and that you can tell which pile an item belongs to before you train anything.

A self-improving GUI agent runs a batch of tasks and keeps every trajectory. It can return that experience in two ways. The context route stores notes in a library and, at each step, appends (in the single-component tests) the five notes whose source screen best matches the current screen to the prompt (procedures are first narrowed to the most similar task by its text), the same retrieve-then-generate loop a RAG system uses. The weights route fine-tunes the model on the same items with a LoRA adapter. The paper first splits each trajectory into four components: locators (where a named element sits on a screen), procedures (a short run of actions that recurs inside one task family), state facts (a precondition that must hold before an action works), and lessons (an action that failed in a state, paired with the one that later worked there).

Then it scores every component on two properties measured from the pool alone. Recurrence is how often an item's key comes back in unrelated tasks on the same app; that is the stapler. State-conditionality is how much the item's usefulness depends on the current screen matching the screen it came from; that is the client-specific signature rule. A linear score, high recurrence pushing toward the weights and high state-conditionality pushing toward the context, picks the destination. In context-engineering terms, this is a rule for which knowledge you write into the model and which you select into the window at run time — and the window is a scarce resource, so moving frequently recurring items into the weights also frees window space.

ComponentWhat it recordsRoute gain (weights − context)Winning destination
LocatorsWhere a named element sits on a screen+4.9 points (paper)Weights
LessonsA failed action in a state, paired with the action that worked+1.5 points (paper)Weights (near the boundary)
State factsA precondition that must hold before an action succeeds−4.1 points (paper)Context (near the boundary)
ProceduresA short run of actions that recurs inside one task family−4.2 points (paper)Context

Route gain = held-out task success with the component in the weights minus success with it in context, on the same items and the same held-out instances, pooled over three backbones on MobileGym; AndroidWorld repeats the same sign pattern with wider intervals. Positive means the weights win. The lessons margin has an interval that includes zero in half of the cells.

The rule held up where it matters. Fitted with one backbone family held out, it recovered the winning destination in 24 of 24 held-out cells. Two interventions moved a component the way the rule predicts: copying every procedure item eightfold before fine-tuning (a stand-in for higher recurrence) moved the procedures' route gain from −4.2 to −0.8, and retrieving state-fact notes by task text alone, which strips their screen condition, cost 3.4 points. Routing each component to its rule-assigned destination beat the better single destination by +3.5 points on average across three backbones and two environments, and beat the reversed assignment by 7.8. The note's effect also changes after training: before fine-tuning, a retrieved note changed the agent's action in 42% of probed steps across all components; once the same items were in the weights, that fell to 16% for locators and 15% for lessons, but only to 22% for procedures. A note gets read when it carries something the weights do not already hold.

Worked example — what happens when a habit and a note disagree. Hold three things fixed: one agent, weights fine-tuned on version 1 of a set of apps, and notes written from version 2, where some screens changed. On version 2, the stale weights alone move the agent −8.4 points below the base agent, and the fresh notes alone move it +3.3. Serving both scores 2.7 to 12.7 points below notes alone, so here the stale habit drags down the fresh note. (When the weights are current, serving notes on top can still help: it added +3.5 points for state facts and only +0.1 for locators.) Now look at the first step where the two versions disagree. For low-recurrence items, the agent follows the note 52% of the time, the weights 34%, and neither 14%. For high-recurrence items it follows the note 43%, the weights 21%, and neither 36% of the time. 36 ÷ 14 ≈ 2.6: for the items that recur most, the agent takes neither the habit's action nor the note's about 2.6× as often. In the office: the stapler moved, a sticky note says where, and the hand still reaches for the old drawer, or does something else entirely.

Where this earns its keep: any agent that improves from its own traces, such as a coding agent, a browser agent or a support agent, faces the same choice, and the default of "fine-tune on the good trajectories" or "retrieve the whole trajectory" mixes the four kinds of knowledge together. The study is limited to GUI agents on Android apps with 7–8B backbones, and its rule weights were fitted on that setting. The transferable part is the method: score each piece of experience on recurrence and state-conditionality before you choose a destination, and keep the volatile, screen-specific pieces in context, where they are cheap to update.

Goes deeper in: AI Agents → Context Engineering → The 4 Fixes

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based