LLM·

EngramEdit edits LLM facts in conditional memory — Knowledge editing in n-gram memory — What does it mean?

The news. On October 7, 2026, researchers from The Hong Kong Polytechnic University and collaborators posted EngramEdit to arXiv. They edit facts in LongCat-Flash-Lite, a 68.5B-parameter mixture-of-experts (MoE) model that keeps 31.4B of those parameters in n-gram embedding tables, and they add separate update vectors keyed by exact n-grams while keeping the pretrained tables and backbone fixed. After 2,000 sequential edits on the CounterFact and ZsRE benchmarks the authors describe editing success as near-perfect, and on the multi-hop MQuAKE benchmark they report nearly 3× the strongest baseline's accuracy when the model reasons step by step. Read the paper →

Picture a librarian who answers questions in two ways. Some answers come from reasoning. Others come from index cards, and the cards are filed by the last few words of your request, not by topic. In an ordinary transformer there are no cards, and earlier editing methods treat the weights of the feed-forward networks as the place a fact is recalled, so changing one fact means surgery on the librarian. Models built with conditional memory do have a card catalog, and an ablation in DeepSeek's Engram work, which the EngramEdit authors cite, found that switching the catalog off hurts factual-knowledge benchmarks much more than it hurts reasoning about text given in the prompt. If facts lean heavily on the cards, a fact can be changed by correcting cards while the librarian stays exactly as trained. EngramEdit does not even rewrite the cards: it clips a correction slip to each card it changes, so the original catalog never changes either. The catch is the filing system. A card is found by phrase, so one fact is spread over several cards, and one card can serve many facts.

Mechanically, a memory row is an ordinary embedding lookup, keyed by an n-gram instead of a single token id. EngramEdit keeps the two-step shape of locate-then-edit. Step one finds what the memory should output; step two solves for separate n-gram update vectors that produce it. For step one, the model first writes four rewordings of each fact, varied in the words just before the subject. A single learnable change is added to the memory output at the last subject token of every wording and optimized until the model predicts the new answer, with the backbone frozen. That gives a target memory representation per wording. For step two, each target is mapped back to the n-grams its wording actually reads, and one linear system is solved so that each shared n-gram gets a single update that serves every wording and every edit in the batch. The system includes a penalty that grows for shorter or more frequent n-grams, because those are read by many unrelated inputs. Each update is stored as its own vector, keyed by the exact token sequence and added to the memory output whenever that n-gram appears, so the pretrained tables stay fixed and hash collisions in them cannot couple two edits.

Inputembeddings
Thecatsatonthemat
Attentioncontext mixing
+ Residualadd & norm
Feed-Forwardtransform
Outputto next layer

Where the edit has to land

Hold the paper's setup fixed. At each token, LongCat-Flash-Lite reads the n-grams of length 2, 3 and 4 that end there, each through four hashed sub-tables (four separate tables, each indexed by a hash of the n-gram), projects the twelve vectors into token-embedding space and averages them with the token's own embedding: 13 vectors in one average. Take the paper's example fact, changing Alex's employer from Lab M to Lab N, and two wordings that end on the subject (illustrative wordings, one token per word): "employer of Alex Kim" and "who hired Alex Kim". At the token "Kim", both read the 2-gram "Alex Kim", but their 3-grams and 4-grams differ. So an edit made only through phrasing A reaches phrasing B through only 1 of its 3 n-grams, whose four vectors are 4 of the 13 in the average (ignoring hash collisions). And "Alex Kim" is exactly the n-gram that every other question about Alex Kim also reads. Generalization pulls the edit toward short, shared n-grams; specificity pushes it away from them. The reuse penalty decides that split. The paper's own baseline shows the cost of ignoring it: fine-tuning only the memory read at the subject token in the original prompt reaches about the same success on the edited prompt as EngramEdit on CounterFact, but generalizes to held-out paraphrases much worse.

ApproachWhat it changesBackbone touched?Main weakness
Fine-tuning (FT)model weights by gradient descent on the new factyeschanges can spread to unrelated facts and skills
Locate-then-edit on FFNs (MEMIT-style; MoEEdit for MoE models)selected FFN or expert weights, solved in closed formyesshared weights can affect unrelated inputs
Memory fine-tuning, subject token onlymemory read at the last subject token of the original promptnomuch weaker on held-out paraphrases
EngramEditper-n-gram updates for 5 wordings, solved jointly with a reuse penaltynoneeds a model built with conditional memory; tested on one model

Two results show that the separate memory updates carry the revised facts. When the authors switch off only the updates belonging to one fact, success on that fact drops sharply, nearly as far as switching off every update, while switching off a matched number of random updates leaves success on the edited prompts and their paraphrases unchanged. And the multi-hop gain appears only when the model reasons step by step: under direct answering EngramEdit is level with plain fine-tuning, but with chain-of-thought it reaches nearly 3× the strongest baseline. The authors' explanation is that each written-out reasoning step can activate the edited n-grams again, so a phrase-keyed edit gets more chances to fire when the reasoning puts the right words on the page. The setup also marks the limits: one model (an MoE with a large n-gram table), up to 5,000 sequential edits, and nothing for models without conditional memory.

Goes deeper in: LLM Internals → Transformer Block → The Feed-Forward Network

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based