GPU·

TACO fine-tunes 32B models on one H100 — One-sparse optimizer state — What does it mean?

The news. On October 1, 2026, researchers from the University of Central Florida and Mohammed VI Polytechnic University posted TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning on arXiv. On OPT-13B, the paper reports 174× less persistent optimizer state than 8-bit AdamW (27.7 GB to 0.16 GB) and 2.9× lower peak training memory (80.6 GB to 27.5 GB). It also reports full-parameter fine-tuning of OPT-30B and Qwen3-32B on a single 80 GB H100. Read the paper →

Picture the school office again. AdamW keeps two notebooks on every single student: one records the direction a student has been drifting (the moment m), the other records how jumpy that drift has been (the moment v). Every weight in the model gets both, in FP32, so the notebooks alone cost 8 bytes per weight. The paper's own arithmetic for a 13B model: the weights take about 26 GB in BF16, the two moment buffers add about 104 GB, and the gradients add another 26 GB — roughly 156 GB before activations, far past the 80 GB of HBM where the model has to live.

The usual fixes shrink the notebooks, but the update still moves every weight. 8-bit AdamW writes them in smaller handwriting (1 byte per number instead of 4). Adafactor keeps a row summary and a column summary of v instead of the full table. GaLore compresses the notebooks into a much smaller set of directions (a low-rank subspace). Each still produces a dense update that can move every weight on every step.

TACO starts from a different question: if each update only ever touches one weight per column, how much history does the optimizer actually need? The answer is: very little. In the basic version, each step, for every column of a weight matrix (every classroom), TACO looks at the gradient, finds the single entry with the largest absolute value, and moves that one weight by a fixed amount in the opposite direction of its sign. Every other weight in the column stays still. Each update value is one of three numbers, +1, 0 or −1 times a step size, which is why the paper calls it ternary, and each column has at most one non-zero, which is why it is one-sparse.

This is not an arbitrary shortcut. Every optimizer is the answer to "what is the best small step?" for some way of measuring step size. AdamW's answer comes from one measuring rule; Muon uses a spectral rule. TACO picks a rule that gives each column its own budget for the sum of absolute changes to its weights. The cheapest way to spend that budget is to put all of it on the weight whose gradient is largest, so the exact best step under that rule is the pick-the-biggest-entry-per-column update. The paper proves a convergence guarantee for the basic, history-free version, and shows that in a simplified fixed-input setting TACO and an Adam-style proxy move within the same set of parameter changes and, when both reach the same output, land on the same solution, although their step-by-step paths need not match; Muon's spectral rule does not share that set. That matters for fine-tuning: the authors cite earlier work showing that switching an AdamW-pretrained model to Muon can hurt accuracy.

One problem remains: picking the winner from a single noisy minibatch gradient is jumpy, like choosing the class standout from one quiz. A dense running average would fix that but bring the full notebooks back. TACO keeps a shortlist instead of notebooks: a running average of the gradient that weights recent steps more (an exponential moving average with decay 0.95) where only the 16 largest entries per column survive, each stored as a 1-byte FP8 E4M3 value plus a 4-byte row index. In this practical version, the winner each step is picked from the fresh gradient blended with that shortlist, not from the raw gradient alone. FP8 is enough because the shortlist only has to rank entries, not scale every coordinate the way an Adam moment does.

The peak-memory number also depends on a systems trick the paper is careful to call standard: each matrix's dense gradient is used and released as soon as backpropagation produces it (PyTorch autograd hooks, no gradient accumulation), and gradient checkpointing trims activations. Vectors and scalars, such as biases and norm parameters, still use ordinary AdamW.

Worked example — one weight matrix. Hold three things fixed: a square 5,120 × 5,120 projection matrix (illustrative; 5,120 is OPT-13B's hidden width), FP32 Adam moments, and TACO's 16 retained entries per column. The matrix has 26,214,400 weights. AdamW's state is 26,214,400 × 8 bytes = about 210 MB. TACO's state is 5,120 columns × 16 entries × 5 bytes = 409,600 bytes, about 0.41 MB. That is 512× smaller for this one matrix, and the ratio grows with matrix height because TACO's cost depends only on the number of columns. Each TACO step changes 5,120 of those 26.2 million weights, one per column. Across the whole OPT-13B fine-tune the paper measures 0.16 GB of persistent optimizer state and 27.5 GB peak, where the weights alone take 25.7 GB.

OptimizerWhat it keeps per weight matrixOPT-13B memory (SST-2 run)Source
AdamWTwo dense FP32 buffers, 8 bytes per weight~104 GB of moments (paper's estimate for 13B, not a measured run)TACO §1
8-bit AdamWTwo dense 8-bit buffers27.7 GB persistent stateTACO abstract
AdafactorRow and column factors of the second momentSmallest persistent state, but 52.6 GB peak from transient memoryTACO §6
GaLoreState projected into a low-rank subspace56.7 GB peak, 31 tokens/sTACO §6
TACO16 FP8 values + int32 row indices per column0.16 GB persistent, 27.5 GB peak, 276 tokens/sTACO §6

The memory saving has a task-dependent accuracy cost. On SST-2 the gap is about a point: 94.2% for TACO against 95.5% for Adafactor and 95.7% for GaLore, at 276 tokens per second against FlashAdamW's 547. On harder tasks in the paper's OPT-13B table the gap to the best baseline is larger: BoolQ 71.5% against Adafactor's 82.0%, and MultiRC F1 61.2 against Adafactor's 75.1. On BoolQ, TACO still beats 8-bit AdamW (62.1%) and FlashAdamW (64.8%), and 8-bit AdamW runs out of memory on MultiRC and DROP. All runs use a 2,000-step budget, and a separate ablation on OPT-1.3B shows TACO still improving at 16,000 steps, so these numbers are not its ceiling. The experiments are all fine-tuning runs; the paper does not test pretraining from scratch.

Goes deeper in: GPU & CUDA → Memory Hierarchy → HBM: Where Your Model Lives

For why bytes per weight decide what fits, see The Size Problem in the quantization module.

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based