A secondary summary circulating on X on 7 October 2026 framed NVIDIA’s LoGRA as cutting average reinforcement-learning (RL) training memory by up to 45.7% on tested reasoning tasks while keeping performance in the reported setups [2]. That figure matches the paper’s own headline claim. The primary source is the arXiv preprint LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches (arXiv:2610.06647, 5 October 2026), by Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz and Yi Dong, with a © 2026 NVIDIA notice on the PDF [1].
This explainer stays with the paper. The tweet is a secondary précis, not evidence. Numbers below are grounded in the abstract, Figure 1, Table 1, Table 2 and the method sections of the PDF.
Why RL post-training memory hurts
Serving a model and improving it with RL are different memory problems. Inference mainly needs weights (and activations for the tokens you generate). RL post-training also needs gradients and optimizer state. The LoGRA authors put the gap in concrete terms for a 7B-parameter Qwen-class model: BF16 weights are about 14 GB, while Adam’s two FP32 moment buffers add another 56 GB—before gradients and activations [1].
System tricks such as FSDP sharding, activation checkpointing and CPU offloading help. Parameter-efficient fine-tuning (for example LoRA) freezes most weights and trains small adapters. LoGRA takes a different route: keep the original weight matrices trainable, but never materialise full gradients for the selected layers. Compress the learning signal into a compact sketch, update from that sketch, and reuse the same compact factors when syncing the rollout policy [1].
The practical point for labs and self-hosters is blunt. Hardware that can run a model may still be too small to RL-improve it under a dense Adam recipe. LoGRA’s claim is that this barrier can be lowered without abandoning full-weight updates on the matrices that matter most for the reported recipes.
The authors’ intuition is worth stating in one sentence: with binary outcome rewards, each generated response often yields only a single correctness signal, so the useful learning direction may need far less information than a full dense gradient can represent [1]. That is a motivation, not a theorem that every RL gradient is low-rank—but it explains why they try sketches first rather than always storing (G) in full.
What a low-rank gradient sketch is (in plain words)
Take one weight matrix (W) of size (d \times k). Its full gradient (G) is the same shape. Storing (G) costs (dk) values. LoGRA picks a much smaller rank (r \ll \min(d,k)) and a random projection matrix (A) of size (r \times k). It represents the gradient by the sketch
[
S = G A^{\top} \in \mathbb{R}^{d \times r}.
]
That stores (dr) values instead of (dk). In the paper’s example, for (d = k = 4096) and (r = 64), an fp32 gradient buffer falls from 64 MiB to 1 MiB for that matrix [1]. The projection (A) is not learned; by default its entries are independent random signs scaled by (1/\sqrt{r}), and it can be refreshed each step so successive updates are not trapped in one fixed subspace.
Importantly, the implementation does not have to build a full (G) and then squash it. Sketches are accumulated during backpropagation by projecting the inputs used in the weight-gradient calculation before they meet the backward derivatives [1].
To apply an update, the sketch is expanded back to weight size. Writing (U) for the (possibly optimizer-adjusted) sketch—(U = S) for plain SGD—the low-rank update is
[
W \leftarrow W – \alpha\eta\, U A.
]
Here (\eta) is the nominal learning rate and (\alpha) is an extra scale factor (1 leaves the proposal unchanged; smaller values shrink it). The result is merged straight into the full weight matrix. There is no persistent adapter branch sitting beside a frozen base, which is the first sharp distinction from LoRA [1].
LoGRA applies this to attention and MLP projection matrices and holds other parameters fixed. The same compact factors are reused for rollout-policy synchronisation: the trainer sends (\alpha\eta U) plus a seed for (A); the generation side regenerates (A) and applies the same update. The communication payload is (O(dr)) values plus metadata, not a full (d \times k) delta [1]. Code is pointed at from the paper in the Molt library: https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra [1].
Predicted-KL step control
Compressing the gradient does not, by itself, decide how large a step is safe. An update that looks modest in parameter space can still move next-token probabilities a lot. Following the trust-region flavour of TRPO, LoGRA adds predicted-KL step control: estimate how much a proposed update would change the policy’s token probabilities before applying it, then choose (\alpha) so the predicted KL fits a budget [1].
The paper notes that predicted KL scales with the square of the step multiplier—halving (\alpha) cuts predicted KL by about four—so the controller can shrink (or, within a cap, enlarge) the proposal to meet the budget [1]. In the main experiments the budget is annealed from (2 \times 10^{-4}) to (2 \times 10^{-5}). This is an estimate used for control, not a guarantee that realised KL will match the prediction on every step. Treat it as a trust-region-style brake, not as a proof of identical behaviour to dense Adam at every update [1].
What the numbers show (Table 1 and the abstract)
Default hardware in the paper is a single node with eight H100 80GB GPUs and 128 CPU cores: four GPUs for FSDP training, four for independent generation engines [1]. Main comparisons use Qwen2.5-Math-1.5B/7B and Qwen3.8-27B on mathematical reasoning (DAPO-Math-7.5K training, MATH-500 evaluation), with a deliberately simple PPO-style setup (one response per prompt, global running reward baseline). The authors say that simple setup is intentional: it isolates the effect of the memory method from gains that a more elaborate RL algorithm might add [1].
LoGRA uses rank-256 RowAdam with mismatch subtraction and predicted-KL annealing; the dense baseline uses Adam with learning-rate annealing. RowAdam is not identical to Adam—it adapts scale per row of the sketch rather than per element—but it keeps adaptive scaling so the comparison is not “sketch SGD versus Adam” in a vacuum [1]. Three seeds. Accuracy peaks in Table 1 use updates 0–200 (Figure 2 plots that same early window); memory and throughput cover the full runs (1.5B/7B were trained for far longer schedules; 27B main MATH-500 runs stop at 200 updates) [1].
Verified from Table 1 / abstract / Figure 1 [1]:
- 1.5B. Average memory 9.18 → 7.18 GiB/GPU (−21.8%). Peak Pass@1: LoGRA 67.87% vs dense 63.77% (also peak Pass@4 82.47% vs 80.33%).
- 7B. Average memory 31.82 → 17.29 GiB/GPU (−45.7%). Peak Pass@1 roughly matched: LoGRA 72.33% vs dense 72.48% (Pass@4 85.13% vs 85.33%).
- 27B. Dense Adam OOMs at the first Adam-state allocation. LoGRA averages 51.54 GiB/GPU (peak seed-averaged 53.88 GiB) and, in the main short MATH-500 window, reaches 71.52% Pass@1 / 81.87% Pass@4. Separately, on the Reasoning-Gym Hard mixture, LoGRA sustains learning beyond 1,100 logged steps on one eight-GPU node—feasibility plus a learning trajectory, not a paired dense baseline at that horizon [1].
Throughput where both methods finish is similar (about 55 vs 54 updates/h at 1.5B; about 26 vs 25 at 7B). The headline systems win is lower memory at 1.5B/7B and feasibility at 27B under the same node budget [1].
The oft-quoted 45.7% figure is the 7B average-memory saving versus dense Adam in that setup—not a universal constant for every model, task or horizon [1].
LoGRA is not LoRA (brief comparison)
LoRA freezes pretrained weights and learns low-rank adapters. LoGRA compresses gradients and writes updates into the full weights for the selected matrices [1].
Table 2 compares rank-256 recipes on Qwen2.5-Math-1.5B for the first ~300 training steps: LoGRA average memory 7.18 GiB vs LoRA 13.21 GiB (peak 8.66 vs 13.38). Under those recipes LoRA’s mean Pass@1 is higher (70.45% vs 68.47%) and its policy-training call is slightly faster; LoGRA’s mean Pass@4 is a touch higher (81.80% vs 81.40%). The paper’s own reading is the right one: this motivates gradient compression when training memory is the bottleneck; it does not establish a general accuracy or speed win over LoRA [1].
Why it matters for labs and self-hosters
If your wall is “the model fits for serving but RL OOMs on Adam state,” LoGRA is aimed at that wall. On the paper’s 27B setup, dense Adam never completes an update; LoGRA runs at ~51.5 GiB average per training GPU on the same 8×H100 node and keeps learning for 1,100+ steps on Reasoning-Gym Hard [1]. Smaller memory on 1.5B and 7B, with Pass@1 that is better (1.5B) or comparable (7B) in the reported MATH-500 peaks, is the supporting evidence that the compression is not empty [1].
For anyone running open-weight post-training, that combination—full-weight updates on the big projection matrices, smaller gradient/optimizer footprint, smaller sync payload, and an explicit KL-budget brake—is worth reading primary, not just the Twitter summary. The paper also reports end-to-end update rates that stay comparable where both methods finish, which matters if you feared that sketching would trade memory for a large slowdown in those configs [1]. Implementation is in the Molt library path cited above; reproduce under your own stack, with your own FSDP/generation split and reward setup, before treating any percentage as transferable [1].
What this does not prove
- That every model, task or training horizon sees a 45.7% memory cut. That number is the 7B average-memory saving vs dense Adam in Table 1 / the abstract [1].
- That LoGRA beats dense Adam (or LoRA) on accuracy in general. At 7B, peak Pass@1 is essentially matched; the LoRA comparison is recipe-specific and not claimed as a general quality win [1].
- A paired dense baseline for the long 27B Reasoning-Gym run. Dense OOMs; the long run shows LoGRA feasibility and a learning curve, not a head-to-head scoreboard [1].
- That predicted KL equals realised KL. The controller uses an estimate; realised divergence can differ [1].
- That LoGRA replaces LoRA for every use case. Different mechanisms: gradient sketches into full weights versus frozen base plus adapters [1].
- Investment, product-roadmap or “NVIDIA shipping this in product X” claims. This draft covers a research preprint and secondary social summary only [1][2].
The Bottom Line
LoGRA is a memory-first RL post-training method: keep learning signals as low-rank gradient sketches (S = G A^{\top}), apply (W \leftarrow W – \alpha\eta\, U A) on attention/MLP projections, sync rollouts with the same compact factors, and scale (\alpha) with a predicted-KL budget [1]. On the authors’ Qwen math / Reasoning-Gym setups with 8×H100, it cuts average training memory by 21.8% at 1.5B and 45.7% at 7B, matches or improves the reported MATH-500 peaks in that window, and makes 27B RL runnable where dense Adam OOMs—including 1,100+ steps on Reasoning-Gym Hard [1].
Read it as a carefully scoped systems result from a NVIDIA-affiliated preprint, not as a universal training free lunch. Primary source first; social summaries second.
Sources
- Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong, “LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches,” arXiv:2610.06647, 5 October 2026 (PDF © 2026 NVIDIA; code pointed to Molt library). https://arxiv.org/abs/2610.06647 — PDF: https://arxiv.org/pdf/2610.06647 — code path referenced from the paper: https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra
- Mark Kretschmann (@mark_k), post on X summarising LoGRA’s reported average RL training-memory cut (up to 45.7% on tested reasoning tasks), 7 October 2026 — secondary summary only; not a primary source for numbers. https://x.com/mark_k/status/2107880399862939929
