Tag: reinforcement learning

NVIDIA’s LoGRA paper shows low-rank gradient sketches cutting average RL training memory by up to 45.7% on tested Qwen reasoning setups—and making 27B RL runnable on one 8×H100 node where dense Adam OOMs. What the method is, what Table 1 shows, and what it does not prove.

No posts to display

Recent articles