AI chips are only as fast as the memory feeding them. “Custom HBM” changes where the work of talking to that memory happens, freeing space on the processor itself. Here is what it is, what NVIDIA claims for its version, called NVHBM, and what is still unproven.
Spotted via Yumzlef (@Yumzlef) on X.
First, what is HBM?
High-bandwidth memory (HBM) is a tower of memory chips stacked on top of one another and placed right next to an AI processor inside the same package. Because it is so close and so wide, it can move data far faster than ordinary computer memory. Large AI models often spend more time waiting for data than doing sums. TSN looked at how much of it Nvidia’s Rubin chips need in Nvidia’s new AI chip needs 288GB of RAM.
At the bottom of every HBM stack sits a “base die”, a layer of logic that connects the memory to the processor. On standard HBM, the processor still carries the memory controller (the circuitry that organises reads and writes) and a wide physical connection, or “PHY”, to reach the stack. Both take up space on the processor’s most expensive silicon.
What makes HBM “custom”?
In custom HBM, the chip designer supplies its own base die. Some of the logic that used to sit on the processor, notably the memory controller, moves down into the memory stack, and the wide standard connection is replaced with a narrower, purpose-built one.
NVIDIA’s version, announced on 26 August 2026, “integrates NVIDIA’s custom memory controller into the HBM base die” [1]. It is offered to customers of NVLink Fusion, NVIDIA’s programme for companies building their own AI chips (often called XPUs). NVIDIA says “Amazon’s Annapurna Labs will be the first to work on NVHBM” [1].
What NVIDIA claims
Company claims, not independently tested. NVIDIA’s blog says: “NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area on XPU compute die compared with standard HBM4E” [1]. HBM4E is the next standard version of HBM. The comparison is with that standard memory, not with any current Nvidia chip. NVIDIA’s technical blog gives the same “up to 30%” and “up to 25%” figures in its summary table [2].
What an analyst estimates
Reported: SemiAnalysis’s estimate, not NVIDIA’s figure. The research firm SemiAnalysis estimates that on Nvidia’s Rubin, “HBM controllers and PHYs take up roughly 16% of the GPU die. On Feynman with NVHBM, we estimate that falls to ~4%” [3][4]. Feynman is Nvidia’s next-but-one GPU generation. It is unreleased: Nvidia’s own roadmap puts it in 2028 with “a custom HBM”, as reported from its March conference [5]. Nvidia has not said Feynman will use NVHBM by name; that link is SemiAnalysis’s.
Why it matters
Space on a leading-edge processor is scarce and costly. If less goes on memory plumbing, designers can add more compute or cache in the same package. It also deepens the ties between chip designers and the memory makers who build the stacks, a supply chain TSN has followed through packaging and Samsung’s memory business.
What this does not prove
- That NVHBM delivers those gains in real chips. The bandwidth, area and power figures are NVIDIA’s own “up to” claims against standard HBM4E, with no independent tests yet [1][2].
- That the 16% and 4% figures are Nvidia’s. They are SemiAnalysis estimates; TSN could not read the paid note behind them [3][4].
- Anything about Feynman’s final design. It is unreleased, and 2028 is a roadmap date [5].
- That AI services will get faster by a set amount. No real-world latency or speed-up figure has been published for NVHBM-based chips.
The Bottom Line
Custom HBM moves the memory controller off the AI processor and into a customer-designed base die under the memory stack. NVIDIA claims its NVHBM gives “up to 30%” more bandwidth and frees “up to 25%” more compute-die area compared with standard HBM4E, and SemiAnalysis estimates the share of die spent on memory interfaces could fall from about 16% to about 4%. Both are promises and estimates until chips ship and are tested.
Sources
- Jesse Clayton, “NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory”, NVIDIA Blog, 26 August 2026 (company claims; read directly). https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/
- Farshad Ghodsian, Jesse Clayton and Varun Nanda Kumar, “NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure”, NVIDIA Technical Blog, 26 August 2026 (company claims; read directly). Note: elsewhere the post also cites “up to a 30% increase in available main-die silicon”; TSN uses the 25% figure that appears in both NVIDIA posts. https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/
- SemiAnalysis (@SemiAnalysis_), X thread, 3 October 2026, 04:00 BST (analyst estimate; read via the X API). https://x.com/SemiAnalysis_/status/2106217837643575607
- SemiAnalysis (@SemiAnalysis_), X thread, 1 October 2026, 22:27 BST (“We estimate that HBM4 controllers and PHYs take up roughly 16% of Nvidia’s Rubin compute die”; read via the X API). The underlying Memory Model note is paywalled and was not read. https://x.com/SemiAnalysis_/status/2105771525605351710
- HPCwire, “Huang Shares Nvidia Roadmap Showing More Chips, NVL1152 Scale-Up, CPO”, 17 March 2026 (press report of Nvidia’s GTC 2026 roadmap). https://www.hpcwire.com/2026/03/17/huang-shares-nvidia-roadmap-showing-more-chips-nvl1152-scale-up-cpo/
- Yumzlef (@Yumzlef), X post, 8 October 2026, 13:51 BST (where TSN first saw the story; read via the X API). It goes beyond the primary sources: it drops NVIDIA’s “up to” from the 30% bandwidth figure in its opening line, and adds results from the author’s own simulation, which the post itself describes as a “made-up model, chip and traffic”. TSN does not use those simulation results. https://x.com/Yumzlef/status/2108178553661477317

