Threadripper Halo Station: Specs, 576GB HBM3E & Price (2026)
AMD's Threadripper Halo Station: 96-core CPU, 4 MI350P GPUs, 576GB HBM3E, 16TB/s bandwidth, targets trillion-param models at 4-bit. Price, date unannounced.
On September 4, 2026, AMD unveiled the Threadripper Halo Station, a liquid-cooled desk-side AI workstation, during its IFA 2026 keynote in Berlin. The system pairs a Threadripper PRO 9995WX CPU (96 cores/192 threads) with up to four AMD Instinct MI350P accelerators, delivering up to 576GB of HBM3E GPU memory. AMD calls it 'the most powerful workstation in the world' and says a 4-bit quantized trillion-parameter-class model can fit within GPU memory. Pricing, availability, and detailed real-world performance have not been announced as of this writing; this article separates confirmed facts from press speculation.
Key Specs at a Glance
| Item | Details |
|---|---|
| CPU | AMD Threadripper PRO 9995WX (Zen 5, 96 cores/192 threads, up to 5.4GHz boost) |
| System memory | Up to 2TB DDR5 |
| Accelerator | Up to 4x AMD Instinct MI350P (PCIe cards) |
| GPU memory | 144GB HBM3E per card, up to 576GB total with 4 cards (IFA demo unit used 2 cards = 288GB) |
| GPU memory bandwidth | 4TB/s per card, up to 16TB/s combined |
| Combined memory (with system RAM) | Up to roughly 2.6TB |
| Cooling | Liquid-cooled, desk-side form factor |
| Price | Not announced (reports speculate $100,000-$150,000) |
| Availability | Not announced (reports speculate 2027 shipping) |
If you want to check whether a specific model will run on your own GPU or Mac, our free VRAM calculator estimates required VRAM just by picking a model, quantization level, and context length. For background on sizing local hardware, see our local AI hardware requirements guide.
How It Compares to the NVIDIA DGX Station
| Item | Threadripper Halo Station | DGX Station (GB300 Ultra) |
|---|---|---|
| CPU | Threadripper PRO 9995WX (96 cores) | NVIDIA Grace-based CPU |
| GPU memory | Up to 576GB HBM3E | 252GB HBM3E |
| GPU memory bandwidth | Up to 16TB/s | 7.1TB/s |
| CPU-side memory | Up to 2TB DDR5 | 496GB LPDDR5X |
| Total coherent memory | ~2.6TB (per AMD) | 748GB |
| CPU-GPU interconnect | Not announced | NVLink-C2C, ~900GB/s |
| Reported street price | Not announced | ~$85,000 (reported MSI resale configuration) |

AMD claims '3.4x system memory and more than 2x memory bandwidth' versus the DGX Station, but these figures describe hardware capacity and bandwidth, not guaranteed inference throughput. A larger memory pool doesn't automatically translate into faster real-world performance — the maturity gap between ROCm and the CUDA ecosystem, along with overhead from model sharding and offloading, can significantly affect actual results. This tension is also discussed in our comparison of local LLM inference engines.
Can It Really Run a Trillion-Parameter Model?
Let's sanity-check AMD's claim that a trillion-parameter-class model can fit in GPU memory at 4-bit quantization. At 4 bits (0.5 bytes per parameter), a model's weight footprint is roughly parameter count x 0.5 bytes. For a trillion-parameter model, that's approximately 500GB just for the weights — consuming most of the 576GB of GPU memory. On top of that, inference requires a KV cache, whose size grows with context length and batch size, leaving limited headroom for long contexts or concurrent requests. How much actually fits depends heavily on the model architecture (dense vs. MoE) and context length settings, so the figures below are rough estimates only. For more on the relationship between KV cache and context length, see our KV cache and VRAM guide.
| Item | Rough estimate | Notes |
|---|---|---|
| 1T parameters at 4-bit | ~500GB | Simple parameter count x 0.5 bytes |
| Total GPU memory (4-card config) | 576GB | AMD's stated figure |
| Remaining after weights | ~76GB | For KV cache and other overhead |
| KV cache at short context | A few GB to tens of GB | Depends on model architecture and batch size |
| Long context / large batch | Can consume remaining headroom quickly | May require offloading to system RAM |
Who It's For — and Who It Isn't
Research institutions and universities that need on-premises infrastructure to evaluate large-scale models
IT departments at companies that can't send sensitive data to external clouds and need local fine-tuning or inference infrastructure
Large enterprises building shared internal AI infrastructure that must serve multiple large models simultaneously
Individual developers or small teams doing experimentation — on-demand cloud GPUs are typically more cost-effective
General business use cases where models in the tens or low hundreds of billions of parameters suffice — a Mac Studio or consumer RTX-class GPU is often enough (see our piece on running huge MoE models with modest memory)
Organizations that need to finalize a budget now — pricing and availability aren't public yet, so planning is premature
What We Still Don't Know
Official pricing (the $100K-$150K figures are speculative, not confirmed)
Whether and when it will be sold in Japan
How well major frameworks (PyTorch, vLLM, etc.) run on ROCm, and real-world throughput
Detailed power consumption, electrical requirements, and footprint
Actual benchmark results with a full 4-card MI350P configuration
Whether the rumored 2027 shipping timeline holds
How much does the Threadripper Halo Station cost?
As of September 2026, AMD has not announced official pricing. Some outlets have speculated a range of $100,000 to $150,000, but this is an unconfirmed estimate, not official information.
When will it be available?
AMD has not announced an official release date. Some press reports speculate a 2027 shipping timeline, but this remains unconfirmed.
Should I buy this or an NVIDIA DGX Station?
Looking at raw GPU memory capacity and bandwidth alone, the Threadripper Halo Station appears to have an edge, but the DGX Station benefits from the maturity of NVIDIA's CUDA ecosystem. The right choice depends on your existing software stack and tolerance for migrating to ROCm — there's no clear-cut winner at this stage.
The MI350P doesn't support CUDA — what software runs on it?
AMD Instinct MI350P runs on AMD's open-source ROCm software stack rather than NVIDIA's CUDA. Major frameworks like PyTorch and vLLM are adding ROCm support, but whether that support matches CUDA's maturity and track record still needs to be verified in practice.
Can it really run a trillion-parameter model?
AMD says a 4-bit quantized trillion-parameter-class model can fit within its GPU memory, but that figure is a rough estimate based on weight size alone — it doesn't account for KV cache overhead or guarantee real-world inference speed. Detailed benchmarks haven't been published, so actual performance remains unknown.
Related free tools (no sign-up, instant results)
Feel free to contact us
Contact Us