Skip to main content
株式会社オブライト
AI2026-09-066 min read

Threadripper Halo Station: Specs, 576GB HBM3E & Price (2026)

AMD's Threadripper Halo Station: 96-core CPU, 4 MI350P GPUs, 576GB HBM3E, 16TB/s bandwidth, targets trillion-param models at 4-bit. Price, date unannounced.


On September 4, 2026, AMD unveiled the Threadripper Halo Station, a liquid-cooled desk-side AI workstation, during its IFA 2026 keynote in Berlin. The system pairs a Threadripper PRO 9995WX CPU (96 cores/192 threads) with up to four AMD Instinct MI350P accelerators, delivering up to 576GB of HBM3E GPU memory. AMD calls it 'the most powerful workstation in the world' and says a 4-bit quantized trillion-parameter-class model can fit within GPU memory. Pricing, availability, and detailed real-world performance have not been announced as of this writing; this article separates confirmed facts from press speculation.

Key Specs at a Glance

ItemDetails
CPUAMD Threadripper PRO 9995WX (Zen 5, 96 cores/192 threads, up to 5.4GHz boost)
System memoryUp to 2TB DDR5
AcceleratorUp to 4x AMD Instinct MI350P (PCIe cards)
GPU memory144GB HBM3E per card, up to 576GB total with 4 cards (IFA demo unit used 2 cards = 288GB)
GPU memory bandwidth4TB/s per card, up to 16TB/s combined
Combined memory (with system RAM)Up to roughly 2.6TB
CoolingLiquid-cooled, desk-side form factor
PriceNot announced (reports speculate $100,000-$150,000)
AvailabilityNot announced (reports speculate 2027 shipping)

If you want to check whether a specific model will run on your own GPU or Mac, our free VRAM calculator estimates required VRAM just by picking a model, quantization level, and context length. For background on sizing local hardware, see our local AI hardware requirements guide.

How It Compares to the NVIDIA DGX Station

ItemThreadripper Halo StationDGX Station (GB300 Ultra)
CPUThreadripper PRO 9995WX (96 cores)NVIDIA Grace-based CPU
GPU memoryUp to 576GB HBM3E252GB HBM3E
GPU memory bandwidthUp to 16TB/s7.1TB/s
CPU-side memoryUp to 2TB DDR5496GB LPDDR5X
Total coherent memory~2.6TB (per AMD)748GB
CPU-GPU interconnectNot announcedNVLink-C2C, ~900GB/s
Reported street priceNot announced~$85,000 (reported MSI resale configuration)
Diagram comparing the Threadripper Halo Station memory layout of four MI350P cards with 576GB of HBM3E at 16TB/s plus up to 2TB of DDR5 against the NVIDIA DGX Station with 252GB HBM3E and 496GB LPDDR5X for a 748GB coherent total, set against roughly 500GB of weights for a trillion-parameter model at 4-bit

AMD claims '3.4x system memory and more than 2x memory bandwidth' versus the DGX Station, but these figures describe hardware capacity and bandwidth, not guaranteed inference throughput. A larger memory pool doesn't automatically translate into faster real-world performance — the maturity gap between ROCm and the CUDA ecosystem, along with overhead from model sharding and offloading, can significantly affect actual results. This tension is also discussed in our comparison of local LLM inference engines.

Can It Really Run a Trillion-Parameter Model?

Let's sanity-check AMD's claim that a trillion-parameter-class model can fit in GPU memory at 4-bit quantization. At 4 bits (0.5 bytes per parameter), a model's weight footprint is roughly parameter count x 0.5 bytes. For a trillion-parameter model, that's approximately 500GB just for the weights — consuming most of the 576GB of GPU memory. On top of that, inference requires a KV cache, whose size grows with context length and batch size, leaving limited headroom for long contexts or concurrent requests. How much actually fits depends heavily on the model architecture (dense vs. MoE) and context length settings, so the figures below are rough estimates only. For more on the relationship between KV cache and context length, see our KV cache and VRAM guide.

ItemRough estimateNotes
1T parameters at 4-bit~500GBSimple parameter count x 0.5 bytes
Total GPU memory (4-card config)576GBAMD's stated figure
Remaining after weights~76GBFor KV cache and other overhead
KV cache at short contextA few GB to tens of GBDepends on model architecture and batch size
Long context / large batchCan consume remaining headroom quicklyMay require offloading to system RAM

Who It's For — and Who It Isn't

Research institutions and universities that need on-premises infrastructure to evaluate large-scale models
IT departments at companies that can't send sensitive data to external clouds and need local fine-tuning or inference infrastructure
Large enterprises building shared internal AI infrastructure that must serve multiple large models simultaneously

Individual developers or small teams doing experimentation — on-demand cloud GPUs are typically more cost-effective
General business use cases where models in the tens or low hundreds of billions of parameters suffice — a Mac Studio or consumer RTX-class GPU is often enough (see our piece on running huge MoE models with modest memory)
Organizations that need to finalize a budget now — pricing and availability aren't public yet, so planning is premature

What We Still Don't Know

Official pricing (the $100K-$150K figures are speculative, not confirmed)
Whether and when it will be sold in Japan
How well major frameworks (PyTorch, vLLM, etc.) run on ROCm, and real-world throughput
Detailed power consumption, electrical requirements, and footprint
Actual benchmark results with a full 4-card MI350P configuration
Whether the rumored 2027 shipping timeline holds

How much does the Threadripper Halo Station cost?

As of September 2026, AMD has not announced official pricing. Some outlets have speculated a range of $100,000 to $150,000, but this is an unconfirmed estimate, not official information.

When will it be available?

AMD has not announced an official release date. Some press reports speculate a 2027 shipping timeline, but this remains unconfirmed.

Should I buy this or an NVIDIA DGX Station?

Looking at raw GPU memory capacity and bandwidth alone, the Threadripper Halo Station appears to have an edge, but the DGX Station benefits from the maturity of NVIDIA's CUDA ecosystem. The right choice depends on your existing software stack and tolerance for migrating to ROCm — there's no clear-cut winner at this stage.

The MI350P doesn't support CUDA — what software runs on it?

AMD Instinct MI350P runs on AMD's open-source ROCm software stack rather than NVIDIA's CUDA. Major frameworks like PyTorch and vLLM are adding ROCm support, but whether that support matches CUDA's maturity and track record still needs to be verified in practice.

Can it really run a trillion-parameter model?

AMD says a 4-bit quantized trillion-parameter-class model can fit within its GPU memory, but that figure is a rough estimate based on weight size alone — it doesn't account for KV cache overhead or guarantee real-world inference speed. Detailed benchmarks haven't been published, so actual performance remains unknown.

Feel free to contact us

Contact Us