AMD opened its IFA 2026 keynote in Berlin on Thursday with a tower. The Threadripper Halo Station pairs a 96-core Threadripper Pro 9995WX with two liquid-cooled Instinct MI350P accelerators — each carrying 144GB of HBM3E at 4TB/s — 2TB of DDR5, and a documented path to four cards and 576GB of high-bandwidth memory. AMD’s claim is that it runs models with more than a trillion parameters locally, with nothing plugged into it but power.
Pricing was not announced. Component-level estimates from reviewers who priced the parts individually put the core configuration north of $100,000, and a fully specified system with storage, power and cooling closer to $150,000. That is not a consumer number and it is not meant to be one. The target is the machine sitting under a researcher’s desk at a lab that will not put its data in someone else’s building — the same slot Nvidia’s DGX Station occupies.
The specification that matters is the memory, not the cores. Frontier-scale inference is a memory-bandwidth problem before it is a compute problem, and 288GB of HBM3E across two cards — 576GB across four — is roughly the amount of fast memory it takes to hold a very large model resident without spilling to system RAM. AMD is explicit about the comparison: HBM3E at 4TB/s is on the order of fourteen times the bandwidth of the LPDDR5X used in the unified-memory desktops that have been marketed for local AI over the past two years. Those machines fit big models. They do not run them quickly.
Why a desk and not a rack
The economics of a $150,000 workstation only work if cloud inference is unavailable rather than merely expensive. For most teams it is available and cheap, and getting cheaper — Microsoft just priced transcription at ten cents an hour. The buyers here are the ones for whom the API is the problem: defence contractors, hospital systems, banks, national labs, and any organisation whose legal department has decided that inference on regulated data happens on premises or not at all.
That is a real and growing market, and it is one AMD has been winning in quietly. Europe’s €388m LUMI-AI order contains no Nvidia silicon at all. Sovereignty buyers care less about ecosystem maturity than about supply chains they can audit, and AMD has spent three years making that argument.
Our take: The interesting number is not $150,000, it is 576GB. AMD has quietly established that the ceiling on local inference is no longer “a 70-billion-parameter model, slowly.” When the deskside ceiling moves, the set of workloads that were only ever going to run in a hyperscaler’s building gets smaller. Not by much this year. But the direction has reversed.
What to watch
- Actual pricing and availability. AMD gave neither. A spec sheet without a ship date is a positioning exercise; a ship date makes it a product.
- Whether four MI350P cards is real. Two cards at 600W TBP each, plus a 350W CPU, is already a serious thermal and power problem for a machine on a floor rather than in a rack. Four is a different building.
- Nvidia’s response. The DGX Station is the incumbent in this niche and Nvidia rarely leaves a memory-capacity claim unanswered for long.
- ROCm. The hardware argument has been winnable for AMD for a while. The software argument is the one that has cost it deals.
The pitch AMD is making is not that this box is cheaper than the cloud. It is that for a certain kind of buyer, the cloud was never an option, and until now the on-premises alternative meant a rack, a room and a procurement cycle. Berlin’s answer is a tower and a wall socket.
