AI

AMD built a deskside box that runs trillion-parameter models with the network cable unplugged.

The Threadripper Halo Station pairs 96 Zen 5 cores with liquid-cooled Instinct accelerators and up to 576GB of HBM3E. The target isn’t Nvidia’s chips — it’s Nvidia’s DGX Station.

N Noah · The Sharp Brief · September 5, 2026 · 4 min read

AMD opened its IFA 2026 keynote in Berlin on Thursday with a tower. The Threadripper Halo Station pairs a 96-core Threadripper Pro 9995WX with two liquid-cooled Instinct MI350P accelerators — each carrying 144GB of HBM3E at 4TB/s — 2TB of DDR5, and a documented path to four cards and 576GB of high-bandwidth memory. AMD’s claim is that it runs models with more than a trillion parameters locally, with nothing plugged into it but power.

Pricing was not announced. Component-level estimates from reviewers who priced the parts individually put the core configuration north of $100,000, and a fully specified system with storage, power and cooling closer to $150,000. That is not a consumer number and it is not meant to be one. The target is the machine sitting under a researcher’s desk at a lab that will not put its data in someone else’s building — the same slot Nvidia’s DGX Station occupies.

The specification that matters is the memory, not the cores. Frontier-scale inference is a memory-bandwidth problem before it is a compute problem, and 288GB of HBM3E across two cards — 576GB across four — is roughly the amount of fast memory it takes to hold a very large model resident without spilling to system RAM. AMD is explicit about the comparison: HBM3E at 4TB/s is on the order of fourteen times the bandwidth of the LPDDR5X used in the unified-memory desktops that have been marketed for local AI over the past two years. Those machines fit big models. They do not run them quickly.

Why a desk and not a rack

The economics of a $150,000 workstation only work if cloud inference is unavailable rather than merely expensive. For most teams it is available and cheap, and getting cheaper — Microsoft just priced transcription at ten cents an hour. The buyers here are the ones for whom the API is the problem: defence contractors, hospital systems, banks, national labs, and any organisation whose legal department has decided that inference on regulated data happens on premises or not at all.

That is a real and growing market, and it is one AMD has been winning in quietly. Europe’s €388m LUMI-AI order contains no Nvidia silicon at all. Sovereignty buyers care less about ecosystem maturity than about supply chains they can audit, and AMD has spent three years making that argument.

Our take: The interesting number is not $150,000, it is 576GB. AMD has quietly established that the ceiling on local inference is no longer “a 70-billion-parameter model, slowly.” When the deskside ceiling moves, the set of workloads that were only ever going to run in a hyperscaler’s building gets smaller. Not by much this year. But the direction has reversed.

What to watch

The pitch AMD is making is not that this box is cheaper than the cloud. It is that for a certain kind of buyer, the cloud was never an option, and until now the on-premises alternative meant a rack, a room and a procurement cycle. Berlin’s answer is a tower and a wall socket.

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.