Tech & Building

Tech & Building

Follow-up: I’m still buying the Mac Studio M5 Ultra 512GB. AMD just made the rest of the market look very interesting.

AMD’s Threadripper Halo Station and NVIDIA’s DGX Station just showed up next to the Mac Studio M5 Ultra decision. A side-by-side on memory, bandwidth, power, and price for anyone weighing a personal supercomputer this year.

I wrote last week that I’m buying the Mac Studio with the M5 Ultra and 512GB of unified memory.

That decision has not changed.

What has changed is the neighborhood. On September 4 at IFA 2026, AMD wheeled out the Threadripper Halo Station: a liquid-cooled deskside tower with a 96-core Threadripper PRO, a path to 576GB of HBM3e, and enough combined memory that AMD is talking about trillion-parameter models living next to your keyboard. NVIDIA already sells the DGX Station in the same price range.

Serious local compute is no longer one product category. It is three different ideas of what a “personal supercomputer” should be.

What I committed to

The M5 Ultra Mac Studio is the machine I can actually live with.

  • Up to 36 CPU cores and an 80-core GPU with Neural Accelerators in every GPU core
  • Up to 512GB of unified memory
  • 1.2 TB/s of memory bandwidth
  • A 7.7 × 7.7 × 3.7-inch aluminum block that idles in the single-digit watts and is rated at 480W maximum continuous power

It will not be cheap. Apple has not published the 512GB price yet. Going from 96GB to 256GB already costs a lot. The 512GB machine will cost like a cheap car. I am buying it anyway, because it arrives this autumn, it is quiet, and I can use it for everything else I already do on a Mac.

Then AMD dropped a server on a pedestal

The Threadripper Halo Station is the opposite personality.

AMD showed a prototype at IFA and says systems come in 2027. The pitch is blunt: “Power frontier-class models. Run hundreds of agents. Locally.”

The configuration they keep quoting:

Spec

Mac Studio M5 Ultra (max)

Threadripper Halo Station

NVIDIA DGX Station (GB300)

Form

3.7" desk cube

Liquid-cooled tower

Liquid-cooled tower

CPU

36-core Apple silicon

96-core / 192-thread Zen 5 Threadripper PRO 9995WX

72-core Grace (Arm)

Accelerator memory

512GB unified

288GB HBM3e now, 576GB with 4× MI350P

252GB HBM3e

System memory

same pool

up to 2TB DDR5

496GB LPDDR5X

Combined memory

512GB

up to 2.6TB

748GB coherent

Memory bandwidth

1.2 TB/s

up to 16.4 TB/s aggregate (HBM + RDIMM)

7.1 TB/s HBM + 396 GB/s CPU + 900 GB/s C2C

Power envelope

480W max continuous

~1,550W of silicon in the 2-GPU show config; more with four cards

1,600W system

Availability

Oct 2026 (512GB)

2027

shipping now via OEMs

Street talk

~$20k+ for 512GB

$100k–$150k estimates

~$93k–$125k

Each MI350P is a 144GB HBM3e CDNA 4 card at up to 4 TB/s and up to 600W. Two of them give you 288GB of accelerator memory. Four give you 576GB. Add 2TB of DDR5 on the Threadripper and AMD’s slide claims up to 3.4× the total system memory of a DGX Station.

That is a different class of problem than “run a 400B model on my desk.” That is “keep a frontier-class model resident, then run a swarm of agents against it for hours without a cloud invoice.” AMD said the quiet part out loud: this exists because inference is no longer a single prompt. It is continuously running agents, and those agents are hungry for memory and bandwidth that ordinary workstation GPUs do not have.

It is also a different class of object. The show unit is a glass-sided tower full of tubing. The two-GPU silicon budget alone is about 1,550 watts before storage, fans, and conversion losses. Four cards and you are in multi-kilowatt territory. In a Swedish apartment that is not a “put it under the monitor” decision. That is a circuit, a noise, and a heat decision.

Three machines, three religions

The interesting part is not who “wins.” It is that the market finally split along actual use.

Apple is betting that unified memory plus efficiency plus an OS people already work in is enough. You buy one box. You do not manage CUDA versions. You do not hear it. You do not rewire the room. You accept that you will quantize, that MLX/Metal is not the entire research stack, and that 512GB is the wall.

NVIDIA is betting that the software gravity well still matters more than the case. DGX Station gives you 748GB of coherent Grace + Blackwell memory, 20 petaFLOPS of FP4, ConnectX-8 networking, and the stack every lab already knows. OEMs are selling it now in the $90k–$125k range. It is a deskside node, not a Mac.

AMD is betting that memory capacity is the new clock speed. 96 Zen 5 cores, eight-channel DDR5, 128 PCIe 5.0 lanes, and a path to 576GB of HBM3e is how you keep a trillion-parameter model and the agent runtime on one machine without renting a rack. ROCm still has to be good enough that researchers will trust it. 2027 is a long time in this market. But the prototype existing at all is the signal.

I keep coming back to power and presence. The Studio is a 3.6 kg object that looks like a high-end DAC. The Halo Station looks like a small server that learned to stand up. One of those belongs next to a display in a flat. The other belongs in a room you have already decided is a lab.

Why this is the interesting year

Two years ago, “serious local compute” meant a 4090 and a prayer, or a cloud bill. Last year it meant 128–192GB unified-memory boxes and the first DGX-class deskside systems. This autumn it means:

  • A Mac you can order that holds a 400B-class open model in one memory pool
  • An NVIDIA tower you can order that claims a trillion parameters and 20 PFLOPS
  • An AMD tower, a year out, that tries to beat both on raw resident memory

The cloud does not disappear. Training still lives there. But the argument that you have to send weights and private data off-box for anything larger than a 70B model is getting weaker every quarter. That is the shift. Privacy, latency, agent uptime, and the simple refusal to meter your own thinking in tokens per dollar.

I am not pretending the Studio is a Halo Station. It isn’t. I cannot fine-tune a frontier model on it the way a lab will on four MI350Ps. I cannot match 16 TB/s of HBM. What I can do, in October, in a quiet room, is run the models I care about, keep the weights at home, and still edit, compile, and live on the same machine.

That is why the order stands.

The Halo Station is the reminder that this category is not finished. Apple made local large-model inference feel normal. NVIDIA made it feel like a product line. AMD just made it feel like a race.

If you have the budget, the circuit, and a 2027 timeline, the AMD box is the most aggressive memory story on a desk. If you need the NVIDIA stack tomorrow, DGX Station is already a SKU. If you want the machine that will actually sit under your monitor this year and not sound like a heat pump, I’m still buying the Mac Studio.

The market for serious compute didn’t just get faster. It got plural. That is the more important news.