On August 25, Apple announced two new chips: the M6, its first built on a 2-nanometre process, and the M5 Ultra, the first M-series part assembled from four dies. The M6 goes into a new Mac mini. The M5 Ultra goes into a new Mac Studio, and its top configuration is a serious machine by any measure: up to 36 CPU cores, an 80-core GPU with Neural Accelerators, a 32-core Neural Engine, and up to 512GB of unified memory read at 1.2TB/s.
That last number is the one worth sitting with. Apple quotes it as 50 percent more than the M3 Ultra, and it is the highest memory bandwidth ever offered in a Mac. It is also about 28 percent more than the 936GB/s that Nvidia's GeForce RTX 3090 has delivered since September 2020 — a consumer graphics card that launched at $1,499 and now trades used for well under a thousand dollars.
Why bandwidth is the number that matters
For running large language models locally, memory bandwidth is usually the wall. Generating a token means reading essentially every active weight of the model out of memory, once per token. Compute helps, but for a model that fits in memory, tokens per second track bandwidth almost linearly. A machine with 1.2TB/s and a card with 936GB/s sit in the same performance class for this workload.
That makes the timeline remarkable. The 3090 set its mark six years ago as a gaming part. It took a 2-nanometre era, a quad-die package and the largest unified memory system Apple has built to move clearly past it. Nearly everything else in computing multiplied several times over in those six years; the cost of moving a byte out of memory barely moved at all.
The money and the power
Apple has not published pricing for the new Mac Studio yet — pre-orders open September 22, and the 512GB configuration arrives in late October. Its predecessor gives a reference point: a Mac Studio with the M3 Ultra configured to 512GB listed at $9,499. If the new flagship lands anywhere near that, the machine costs roughly ten times what a used 3090 does, for about a quarter more bandwidth.
Power is the counterweight, and here the new silicon genuinely earns its keep. A 3090 budgets 350 watts for the card alone, and Nvidia recommended a 750-watt supply for the system carrying it. Apple does not publish chip TDPs, but a whole Mac Studio at full load draws less than that single-card budget, and idle and per-token figures favour it further. If a machine generates tokens all day, the energy line item is real.
Capacity is the other honest advantage. The 3090's 24GB caps what fits without splitting a model across cards or quantising hard. 512GB of unified memory runs models with hundreds of billions of parameters on one desk — Apple's announcement says exactly that, and it is the configuration's whole reason to exist.
What this says about hardware lifecycles
We refurbish and resell enterprise hardware, so we notice when the market proves our point. The number that governs this decade's defining workload was reached by a consumer part in 2020, and that part is still competitive on it — not nostalgic, competitive. Per gigabyte-per-second, there is no cheaper way to buy memory bandwidth than the used market.
The lesson carries past GPUs. Servers, switches, storage: depreciation schedules assume capability decays on a calendar, and it does not. Hardware leaves the field because budgets and fashions move, usually years before the physics does. That gap is where our business lives — and, for buyers, where the value is.
If you are building local AI capacity on a budget, talk to us about tested, warrantied GPU-capable systems — or start with what the used market already offers.
Image: Apple.
Highlights
- M5 Ultra top configuration: 36-core CPU, 80-core GPU, 32-core Neural Engine, 512GB unified memory at 1.2TB/s
- GeForce RTX 3090 (September 2020): 24GB GDDR6X at 936GB/s, 350W board power, $1,499 at launch
- Six years on, the bandwidth gap between them is about 28 percent
- Mac Studio pre-orders open September 22; the 512GB configuration follows in late October