Hacker News front pageBy Federico Viticci39 min readintermediate
M5 Ultra Mac Studio Review
Summary
The M5 Ultra Mac Studio (256 GB RAM) uses a quad‑die M5 Max architecture with an 80‑core GPU and 1.2 TB/s memory bandwidth, delivering ~70 % faster prompt‑to‑first‑token and token‑generation rates than the M3 Ultra. In the author’s tests Qwen3.8‑Flash‑Next hits 100 tokens/s on short prompts and 60‑85 tokens/s with 64‑256 KB context, making local AI agents (Open Minis, Hermes, Codex) feel snappy a…
- Quad‑die M5 Max + UltraFusion gives 4.5× AI GPU compute vs M3 Ultra.
- Memory bandwidth up from 819 GB/s to 1.2 TB/s (≈50 % increase).
- Benchmark: ~70 % faster overall response time vs M3 Ultra; 100 tokens/s on short prompts, 60‑85 tokens/s with large context.
- Enables always‑on local agents for research, note‑taking, and code assistance without cloud fees.
Local AI workloads are becoming practical on consumer hardware, and the M5 Ultra shows that a Mac can replace expensive GPU rigs for continuous, privacy‑preserving assistant tasks. The performance jump directly translates to smoother multi‑turn interactions and larger context windows, which are cri…
6/10



