Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model would generate tokens at ~200 tokens/sec.
Macs don’t have very good prefill, though. So it’s important to use a model serving stack that has excellent prompt caching and use a harness that won’t bust the cache.
I’m cross-shopping DGX Sparks and M5 Studios, and having a hard time deciding because they have exactly opposite characteristics for prefill and decode.
Macs don’t have very good prefill, though. So it’s important to use a model serving stack that has excellent prompt caching and use a harness that won’t bust the cache.
I’m cross-shopping DGX Sparks and M5 Studios, and having a hard time deciding because they have exactly opposite characteristics for prefill and decode.