Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model would generate tokens at ~200 tokens/sec.

Macs don’t have very good prefill, though. So it’s important to use a model serving stack that has excellent prompt caching and use a harness that won’t bust the cache.

I’m cross-shopping DGX Sparks and M5 Studios, and having a hard time deciding because they have exactly opposite characteristics for prefill and decode.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: