Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!
I've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done.
I'm definitely having full eval suite on already if they get overloaded later on.
Self-replying as this is no longer correct - there are two tiers, one matching Moonshot pricing with comparable latency/throughput, and a fast mode at $4.50/M (still cheaper than Opus) with 3x the throughput.
Comparing to Opus 5: Claude Opus 5 (Uncached Input $5/M Cached Input $0.50/M Output $25/M) but you also pay a premium on Cache write 25% for 5m and 100% for 1h.
I have to say cc opus 5 is abysmal. It talks to itself incessantly, gets stuck in minutia, fails to understand problems clearly and makes steering mistakes constantly. It also has a weird behavior where it says “ok I know exactly what to do and I will start now,” then sits waiting for user input. If you’re not on the ball you’re constantly losing 5m/1h cache. Just give me back 4.6.
You can even use /model to use Opus 4.5 if you want. Broadly though I find Opus 5 and most of their point releases (4.7 being the exception) to be excellent and upgrades.
I love fireworks.ai! They launched it couple of hours ago and we have it now already live on our platform for our users. Just a shame they deprecated the on-demand flux models :( Where do I get my fix for image gen now?