- 89 tokens/sec generation with speculative decoding
- 81 tokens/sec at 20k context
- 474 tokens/sec ingestion at 20k — about 42 seconds
- 10.1 GiB peak VRAM with a 24k context window
ROCm 7.2.3 · PQ2_0 · Qwen Q4 MTP, draft length 2
---- versus ----
Qwen3.8-27B IQ3_S · Radeon RX 7900 XTX
- 79 tokens/sec generation on a short coding prompt
- 61 tokens/sec at 60k context
- 53 tokens/sec at 95k context
- 558 tokens/sec ingestion at 60k — about 108 seconds
- 19.9 GiB peak VRAM during coding tests with a 100k context window
what kinda speeds do you see on 6700 XT? i'm always conflicted on investing time chasing speed-vs-quality tradeoffs. I've got a 7900 XT (about double the IO throughput). I'll probably end up giving it a go when I find time.
don't see it in my AWS bedrock model list yet, but boy has bedrock mantle been annoying today with the errors/downtimes, with NO status page entries >_<
I think I have to look at setting up one of these AI gateways for work, since AWS bedrock is such a PITA to hook a harness up to over IAM roles. Also I still can't believe bedrock hasn't released any open models in months (so there's paranoia that I'll want to swap in another provider).
Really though I'm hoping anthropic fixes the oppressive claude verbosity. I saw someone refer to being "clauderboarded" and my brain cannot let go of this as Claude's tokens bombard me.
We had this exact problem so we solved it for ourselves. Happy to help if you run into any issues.
I have a /hmmm command I use for the second part that works reasonably well: "Stop using jargon and speak coherently. State it more simply and concisely, like one human talking to another. Make it like google dev docs style. More dead prose. No aphorisms, no flourishes. Simple."
Oh. The UI screenshot on github is... not actually in the github repo? It's platform/hosted only? There's my first awkward discovery, but makes sense in retrospect.
I had to go down to UD-Q3_K_XL for Qwen 3.8 27B to get it to fit in VRAM and be usable, but I worry I'm gutting its intelligence somewhat. I too am interested in faster + more-usable alternative that can exchange blows with the Q3-dumbed 27B.
I'll be comparing the 9B vs Ling 3 Tiny (8B-A1B) as a scout model. Ling tiny is so fast but can be a little too dumb. Hope the 9B strikes a good middleground even if dense/slower.
ended up disabling ornith 9B. Oddly Ling 3 Tiny is pretty dang capable if its thinking is unleashed (tons of output tokens, maybe 5X the tokens but it's so fast it's maybe only twice as slow as a smarter model). This is a really interesting space if these super small models keep improving. SO FAST! Want to use!
I've been waiting on this model to show up on the Deep SWE benchmark results and treat its absence/delay as an indication of how slow and unusable it is for good results. I bet it thinks to the moon on some of those complex challenges.
---- versus ----
Qwen3.8-27B IQ3_S · Radeon RX 7900 XTX
Vulkan · GSQ-RCO IQ3_S · MTP, draft length 2 · vision projector loadedreply