Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Imagine this, and sucesor models on Cerebras or other silicon...


That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.

If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.


Very exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s

https://news.ycombinator.com/item?id=49308715




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: