Hacker Newsnew | past | comments | ask | show | jobs | submit | Tactical45's commentslogin

The fact that it doesn't work well today doesn't invalidate the user need. Not everyone wants a terminal style interface. Traditional UX makes more sense especially if say you are accessing existing plugins and say kicking off a scheduling task - no way would I want to do that via a chat interface if I have the option.


Agree, and it allows ordinary users to have a taste of both worlds, instead of being like: "eww, what is this black box and why doesn't my mouse work"


I think the notion that it's the worst piece of software is a bit of a stretch here. I've used various streaming music players and Spotify has been the best by far. Doesn't mean it's perfect.

I get that not all edge cases (eg swipe gestures) are bullet proof but that doesn't make it terrible software.


It's objectively really bad at a lot things. I've always had the feeling that the music is not even a primary feature of their product.

The last time I used it, I uninstalled once I realized that my kids could watch Logan Paul YouTube videos through the Spotify app. These masquerade as podcasts, but like a Trojan horse they are just vods of YouTube content. And at the time there was literally no way to disable this. No way to either turn off podcasts or specific content through parental controls or otherwise.

I just wanted to give my kids a way to listen to music while reading. I have since switched to our local community radio station.


WinAmp was smaller, faster, and more customizable thirty years ago.


This response is not relevant to the this comment


At what cost difference?


I don't think it matters, if it's for local/on-device usage.

The cost is similar vram footprint I guess (?)


loading the model would be similar vram footprint, correct, however the size of KV is based on 'active' params, not total params. So while at 1k ctx both will be in the same ballpark vram footprint-wise, at 100k the story will be very different. 27B at q4 kv for 256k takes ~8gb, while 35B at q4 kv around ~3.5gb, so at full precision kv those would be ~32gb and ~14gb (all ballparks, if you want exact numbers, its not hard to test).

As for the "cost", here i think the interesting arguments are around speed vs accuracy/"getting the job done", not literal $ cost per token.


I am asking mostly for running on a 3090.

I think the tps difference between them (both fitting in vram) won't be more than 2x in practice.

I would happily take 20tps over 40tps, if the model gets 3x more correct answers.


Speed is (for the most part) active-parameter based, so a 30B-A3B model is roughly 10x the speed of a dense 30B (realistically closer to 8x) in the case when both fit. That's the proposition of MoE and why everyone is trying to make massive models with very few active params, so that they are still fast while having access to a lot of knowledge (at the cost of reasoning, as reasoning ability 'for the most part' comes from active params).

You can test this by running this nemo or 35B on the 3090. I have and its very fun (but sadly a worse model than 27B, so I usually keep 27B on my 3090)


I remember running both qwen 30b-a3b and 27b on my 3090, and on the initial test, the 27b was only like 2x slower.


Ran a quick test so that we both have accurate numbers, without MTP* at 10k ctx 27B hovers around 42 ts in llama.cpp, 35B around 135 ts. So not the 8x I assumed, just over 3x, but thats still a big difference.

For the sake of testing I turned MTP off, as that heavily depends on what the generated text is (structured text like code is very often a lot more predictable, therefore bigger boosts) and the quality of the quantization, as drafters learn how to "mimic" the full precision generated tokens, so when you layer the fact that MTP is a 'guesser' of the main model's next token, and quantization affecting what exact token is generated, it'd make comparisons like this needlessly noisy.


Thanks for sharing.

Yeah, that was my experience too (2x or 3x is indeed considerably faster), but not workflow-changing faster at 45tps baseline, especially for asynchronous tasks (which is my goal with a local 3090, to just let it do things non-stop, without my intervention).

What was the result with MTP?

Isn't MTP "losless"? The result should still be relevant when averaged across a fee queriers across different domaine I guess.


MTP is lossless in the sense that running a model with and without (at temp=0, meaning no randomness) will produce identical results. It's true that with enough samples across domains and runs with MTP it should even out around concrete numbers, but I don't have time currently for long tests. On a quick test (before I remembered MTP is on), 27B was around 60-70 ts and 35B around 180-200 ts, both going up and down but mostly in those ballparks, which is inline with the ~3x from not using MTP.

One somewhat related thing is that, without drafters (the models doing just generation) ts tends to slowly go down, while with drafters, as the model is "checking drafts" instead of pure generation, even when the avg ts has fallen to say 50 ts (27B, ~128k ctc for example), there still tend to be random spikes to up to 70 (mostly during coding), which is always funny to me but also makes models a bit more "resilient" to the slow degradation of ts.


Don't be dramatic. Gemini is fine. I have the Google voice setup all throughout my house and haven't had any issues major enough that would want me to keep the original assistant.


It's not dramatic, it's just something you personally haven't noticed. A quick search regarding Gemini weather hallucination shows it's both widespread and something not mitigated by recent model releases. It makes up other basic information or can't perform tasks that a simple assistant can do properly (e.g. mixing up calendar events, emails, can't navigate etc). Why defend what is clearly a move to increase a department's KPIs when you as a consumer receive less in return?


Besides mispronouncing words most days, it literally said it was snowing and Florida-hot in one sentence this week. It's the first time it's been that wrong though.


Gemini has no idea what to do when I tell it to 'Navigate home'. Or 'Next song'.

It's objectively worse at every basic command.


Gemini defaults to performing search from within its own weights until you curse it and damn it sufficiently to do basic RAG. Utter trash on this limitation/behavior alone but there are so many other hallucinatory shenanigans it embraces so why limit yourself to one.

Compare and contrast with asking Claude or ChatGPT a question directly and get back to me.


I never thought I'd see a Gemini apologist.


You've hit the nail on the head, it's a conditional defenition, rather than an absolute one.


OpenAI or Anthropic are not anywhere as well funded as Google. Apple already has everyone in their pocket via the ecosystem, they just have to not crap the bed. They value stability over the competitive component here.


More like 7-8 digit salary, in the case of the CEO's direct reports at Uber


And speed light cameras


What is a speed light camera?


I'm working on Travi, an AI-powered travel companion that helps travelers effortlessly discover the best attractions to check out based on their interests, and experience them through rich, immersive audio narratives.

You can check it out at https://TryTravi.com


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: