Hacker Newsnew | past | comments | ask | show | jobs | submit | ckocagil's commentslogin

I'd rather go the other way and entirely decouple my online identity from Google (or Apple/Meta/Microsoft). Much better than constantly worrying about some crappy AI deciding to ruin your life with no recourse.


Are you talking about router firmware or NIC drivers? There's a slim chance of vendors opening drivers, but firmware is never going to happen.


>The stated goals of Wi-Fi 8 include a 25% increase in throughput at different signal-to-interference-and-noise ratio (SINR) levels, reduce latency by 25% for the 95th percentile scenarios with latency

So it is getting faster (in both senses) and doing so in scenarios that actually matter.


Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.


Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.


The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.


Isn't Q8 way overkill these days? I see many graphs showing Q4 or Q5 having less than %1 deviation. Nvidia's NVFP4 Qwen quantization should be even better due to its better training methods.


Careful with those graphs, they're usually evaluating the model on KLD on relatively short transcripts. When you're running with 100k token contexts and the model running close loop a difference that looks small in terms of KLD may be quite substantial.

I'm not aware of any great benchmarks that work by giving it a live agentic harness and a number of realistic tasks that take most of the context window to accomplish and evaluate success rate and tokens to completion... but that's what you'd really want to use to judge different quantization levels.


Q8 isn't overkill if you have sufficient RAM to fit the whole model, and you care about quality. There's a number of people who have enough hardware to fit exactly one 27B to 35B size Q8 model and not more than that, so if you can fit the whole thing in Q8, no reason to use Q4 or Q6.


When orgs/bencmarks claim 1% deviation, in most cases that means measuring perplexity loss on datasets like wikitext or c4. Even if the loss is calculated via KLD or similar, its not a good proxy for whats actually degradaing at the task level across an entire rollout.

And for MoEs, very small amounts of loss can mean you're flipped to entirely different experts (this is also a problem more broadly with numerical stability issues too).


Qwen3.6 below Q8 often can't exit a reasoning loop (until it hits max output token count), forgets to insert a tool call, often mistakenly inserts them inside the thinking block... It's still usable though.


1.01 over 30k tokens is over a googol (a large number with 100 zeroes)


It depends on model size I think, but yeah, from my understanding at ~30B and below Q6 or even Q4 will get you 95%+ of the way there


aka a stroboscopic measurement,

but I don't think it will work well for this case.


It's just higher nyquist zones.


Late by a decade or more (JSR310 was released in 2014), but still a good development. I've tried convincing colleagues to use js-joda in the past but they thought they were keeping it simple by sticking to moment.js. They weren't.


I remember running into the moment.js issue where my package size doubled by adding a date library: https://github.com/moment/moment/issues/3376


Isn't the US stock market betting on AGI and superintelligence in the next decade? Maybe a lower population won't be as big of an issue, or even an advantage.


That's a neat software solution. My first inclination would be to grab a soldering iron and replace the crystal with either a TCXO or a socket to provide an external clock disciplined to the 1PPS.


That's also the most tiresome part of driving and has the least risk due to low speeds. Easy win for FSD. But for all other cases it becomes a complicated ethical question.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: