Hacker Newsnew | past | comments | ask | show | jobs | submit | cbg0's commentslogin

Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)

I use AI on a daily basis and I still found all of your takes in this thread to be silly hyperbole.

You'll be surprised (or not) to find out a lot of people aren't using anything. Claude gets asked to build it, then some other clown uses their copy of Claude to do the PR review and then it gets merged.

> regardless of what you think about coding is done, agents and your harness is the best tool.

I guess you're all in on that Anthropic IPO with the "coding is solved" approach.


There will always be people that write code with pen and paper. Its going to be an art form though rather than something productive.

There are AI players already inside Follower Dungeons. (not LLM-powered though)

Assuming you are right, the "build-up" of more physical datacenters isn't really the difficult part, training frontier models is. The open-weight Chinese models are several months behind frontier models at this stage and as long as you're not interested in burning cash to stay in the lead you can just wait it out. Many signs (upcoming ridiculous IPO figures and benchmarks) point to the current LLM paradigm reaching a plateau.

I think that currently the EU is focusing on strengthening itself from a military perspective due to US policy shifts, so the EU just doesn't have the funds to divert to AI without impacting day to day life of people, thus prioritizing safety and perhaps energy production, as electricity is considerably more expensive here than in the US or China. If you ask the people who vote, they'll put safety, grocery and gas prices way above EU's LLM capabilities.


> so the EU just doesn't have the funds to divert to AI without impacting day to day life of people

The US Government is not a major funder of frontier labs (who are almost exclusively VC and private money funded)...

> as electricity is considerably more expensive here than in the US or China

Yes, there are government policy reasons behind why this is the case. Some would argue that these policy decisions are a net negative for the EU.

> If you ask the people who vote, they'll put safety, grocery and gas prices way above EU's LLM capabilities.

Fair point, these are good things to focus on. However, EU citizens are going to love it when there are no frontier-tech opportunities left in their countries, with everything in this area having to be imported from the US and China (note that just because Chinese firms make open weight models available now, this could easily change very quickly in the future for a variety of reasons).


> just because Chinese firms make open weight models available now, this could easily change very quickly in the future for a variety of reasons

China is not providing "open weight" models just because they think it's good for humanity. It is a competitive strategy to stay relevant at the forefront of the industry. The reality is that China is severely compute-constrained mostly due to Nvidia chip and ASML embargo. Nearly all available compute capacity has to go toward training, leaving very little for inference. It is why they release their models openly allowing third parties in EU and USA to host inference for them.

This will change when China reaches a critical scale of domestic AI chip production, which is inevitable.


Yer this is one such 'reason' for China stopping the release of open weight models in the future, there are many others...

> The US Government is not a major funder of frontier labs (who are almost exclusively VC and private money funded)...

There isn't a lot of VC funding available in Europe compared to the US, so without government intervention catching up will be very tricky.


The problem is that EU has no money. France is virtually broke. Germany and eastern/central Europe is hemorrhaging its most valuable industries to cheaper Chinese competition. The whole bloc is carving out $100 billion military aid packages for Ukraine basically every year with no end in sight. And now EU is getting into a trade war with China despite having little leverage. Energy costs were bad even before the Ukraine war and Iran, and now they're at the boiling point.

It's why the USA and Chinese vultures are swarming. It's a bad situation.


France isn't "broke", its debt-to-GDP ratio is actually lower than the US right now and it's not in a situation where it can't pay its loans.

There is currently no real trade war EU-China; Chinese EVs are streaming into the EU while the US is blocking them completely for one example.

While Germany and more specifically Volkswagen is having some issues, EU industrial production actually had growth last year.

The $100 billion aid packages to Ukraine are also a bit overblown based on public figures https://www.kielinstitut.de/media/news/ukraine-support-track....

I feel like your comment pulls on some real threads, then blows things out of proportion.


Yes, I'm sure that EU doing its best to ignore building any technology frontier capabilities in its population will lead to a good solution for this lack of VC funding... ;-)

No shortage of visionaries in EU leadership these days :-)


You've posted a link that doesn't support your statement.

If you click through to [1] that seems like a clear downwards trend (beyond the usual noise) about two weeks before the release of Opus 4.7, Opus 4.8, and Opus 5.5. Opus 5 is the only launch that looks clean without the previous model being nerfed beforehand

https://marginlab.ai/trackers/claude-code-historical-perform...


From that site:

> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).

Changing the harness can have a big impact on performance even when leaving the model completely unchanged.


Sure, maybe it isn't the model getting nerved but the harness getting updates that make it better with the new model but substantially worse with the old (at that point still current) model.

The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release


Yes, but then the model wasn't nerfed, the harness/overall product just had a plain old regression.

This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.


This has always been Reddit, aside from the niche subreddits.

I don't think so. I've been a Reddit user for 12+ years or so. Obviously when you make the platform easy to use for anyone, you also invite more casual users who just want to meme.

Reddit doesn't even make any attempt to control quality since user engagement numbers matter more than quality.

Reminds me of Quora which over optimized the site for engagement growth but eventually became completely useless for this reason.



The big default subreddits were garbage 12 years ago, yes. But the rest of the site was not.

Now in 2026, even small niche subreddits are facebook / tiktok tier with low effort meme posts etc


> even small niche subreddits are facebook / tiktok tier with low effort meme posts

Completely agree.

There's a few niche subs I hang out in but, for me, the standout in terms of quality decline is r/DIYUK. The difference between that sub in 2019/2020 vs today is stark.

It's pretty common to scroll through several pages of repetitive one-line "jokes" before getting to anything remotely resembling a real answer.

Mods seem relatively inactive and there are only 4 of them anyway, which is nowhere near enough to keep on top of the volume of posts and comments nowadays. Really it needs either a lot more mods, or some form of LLM-backed automation, to weed out the low effort comments.

It's become a bit of a wasteland in terms of a lot of the regular contributors who used to give good, well thought out answers, have tapped out. I complained about this the other week and got downvoted and accussed of "showboating" which, honestly, I'm not even sure what the commenter who said that meant in the context of the discussion that we were having.

It's a real shame because there aren't really any good UK-focussed DIY-oriented discussion forums any more: all the others degraded and became toxic, which is where r/DIYUK won, because it resisted that trend for longer. No more, sadly. DIYnot always stood out as a place that was more about battling egos than providing meaningful help and advice and, although r/DIYUK hasn't gone in quite that direction, it is sadly headed down an equally useless and destructive path quality-wise.


The subreddits that I frequent have turned into a low effort meme fest.

No original OP posts are allowed, only URLs from the same old websites.


Yes, and the person you responded to posited that this is because they used to be niche. You claimed this was not the case, they provided some URLs that demonstrate semi popular subreddits have always been shit.

This post sounds like you are still disagreeing but nothing you've said points to another conclusion


Your reddit experience is from mods and their rules and how they enforce it, not the site overall.

As long as mods are human volunteers, can they ever keep up with scale of reddit? I regularly see submissions (on old.reddit's /r/all) that got to 10k+ upvotes and 200+ comments in hour or two before it was taken down for breaking given subredit rules.

Yeah, look at AskHistorians. Its still pretty awesome.

Every interesting question just shows a bunch of deleted comments, so I stopped looking at it a while ago.

There is another sub which tracks answered questions. Its a side effect when focussing on quality.

Not really, if you want to ask a question. 9/10 times it gets locked for breaking some obscure rule. If anyone answers, that answer is just deleted immediately as well.

This type of subreddit specifically is almost entirely useless since the advent of AI.


It depends. I have compard some answers with my discussions with different llms and i still find the answers in the sub qualitatively better.

Site overall. Mobile app makes it easy to type a few words and hit send and hard to write well thought out posts.

When every mod does the same it becomes the whole website.

The site is the sum total of all that

I've been on Reddit since late 2005. Before there were comments or subreddits.

My memory is fuzzy, but I remember it getting pretty shallow and memetic since before the Digg diaspora in 2010. And I think it was the popularity of image posts that did it; it was better before image links became popular.

However, Slashdot was already highly memetic since the late nineties, and it was not image-based, so I may be wrong about the effect of pictures.


> Reddit doesn't even make any attempt to control quality

But mods can say they censored you because of low quality contribution to reddit. :)

I do agree that the karma system does not work. I was able to farm 77k karma in two years, after I created a new account there (the old account I had before for +14 years was forced to change the password, so I created a new account - and used EXACTLY the same password, to prove a point to reddit that they are incredibly stupid. Sadly, reddit never understood the problem, even though it should be obvious that, if your passwords are lost yet users retain the same password after being forced by reddit to change it, something in reddit's insinuation simply does not add up here.)


I once cracked a pretty good joke on a high-traffic thread and got about 30k upvotes for that one post. None of my serious posts in serious subreddits ever got close. I don't think Reddit karma is a useful indicator of anything.

you weren't around back in the days that it was just a Ron Paul fanboy forum and a lot of r/atheism echo chambers (I say this as an atheist who got keenly annoyed by the amount of self-fellating going on there) - this was pre-sub-reddits, when kn0thing was still active on the site and a force for some pro-social moderation, and Digg was still their biggest competitor (until it was flooded with the HD-DVD key leak and a lot of people left it out of annoyance)

This got me wondering whether you could train an automod to act like Dang.

Dang is more or less what keeps HN from becoming Reddit.


Don't think so unless all moderate actions are public.

I don't know about that, political slop is sometimes manually unflagged here.

People would find exploits in the automod. For example if a trivial keyword filter flags comments mentioning Palestine but not ones that mention Paluhstein, or "the country of Isn'treal"

This is called Algospeak and has infected a lot of people's everyday language. People say "unalive" instead of "kill" or "murder", because the internet has drilled into them that you have to use that word.


An LLM would easily understand that, no?

If trained enough, but the commenters will just start calling it watermelon instead. So you train your LLM to block comments about watermelons, so they say thing. So you train your LLM to block comments about things. Congrats you blocked all comments.

Once-upon-a-time reddit was the cool new alternative to digg

> This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

It's been out for an hour and you've already concluded this?


You don't have to run Claude/Codex in auto-approve mode, you can manually approve its interactions with your machine without having to copy your code back and forth between the website and your local files.

Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: