Hacker Newsnew | past | comments | ask | show | jobs | submit | AnodicElegy's commentslogin

"OpenAI’s primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning."

Yes, and it would greatly aid alignment if the user could monitor the chain-of-thought! The open models, including very powerful ones like Kimi K3, are delivering this. The fact that OpenAI and Anthropic are not clearly indicates that a commercial consideration (avoiding distillation) takes precedence over alignment -- regardless of how much they bloviate about the latter.


If someone uses an LLM to write and is able to tailor their writing such that it isn't obviously written by an LLM, then I'm fine with it! But two of my otherwise-favourite news sources -- the Hacker News front page and FT Alphaville -- are inundated by articles where the LLM usage is blindingly obvious.

This update really gives OpenAI a boost. Not saying there's anything inaccurate or untoward about that, but the timing is unfortunate. It would have looked better had it been done prior to the Fable 5.1 and GPT 6 releases. I guess AA would say that there's no perfect time to do these updates, given the rapid fire pace of releases!

The timing is related to the fact that their benchmark was saying it was the same as Sol, and below Opus 5, when anecdotal reports and other benchmarks strongly disagree. It looked bad for them for their benchmark to disagree with people's lived experience so hard.

Kind of reminds me when GPU benchmarks used to game the drivers to maximize the FPS. If the benchmark slightly changes the camera view is that cheating? Or is it calling out the cheaters?

It's very different than it was earlier today, so pretty sure it's v4.2.

Any angler will know that, in some lakes, it seems like every fish is full of worms. I've always dealt with that by frying the crap out of them -- no medium-rare bass for me, sorry. Alternatively, I'm sure they're fine after a few weeks in the freezer.

Rivers seem to be better, with worm presence (or at least, obvious worm presence) being less common. Additionally, I've never seen worms in bottom-feeding fish like catfish and carp.


>Additionally, I've never seen worms in bottom-feeding fish like catfish and carp.

no just the heavy metals you can see...


Everything in moderation! With the big predators you get mercury, with the bottom-feeders you get persistent lipophilic compounds like PCBs and PFAS.

I could just fish for perch, but what's the fun in that?

(I'm mostly kidding, I love fishing for perch too.)


If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.

This is pure speculation, and he really doesn't know what he's talking about:

"The main problem is that DXO also contains an amine, and that amine can react during the process. So the first step would be to temporarily protect the amine so that it doesn’t interfere."

Anyone who's taken introductory org chem knows better than this. That's a tertiary amine, so you're not going to be protecting it. Not that the compound can't be made -- it probably has been already. I'd check on Reaxys but I don't really want a morphinan in my search history...


Confirmed: the compound proposed by the author has been made and studied since at least 1992 (https://pubs.acs.org/jmcmar/article-abstract/35/22/4135/7112...). It had activity interesting enough to be published, but like almost all active molecules, it wasn't interesting enough to push into the clinic.

Some other inaccuracies I noticed in the post:

"DXO is more lipophilic, meaning it’s attracted to fats and lipids."

Dextrorphan is more polar, and therefore less, not more, lipophilic than dextromethorphan. It is still able to cross the blood-brain barrier because it is still a fairly lipophilic molecule.

"My first thought was replacing the methoxy group with 3-fluoromethoxy. On paper, this seemed like it could be an effective way to interfere with CYP2D6-mediated demethylation, but synthesizing a fluorinated DXM analogue would be difficult and costly."

There is no such thing as a "3-fluoromethoxy" group. What I think the author meant was "trifluoromethoxy" ("3-" and "tri" have very different meanings in chemical nomenclature). The trifluoromethoxylated derivative of dextrorphan has, in fact, been synthesized and patented (https://pubchem.ncbi.nlm.nih.gov/compound/24997117).

"The structure of DPO is (+)-3-(isopropoxy)-N-methylmorphinan, with the 3-methoxy group of DXM replaced by a 3-isopropoxy group."

The "+" indicates the direction that a chiral nonracemic compound rotates plane-polarized light in solution. While there are computational methods to predict it, it can't be assumed that the derivative of a "+" compound will also be "+". Better to use "R" and "S" (CIP nomenclature), since they are determined entirely by the 3D structure of the chiral compound and don't require a physical measurement to determine.


This stood out:

"Artificial Analysis Intelligence Index v4.1.1

61.2"

So on the Metacritic of LLM benchmarks, it's.. basically where everyone else is (except for Fable 5.1, which is a bit ahead).


On their Agentic Index, GPT-6 Astra (both max/xhigh) has the same result as Qwen3.8-27b. Weird.

May be Qwen3.8-27b is AGI, too.

Where did you see this? I haven't been able to find any benchmarks.

It was in the link in the parent of the thread to which I replied. But you can find it on Artificial Analysis's website now.

Zitron has staked his bear position and isn't budging, so regardless if he's been wrong and wrong again, he'll be remembered for calling the bubble if/when it pops, if only because so few in the media have done so without equivocation.

Fable 5.1 is actually more expensive than 5.0 when run on the Artificial Analysis suite:

https://artificialanalysis.ai/#intelligence-efficiency-tabs


The tests measure Fable 5.1 (with fallbacks). The increased cost can come from Fable 5.1 triggering fallbacks less; which means less (cheaper) Opus when AA ran it.

IPO is nearing, must squeeze users more...

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: