Hacker Newsnew | past | comments | ask | show | jobs | submit | avereveard's commentslogin

btw that how mmlu score things to answer question instead of producing all the answer token they look at logprob of a b c d keys in 2020 making this technique old as dirt in nlp

This technique is so obvious to anyone who spends more than a minute with multiple choice tasks. It's wild they're claiming it as a feature.

Aws cto is basically opening every presentation with a everything breaks all the time slide, idk where you get that idea they oversell the cloud resilience

They could offer controlled "failure of service" service where they randomly take stuff offline and you have to pay to get it back if you don't have a backup/recovery strategy.


The customer CEO never watched that presentation.

Opus specifically talks as if having a stroke. 4.6 was the last version that was pleasant to work with

I couldn't even stand 4.6 by the end. 4.5 was very usable, 4.6 started alright and somehow got more annoying. 4.7 made me quit my subscription and ditch their services entirely - it's an honourary member of my quite short "Coworkers I'd like to throttle if I didn't work remotely" list. Infuriating to instruct or communicate with.

Yep the only way out is hooks to forbid what can be detected by ast and second model to prune comments, flatten pyramids of fallback, and squash the test suite removing quirks maintaining wanted behaviors.

A tool aggregation layer made of code is a good way to save tokens

Composition is an issue only as long as one keep demanding tool calls in json. If tools are goal predicates in prolog, it's easier.


Have your tests been ran on vpn with vpn cloaks?

I explicitly say, that I won't be surprised if the results change from Russian territory. I already put way too much effort to check the validity of an internet comment, but if anyone is interested, they can check.

Yes nobody is producing early 90s punk with the same intensity and dinosaurs never died, sorry nofx, so now I have my personal playlists.

Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.

No but non enterprise users will receive injected prompts.

See how emany streaming now created a plus plus version while adding ads to what was the premium version.

And the funny thing they just needed to increase base rates for users to start to ask for an ads supported version.


This era reminded me of late 90s web culture, and it's heading the same way.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: