> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
You say that they pass the Turing test yet every post on HN complains about the way LLMs write so clearly they haven’t passed it yet because we can still tell it’s a bot.
The Turing Test is about being able to figure out if “someone” is a computer during a short conversation, not “millions of articles”.
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
We don't train current LLMs to mimic the average human's writing style. We train it to be smart, helpful, and knowledgable, too knowledgeable for a human. We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.
I think of the Turing test as one of the starting lines, along with image recognition ("a summer break project for a group of grad students" resisted being solved for decades).
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
I just asked Fable 5 max to create a 2d game about caterpillar climbing a tree and eating fruits.
The game looks good - animations, 8-bit aesthetics, procedural tree branching, but the tree's branches are dead ends. You can't go back once you started climbing a branch.
Yes, LLM doesn't have a reliable way to test it's game yet. All the screenshots, and playwright tests will never be enough to test even a simple game.
But can we call a machine doing such mistakes a general intelligence? It has no embodied intelligence. No way to experience time the way we do. All it has is text. Yes, they can have images, sound too, but no big models (at least those we are supposed to use for coding) currently are native with video as far as I am concerned. And I am not sure that just video without embodied experience is enough to understand the world the way humans do.
Of course we can get incredible results from machines that have a very different experience of the world than we do. But is this a general intelligence? I guess "general" is supposed to mean being able to do everything any human can do (minus the skills requiring a body)?
Current Gemini Flash models can take video input. They're not hyper-specialized coders, but they're better than the competition on many tasks. They seem to be better with spacial reasoning, as well - they are the best choice for OpenSCAD, for example.
I think it's Karpathy who coined the term “jagged intelligence”. LLMs are both extraordinary smart in domain they have been explicitly trained on (like Math) and positively dumb on things they haven't.
Would you also consider a database of questions and answers smart? LLM are basically lossy text compression databases with a clever query method. Useful for sure but it’s not thinking, it’s recall.
I wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.
You need to be smarter (or rather: more knowledgeable in the problem domain) than the model to be able to use it efficiently. Hallucinations are still a problem occasionally but a bigger one is failure of imagination. Even Claude Fable lacks a holistic understanding of many domains it wasn't obviously trained on. The biggest problem with AI (if we assert that LLMs can be the basis of AI) is that these models will make mistakes that exist in an entirely different category of the kind of mistakes humans will make.
As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
A simple example would be "work to rule": in many professions work processes are heavily regulated (whether by law or by corporate guidelines) but the unspoken assumption is that you know which rules you should ignore and which ones you actually need to follow - but if you tried to find this out by asking "is this a rule I need to follow or not" you would get the clearly incorrect answer that all rules must be followed; of course if you did follow all the rules (aka "work to rule") you would be disciplined for failing to meet quotas (because you can't be disciplined for following the rules).
I think that's the crux. As soon as something turns into a commodity, it will be priced as a commodity. It stops being economically valuable and can only be economically viable at scale.
Take mathematical calculations, styrofoam, LCD screens, ice cubes, embroidery (try searching for "computer work blouses"!), texts, navigation. The thing itself costs nothing, the service and theater around it becomes everything.
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
Which brings up another point in that we should have a term that distinguished between super intelligence in some capacity and godlike super intelligence.
ok so first the industry takes the AI term, uses it with a very vague resemblance to what it used to mean, for marketing purpose, then they latch on to the AGI term to talk about what people used to consider AI. Now there is no sign of actual AGI happening any time soon, so we're going to reinterpret AGI to mean something diminutive like a chatbot?
Don't you see a problem here? Terms are used to describe the world and need a semblance of stability so we don't end up in a race to the bottom just so investors can feel good.
> Now there is no sign of actual AGI happening any time soon
Do you genuinely hold this position, or do you not realize how far the goalposts have shifted?
In 2022, prominent AI critic Gary Marcus offered to bet $100,000 that we wouldn't have AGI by 2029. https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi... Because the definition of AGI is unclear, he defined that AGI would be achieved if an AI model could do THREE of the five following tasks:
- In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.
- In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.
- In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).
- In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
- In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
Today's AI models can do FOUR of these five.
Using 2022 goalposts, we already have AGI. We blew past these goalposts months ago, and nobody noticed.
That’s his definition and in no way a universal one.
In the meantime, it’s still very easy to differentiate between an AI and a human in a chat. You just need to know the quirks of these systems. Like counting letters, hitting the safeguards, etc.
So call me when one of them can pass the Turing test against me and then we can talk about AGI
There is no accepted definition of intelligence that is usable for classification of intelligence across the sciences.
On top of that for your strong feelings, you don't have the conviction to write down a strong definition of intelligence yourself, which allows you to accelerate the goal posts up to light speed. The fun thing about writing out a formal definition is suddenly almost everything or almost nothing, including a lot of humans, has intelligence.
Not basing intelligence on your feelings of the moment makes it a hard thing to define across everything intelligence applies to.
GPT-4.5 passed the Turing Test with a 73% human rating, outscoring actual human subjects. That is, human evaluators considered the AI more human than an actual human, 73% of the time. LLaMa-3.1 was judged to be a human 56% of the time.
The 'strawberry' test was fixed years ago with the invention of CoT; models only fail that test today when thinking is disabled.
It’s interesting that they haven’t declared AGI yet, even as a PR stunt. It can be like a “pre-revenue” tactic. They are pre-AGI so investors can still pour money.
> This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
> Which I think is a fair interpretation of the term
You can't be serious... Case in point - I just asked GPT Sol High to give me a weekly update of local ai changes.
Here's it's first update, which is complete and utter garbage; i.e. it's a lot of words that says absolutely nothing.
That's just a random word generator; AGI? Not even remotely in the ballpark.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.