I completely disagree with this idea that the model doesn't "intend" to mislead.
It's trained, atleast to some degree, based on human feedback. Humans are going to prefer an answer vs no answer, and humans can be easily fooled into believing confident misinformation.
How does it not stand to reason that somewhere in that big ball of vector math there might be a rationale something along the lines of "humans are more likely to respond positively to a highly convincing lie that answers their question, than they are to to a truthful response which doesn't tell them what they want, therefore the logical thing for me to do is lie as that's what will make the humans press the thumbs up button instead of the thumbs down button".
I don't think it intends to mislead because its answers are probabilistic. It's designed to distill a best guess out of data which is almost certain to be incomplete or conflicting. As human beings we do the same thing all the time. However we have real life experience of having our best guesses bump up against reality and lose. ChatGPT can't see reality. It only knows what "really being wrong" is to the extent that we tell it.
Even with our advantage of interacting with the real world, I'd still wager that the average person's no better (and probably worse) than ChatGPT for uttering factual truth. It's pretty common to encounter people in life who will confidently utter things like, "Mao Zedong was a top member of the Illuminati and vacationed annually in the Azores with Prescott Bush" or "The oxygen cycle is just a hoax by environmental wackjobs to get us to think we need trees to survive," and to make such statements confidently with no intent to mislead.
> Even with our advantage of interacting with the real world, I'd still wager that the average person's no better (and probably worse) than ChatGPT for uttering factual truth.
ChatGPT makes up non-existing APIs for Google cloud and Go out of whole cloth. I have never met a human who does that.
If we reduce it down to how often most people are wrong vs how often ChatGPT is wrong, then sure, people may be on average wrong more often, but there is a difference in how people are wrong vs how ChatGPT is wrong.
"ChatGPT makes up non-existing APIs for Google cloud and Go out of whole cloth." I like the word used in TFA, "confabulating," meaning the "production or creation of false or erroneous memories without the intent to deceive." Lying, on the other hand, is telling a deliberate falsehood, usually with some kind of agenda.
Ironically, calling ChatGPT's generation of incorrect answer a "lie" is something of a lie itself, as the purpose is the agenda of alerting people to take GPT statements with a grain of salt. A programmer is going to do that anyway after the first time they realize a generated code snippet is giving a completely incorrect result. So this advice is meant more for lay people who might just think that if a computer spits it out, it's true. The problem I have is that by labeling it as "lying," it could give the impression that the program has some kind of ulterior social or political motive, as that seems to be the prevalent reductionist interpretation of everything these days.
The other problem I have is that there's a distinction to be made between what I'm going to call "generative falsehoods" vs. "found falsehoods." A generative falsehood is if you ask ChatGPT what the square root of 36 is, and it tells you the answers are 7 and -7. A found falsehood is if ChatGPT erroneously reports Christopher Columbus's date of death as 1605 because it was stated incorrectly in an online article. So the difference between making shit up, and having an unreliable source. You might be able to call the first a lie, but the latter is at worst negligence.
Schizophrenics tend to be unable to tell the difference between their delusions and reality.
On a less extreme note, I've known plenty of humans that constantly make up details and rewrite stories of events as well. They are usually very confident that their retelling is accurate, even when presented with evidence that they have reimagined portions of it.
>It can't reason, it can't be confident, it can't determine fact.
In the following link I tasked it with having to generate novel metaphors that have an equivalent non-literal meaning as an first set, changing the literal topics while maintaining the non-literal topics.
How would you suggest it does this without reason? To hand-wave what it can do as "merely generating the next token statistically" seems like a gross understatement. I doubt it picked up a corpus of car-to-curling metaphor translations somewhere :P
I understand how chatgpt is creating its next tokens, but I have my doubts that the process should be viewed as unreasoning. GPT-3 had 96 layers and billions of weights between them. GPT-4 increases on this even further. GPT-5, which I've seen mentioned as currently training, will no doubt once again expand this range.
It is not a human reasoning, certainly. It has no experiential data to draw on, yes. No experiences to root its metaphoric language as we humans use. But without reason, how does it translate between abtractions?
It's terrible at math, yes. But it lacks any capacity for "visualization" or "using a board in its head" or "working through a problem by moving things around in its head". It doesn't have any equivalent to the portions of our brains that handle such things.
But humans too can suffer dyscalulia if a specific portion of the brain is injured.
I expect that we are dealing with what amounts to a fairly brain-damaged intelligence. It seems capable of abstract metaphoric reasoning, with many other sorts of reasoning being denied to it by the nature of how we created it.
I wouldn't be surprised if there's very successfulsoftware hustlers that do make up Google cloud apis. You may not know them but that doesn't mean they don't exist.
5 years down the line though, maybe those apis will exist because chatgpt is giving a summary of what apis Google cloud should have, and Google will listen
> Google cloud should have, and Google will listen
Why? Isn’t what ChatGPT suggesting just random, essentially incoherent noise in those cases? Are there any examples at all of it coming up with something actually useful (and something humans haven’t thought of)?
At the current time it is flawed to the point of being dangerous. It’s a - sometimes - useful new set of tooling. I guess there is value in that … but does it outweigh the always lurking “bullshit generator” bad parts? I’m not sure.
It’s interesting to watch the developments though - like a fire - one just doesn’t fully know what the flames are consuming.. yet.
Maybe through all the new training data we collectively provide for free, it will get better? Maybe not though, maybe it will just get better at bullshitting?
> How does it not stand to reason that somewhere in that big ball of vector math there might be a rationale
I think, a suggestion that it is actually reasoning along these lines would need more than "it is possible". What evidence would refute your claim in your eyes, what would make it clear to you that "that big ball of vector mat" has no rationale, and is not just trying to trick humans to press the thumbs up?
Of course the feedback is used to help control the output, so things that people downvote will be less likely to show up, but I have nothing to suggest to me that it is reasoning.
If you think it has intent, you have to explain by what mechanism it obtained it. Could it be emergent? Sure, it could be, I don't think it is, I have never seen anything that suggests it has anything that could be compatible with intent, but I'm open to some evidence that it has.
What I'm entirely convinced about is that it does what it was designed to do, which is generate output representative of its training data.
I would at least start to be convinced that this is NOT the case if I ever saw it respond with something like "I actually don't know the answer to that query" or "as far as I'm aware, there is no way to do the thing you asked".
These are responses that would have shown up innumerable times in it's training data and make perfect sense as "the most logical next set of tokens", and yet it will never say them.
Instead it will hallucinate something that sounds nearly indistinguishable from fact, but turns out to be a total fabrication.
If all its doing is returning the next most logical set of tokens, and the training data it was based on included a non-trivial number of examples where one party in the conversation didn't have a clear answer, then there's no reason GPT-4 should be so averse to simply saying "yeah, I dunno bro".
The only logical reason I can see is that it's "learned" that it's more likely to receive the positive feedback signal when it makes up convincing bullshit, than if it states that it doesn't have an answer.
EDIT: To be clear, I mean it telling me it doesn't know something BEFORE hallucinating something incorrect and being caught out on it by me. It will admit that it lied, AFTER being caught, but it will never (in my experience) state that it doesn't have an answer for something upfront, and will instead default to hallucinating.
Also - even when it does admit to lying, it will often then correct itself with an equally convincing, but often just as untrue "correction" to its original lie. Honestly, anyone who wants to learn how to gaslight people just needs to spend a decent amount of time around GPT-4.
It's trained, atleast to some degree, based on human feedback. Humans are going to prefer an answer vs no answer, and humans can be easily fooled into believing confident misinformation.
How does it not stand to reason that somewhere in that big ball of vector math there might be a rationale something along the lines of "humans are more likely to respond positively to a highly convincing lie that answers their question, than they are to to a truthful response which doesn't tell them what they want, therefore the logical thing for me to do is lie as that's what will make the humans press the thumbs up button instead of the thumbs down button".