Humans don't need a lot of training and finite tuning to make only legal moves.
An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.
An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).
How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?
I don't know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.
More importantly, beginner human players don't exhibit that tendency. The history of the position doesn't bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.
Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.
> If you even the playing field and give the LLM the position at each turn, it does not make mistakes.
It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.
Yes, we can come up with all sorts of weird situations where you can get it to be confused. But what I'm saying is it's _trivial_ to give it a simple prompt that prevents it from ever making any errors, and so I don't think it's this big LLM gotcha (of which there are many!)
Nothing is forcing the LLM to play 'blind'. If it's smart, it should be able to create its own representation of the chess board and update it with every move, just like a human would. Any chess engine that's sensitive to how the moves are formatted is clearly not very capable.
A human wouldn't do that, they'd look at the board. I'm not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.
That's what a human will do if they already have a physical board to look at. But if someone, say, posed you a chess exam question via FEN notation, or as a sequence of moves in algebraic notation, you'd sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in whatever format enables it to easily keep track of game state and legal and illegal moves. If it fails to do so, that's a sign of its own limited understanding of chess as compared to a human.
The LLM would only be playing 'blindfolded' if you somehow forbade it from making notes (as you effectively do by literally blindfolding a human, given how limited human working memory is). But you are not doing that. The LLM is free to keep track of the game state via whatever means it chooses.
None of this is about superhuman ability. Any human who understands a given chess notation can convert it to a visual representation of a chess board and then use that representation to choose their next move, with their usual level of performance.
I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you're deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn't a very good thing by which to evaluate LLM capabilities. If there's no argument you'll accept, I am happy to move on.
You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.
> I just think this isn't a very good thing by which to evaluate LLM capabilities
I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)
> If there's no argument you'll accept
It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your very brief comments so far. I could equally well say the same thing to you!
> generally read the rules of a game and then follow them
How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
Chess.com has to accommodate people who haven't learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they're often outlining sequences of multiple "pre-moves" during the opponent's turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.
An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
Do you know what the "General" in "Artificial General Intelligence" means? It specifically means that the AGI adapts to novel domains that it hasn't been trained on - its training generalizes to real world problems.
That doesn't mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.
On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.
I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe
> AfD is not democracy. AfD is antidemocracy even if it would get elected democratically.
Here's a fun thing language-game to try: expand your definition of "democracy" to encompass your partisan political positions, don't tell anyone, yell from the rooftops that your opponents are "anti-democratic" and banning them is "democratic."
Then observe how you only convince yourself of your own righteousness, and further alienate the people you wished to persuade.
This is certainly what could be the case, except it isn't. It was decided in 1949 what the current understanding of democracy in Germany is. You could of course decide to define democracy as broad so that it also encompasses the Athenian democracy, but that is not what we currently mean when we speak of a liberal democracy.
If party A, being labeled as communist by party B, and party B, being labeled as fascistic by party A, both agree, as well as all other parties in-between, that you are a threat to democracy, maybe you are.
If you are looking for a party that labels itself democratic and all others anti-democratic, you need to look at the AfD themself.
We have a trade war currently. Direct intersection of those 2 concerns. Some provinces are more open to trade with US while others become manic at the mention of it.
Having 2 groups working on related problems independently does not solve anything.
You design a product/machine 100% digital, you create a digital twin of it, you train a robot ml on this machine, you upload it to the robot who is sitting in the factory and it analyses the machine, reruns a simulaton on how to fix it and executes it.
I would say in 50 years max this is a solved problem. I estimate 30 years and would go down as early as 20 years.
> I would say in 50 years max this is a solved problem.
AI security will be perfect in 5 years, it's in good shape already - my agents have never hacked anybody, those lab LARP-ers better learn something about security and isolation.
> you upload it to the robot who is sitting in the factory
You assume no guardrails, in that case even script kiddies can do more damage than AI.
Or do i miss the point you are trying to do?