Hacker Newsnew | past | comments | ask | show | jobs | submit | Gregkion's commentslogin

And you know compare Eliza with what an LLM can do today?

Or do i miss the point you are trying to do?


Thats just absolutly not true.

A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

And there is a relevant and significant difference between the expectation of an AGI and an ASI system.


Humans don't need a lot of training and finite tuning to make only legal moves.

An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.

An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).


How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?


I don't know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.


More importantly, beginner human players don't exhibit that tendency. The history of the position doesn't bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.


Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.


> If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.


Yes, we can come up with all sorts of weird situations where you can get it to be confused. But what I'm saying is it's _trivial_ to give it a simple prompt that prevents it from ever making any errors, and so I don't think it's this big LLM gotcha (of which there are many!)

I've not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.


I haven't tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: https://news.ycombinator.com/item?id=49720751

Nothing is forcing the LLM to play 'blind'. If it's smart, it should be able to create its own representation of the chess board and update it with every move, just like a human would. Any chess engine that's sensitive to how the moves are formatted is clearly not very capable.


A human wouldn't do that, they'd look at the board. I'm not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.


That's what a human will do if they already have a physical board to look at. But if someone, say, posed you a chess exam question via FEN notation, or as a sequence of moves in algebraic notation, you'd sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in whatever format enables it to easily keep track of game state and legal and illegal moves. If it fails to do so, that's a sign of its own limited understanding of chess as compared to a human.

The LLM would only be playing 'blindfolded' if you somehow forbade it from making notes (as you effectively do by literally blindfolding a human, given how limited human working memory is). But you are not doing that. The LLM is free to keep track of the game state via whatever means it chooses.

None of this is about superhuman ability. Any human who understands a given chess notation can convert it to a visual representation of a chess board and then use that representation to choose their next move, with their usual level of performance.


I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you're deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn't a very good thing by which to evaluate LLM capabilities. If there's no argument you'll accept, I am happy to move on.


You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.

> I just think this isn't a very good thing by which to evaluate LLM capabilities

I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)

> If there's no argument you'll accept

It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your very brief comments so far. I could equally well say the same thing to you!


The actual question is backwards: how do we keep the prompt and context small enough so the LLM doesn't start hallucinating basic rules of chess.


How much support do we as humans need to get rules right?

I'm an expert in my field, read my comments, my gramma is shit.


So we humans are not a general intelligence then?

And the stuff i'm using LLMs daily is just fake?

I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.


> So we humans are not a general intelligence then?

No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us.

> And the stuff i'm using LLMs daily is just fake?

This misses the point completely.


> generally read the rules of a game and then follow them

How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.


Chess.com has to accommodate people who haven't learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they're often outlining sequences of multiple "pre-moves" during the opponent's turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.


Do you play chess ? Do you even know what an illegal move is ?


If you have something to contribute to the discussion, just say it


An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.

And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.

This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing


Do you know what the "General" in "Artificial General Intelligence" means? It specifically means that the AGI adapts to novel domains that it hasn't been trained on - its training generalizes to real world problems.

That doesn't mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.


On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.

I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe


And you want to indicate what with your smile?


No but a democracy should protect democracy.

AfD is concluding with russia, a country which poisned people on EU ground.

AfD also has plenty of real nazis doing nazi shit and AfD is not excluding them.

AfD is not democracy. AfD is antidemocracy even if it would get elected democratically.


> AfD is not democracy. AfD is antidemocracy even if it would get elected democratically.

Here's a fun thing language-game to try: expand your definition of "democracy" to encompass your partisan political positions, don't tell anyone, yell from the rooftops that your opponents are "anti-democratic" and banning them is "democratic."

Then observe how you only convince yourself of your own righteousness, and further alienate the people you wished to persuade.


AfD is a facist party and anti-democratic.

AfD fights against mechanism we have in place to protect the democracy.

Has nothing to do with my own righteousness. AfD could easily do things against all of this garbage and they don't.


This is certainly what could be the case, except it isn't. It was decided in 1949 what the current understanding of democracy in Germany is. You could of course decide to define democracy as broad so that it also encompasses the Athenian democracy, but that is not what we currently mean when we speak of a liberal democracy.

If party A, being labeled as communist by party B, and party B, being labeled as fascistic by party A, both agree, as well as all other parties in-between, that you are a threat to democracy, maybe you are.

If you are looking for a party that labels itself democratic and all others anti-democratic, you need to look at the AfD themself.


There is a group of people in canada which takes care of QoL and another group of people taking care of external relationships and trade.

And normally this helps a country?


We have a trade war currently. Direct intersection of those 2 concerns. Some provinces are more open to trade with US while others become manic at the mention of it.

Having 2 groups working on related problems independently does not solve anything.


You design a product/machine 100% digital, you create a digital twin of it, you train a robot ml on this machine, you upload it to the robot who is sitting in the factory and it analyses the machine, reruns a simulaton on how to fix it and executes it.

I would say in 50 years max this is a solved problem. I estimate 30 years and would go down as early as 20 years.


> I would say in 50 years max this is a solved problem.

AI security will be perfect in 5 years, it's in good shape already - my agents have never hacked anybody, those lab LARP-ers better learn something about security and isolation.

> you upload it to the robot who is sitting in the factory

You assume no guardrails, in that case even script kiddies can do more damage than AI.


Does that really matter?

We and a LLM are doing the same thing: trying things out while skipping unrealistic/wrong things.

If an LLM can do the same thing as we do but a lot faster, t already won and this is just another example.


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: