State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that's difficult to engineer out.
I'm not sure why anyone is expecting stochastic systems to be deterministic.
Chess is a deterministic game won by a combination of known movesets and constrained multi-level forward search.
LLMs do neither of these things. They don't reproduce training data exactly, their next response is more 'inspired by' prompts and its own memory than produced deterministically, and they don't have the capability to do general forward search on their own.
So when you ask an LLM to play chess you're getting the equivalent of a very compressed and lossy JPEG of chess rules and strategies with added per-turn random noise.
They also don't have the ability to design their own chess engine, although it would be interesting to see what happens if you ask for one.
For me the useful intuition is that LLMs haven't somehow magickally learned to implement any of the algorithms we know that we have used to make strong chess engines: alpha-beta minimax and Monte-Carlo Tree Search on the one hand, and obviously the ability to learn accurate evaluation functions by self-play.
I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to do that in the process of learning to represent the entire text on the web. On the other hand they have gotten say very good at machine translation without being trained exclusively (and I select the preceding word carefully) on machine translation.
Edit: I'm saying this because there is this idea expressed by e.g. Ilya Sutskever, that in order to predict the next token accurately an LLM has to learn something about all of underlying reality. See for example this interview with Dwarkesh:
Where Sutskever claims that "Predicting the next token well means you understand the underlying reality that led to the creation of that token".
If that were true, we should have seen LLMs play good chess by now. There is a huge amount of data on playing chess floating around on the web in the form of algebraic chess notation and if LLMs were capable of learning the "underlying reality" of chess, they would already have. They haven't. Because they can't. What Sutskever is saying flies in the face of literally hundreds of years of statistical modelling, which is to say, building predictive models that, very explicitly, do not have to understand any "underlying reality" and only have to be good at modelling a dataset.
>If that were true, we should have seen LLMs play good chess by now.
Not at all. LLMs learn by imbibing a mass of relationships as isolated fragments of information. There is a certain amount of sorting and indexing that happens during the training phase. There is also a certain amount of compute executed on these relationships during inference. LLMs can model processes that fit within the compute budget. Language translation works well because language is lookup-heavy while being light on compute.
Chess is a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Humans cut through the compute requirements by reinforcement and learning intuition. LLMs don't get reinforcement on chess so they must compute during inference a unified model of chess. Developing a strong model of chess from raw fragments of information is simply not in their compute budget.
>> Not at all. LLMs learn by imbibing a mass of relationships as isolated fragments of information.
You gotta be careful how you use the word "relation" here because there's an informal meaning (I'm related to my cousin) and a more strict, formal meaning, that is used in computer science e.g. in the "Relational Calculus" etc. In the formal sense, the one relation that LLMs learn during training is the co-occurrence of tokens in a corpus of text, what's called more technically a "collocation" relation. Nothing says that this is enough to play chess, so I'm indeed doubtful that they can.
But they have "learned to implement any of the algorithms we know that we have used to make strong chess engines". Ask Claude Code to write you a chess engine. Your objection is that they don't implement MCTS in the neurons themselves? Neither does a human, we use a C compiler when we want to play chess using MCTS.
That's a separate question from whether an LLM (unaided by a C complier) can learn to play chess as well as human (also unaided by a C compiler). Certainly humans can't become grandmasters only by reading chess transcripts on the web, and certainly humans require many "thinking tokens" during a game to play effectively. Do you know for sure that a transformer can't reach grandmaster level if it is allowed to learn by playing games (as humans do) and is given a sufficient number of thinking tokens during the game? It seems near certain that they could, if someone wanted to spend the money (and I don't see why anyone would.)
Sutskever's claim is that in order to predict the next token a system must learn something about the "underying reality" that produced the token. In the context of chess that means that the LLM must learn something about playing chess (since tokens are the moves in a game of chess). My argument is that contrary to what should be expected if we take what Sutskever says to be true, they don't seem to have.
Yes, I do mean that the LLM's weights are set so that it will execute minimax or MCTS when it needs to. That has nothing to do with whether humans can do the same or not.
I don't disagree that a Transformer could learn to play chess if it was explicitly trained to do that. My argument is that LLMs, trained to predict the next token, have not learned to play chess. That's LLMs, not Transformers.
Just to make sure this is not taken as splitting hairs, the point is that there's all sorts of claims made about what LLMs learn when they train on text. For example, there was a claim by Sundar Pichai that one of their models had learned to translate Bengali without explicitly being trained to do so. It later emerged that Bengali was indeed included in the model's training set [1]. It's not clear whether that included parallel texts, e.g. between Begnali and English or another intermediary language, in any case Sundar Pichai's claim was that the ability to translate Bengali was "emergent".
So I'm interested in understanding the extent to which these "emergent" abilities are real or not. With chess, given the amount of textual data tracing games that floats about on the open internet, I would totally except some ability to play chess to "emerge". Maybe the reported 700-800 ELO level is even that sort of ability. Maybe we should only expect LLMs to learn to play at the level of an untrained, casual player. Maybe not. I have no idea.
On the other hand, the fact they keep making elementary mistakes like illegal moves must be taken to mean that, so far, LLMs haven't learned to play chess.
But it speaks in words, therefore it must be super duper extra smart!!11 /s
Sarcasm aside, I think this is an easy cognitive trap to fall into. It does sometimes feel like the LLM must have some world model because it converses somewhat coherently. Examples like this failure to understand chess, or to count the number of Rs in "strawberry", seem difficult to explain if the models are intelligent. But that doesn't stop people believing they are anyway. I think there must be something about the conversational interface that fools us easily. I wonder if people trained in interrogation techniques are also fooled?
I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.
If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.
FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word. Have at it, but don’t let anyone know!
You hung up a “services offered” flyer on every town square around the world and want to control what people do with the information after you’ve shouted it into the void? Don’t post if you don’t want every tom, dick, or jAIne reading your stuff.
There are two internets: the one on your phone (ruined by outrage and mobile slop) and the one on your computer (ruined by marketing slop). They’re both dangerous and disappointing for different reasons - but mostly because independent voices got drowned out.
Cutting them down isn’t the worst idea but it would be nice if we could just elect people that wouldn’t allow this in the first place. Knowing that won’t happen, the only chance reasonable people have to stop this is to turn these systems against the decision makers until they yield.
Some kind of Flock@home volunteer surveillance network that tracks politicians, records LEO interactions (including those involving masked unaccountable agents), and provides the only accountability that the public has left.
What do you mean "allow"? Police chiefs are not elected positions and they often find funding outside of the town's budget for these automated bill of rights violation machines.
Sheriffs are elected (except in New York City, Rhode Island and Hawaii) and responsible for a county whereas a police chief manages law enforcement in a city or town and is appointed by its local government.
> Cutting them down isn’t the worst idea but it would be nice if we could just elect people that wouldn’t allow this in the first place.
Is this really better though? First-party community maintenance is more direct, cheaper, and doesn't rely on fragile central authorities (granted, it _does_ rely on a difficult pattern of community discussion and collaboration to satisfy tribal consensus, etc).
I'm broadly OK with cameras being in public, but I have much more trust in direct action than I do in elected "officials".
On medium time-scales, it seems to me that a great solution will be to have community-maintained cameras kinda all over the commons, but such that the camera eye can be accessed publicly (ie, removing the information asymmetry which gives the undesirable actors (police, etc) greater view than the actual community participants).
Elon Musk fully tweaked out and I believe banned his twitter account. There’s some people who think he bought out twitter purely for this power, to unilaterally ban people who annoy him. He also offered some money for the kid to take it down, which he refused. It was really not enough money if I remember correctly, maybe 10,000 dollars. Me personally, if I was squeezed to give up my conviction, I’d at least want to maximize my profit.
I agree with you about juniors and remote work - but like everything else: just figure it out.
Your customers aren't sitting next to you several times a week, yet they have no problem paying you and expecting you to get your job done. Sure, you might need to meet in person a couple of times a year but if you're smart enough to use a computer you can figure out how to work with someone in a different location. The issue is probably not the distance.
Yes, much like the “we’re all family but we’re reducing headcount by…”
Treat people like professionals: admit that you did a bad job but that you’d rather protect yourself and not give up any of your salary/benefits and instead you are going to layoff your “family”.
It’s more than one thing. Yes there’s definitely some populism and maybe some astroturfing from outside groups - but actual electricity costs are going up and the average consumer can’t really get anymore efficient, so they’re indirectly subsidizing this stuff. Computers and related hardware have gone up in price. And then you’ve got all of the “AI layoffs” that also result in real-world pain.
American electricity costs a very low. What they've gone up is nothing its a drop in the bucket compared to the damage from the tariffs and middle east war. Populists settle on this issue because its malicious. They are Accelerationists
I see some similarities to 3D printing here. It’s great that everyone can make their own toothbrush holder (or whatever) but I’m probably not going to pay for someone’s weekend project.
I’m “seeing” more devs stepping into the SendCutSend stage where they’re cleaning up/fixing/productizing vibe coded projects so maybe there will be some new demand in that space?
3D printing is a good comparison - it allows almost anyone to make things, but in the end very few do.
Another example is when the WWW first became available, and suddenly everyone COULD be a publisher (browsers even included built-in HTML editors), and for a while MySpace pages proliferated until the excitement died down and people went back to being media consumers.
I expect we'll see the same thing with consumer use of generative AI. Suddendly everyone is generating 3-D worlds/games with Fable because they can, but I expect that just as with the web the novelty will wear off and they'll leave it up to the pros.
Professional use of GenAI, and coding in particular, is certainly here to stay, but it seems we're still in the early experimental/hype phase. At least tokenmaxxing has passed, and it seems most companies are now paying attention to, and limiting, how much they are spending, but it doesn't seem we've yet progressed to the stage where companies are paying attention to what they are actually getting out of it - is the money spent showing up on the bottom line in the form of increased revenues.
A comparison I find useful here is Excel (and spreadsheets in general). Those enabled huge numbers of non-programmers to build software-like things, while the demand for expert developers grew enormously at the same time.
It’s terrible and depressing work to take vibe coded garbage and make it a real product. There will be demand, but good engineers won’t want to touch it. And people paying will think they did the hard work so why pay a good rate?
These are all very neat things and I’m very pleased to see the hacker spirit move into AI, but I gotta admit - some of this just comes across as “what did you print with your printer?”
Another recent thread mentioned that AI has helped devs build better “shop jigs.” This seems to be where the rubber is meeting the road for AI-powered development. So maybe more people will be developing custom tooling for their own little problems but you still need reliable, deterministic, interchangeable tools to realize the value of all these shop jigs.
reply