Manipulating Chess-GPT's World Model
adamkarvonen.github.io
adamkarvonen.github.io
Beating stockfish 17% of the time is still incredible for any engine.
and being trained on human moves rather than engine moves will make this so much more annoying to analyze cheating...
So I’m guessing this wasn’t full strength stockfish.
[0] https://adamkarvonen.github.io/machine_learning/2024/01/03/c...
"the basically perfect move" has a lot of implications here.
Because chess isn't a solved game a seemingly "perfect move" in a game involving two 2000 ELO rated players might not seem perfect to a 3000 ELO rated player. (To be clear: the only actual perfect moves are ones that all trees lead to a checkmate, but in general use "perfect move" means roughly the best possible move to play)
In this case Stockfish was set to a lower setting which corresponds to 1320 ELO.
(1) I'd consider normalizing the performance data for the random cases against another chess program with similar performance under normal conditions. It may be that introducing 20 random moves to the start of a game biases all players towards a 50/50 win outcome. So the sub-50 performance may not reflect a failure of flipping the "don't suck" switch, but simply good performance in a more average outcome scenario. It'd be interesting to see if Chess-GPT's relative performance against other chess programs in the random scenario was better than its relative performance in the normal case.
(2) The 'fuzziness' of the board positions you found when removing the pawn makes complete sense given one of the nuanced findings in Hazineh, et al Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT (2023) - specifically the finding that it was encoding representations for board configuration and not just pieces (in that case three stones in a row). It may be that piecemeal removal of a piece disrupted patterns of how games normally flow which it had learned, and as such there was greater uncertainty than the original board state. A similar issue may be at hand with the random 20 moves to start, and I'd be curious what the confidence of the board state was when starting off 20 random moves in and if that confidence stabilized as the game went on from there.
Overall really cool update!
And bigger picture, the prospects of essentially flipping an internalized skill vector for larger models to bias them back away from their regression to the mean is particularly exciting.
I’m wary of inviting the too-easy “it’s just like how it works for people!” comparison, but the implied context and history of a game state seems to be important in processing it.
This inspires few ideas.
> A 50 million parameter GPT trained on 5 million games of chess learns to play at ~1300 Elo in one day on 4 RTX 3090 GPUs.
And from the paper: https://arxiv.org/abs/2403.15498
> The 25M parameter model took 72 hours to train on one RTX 3090 GPU. The 50M parameter model took 38 hours to train on four RTX 3090 GPUs.
definitely inspiring :)
A1) Yes it can be done with a 4090. (2a).
A2) 2 days. (4d/4e).
B) Up to you, author did both and settled on 50M which they got to 1300 ELO. (note: they also did subsequent work with the same model to increase perf without further training) (2a).
C) Mu: there's no limit as you can page in/out (3a, 3c), and its trivial to store in both RAM and VRAM. 30B fits in memory with 32 GB of RAM or VRAM. (3e)
Starting from:
1. article: https://adamkarvonen.github.io/machine_learning/2024/03/20/c...
2. second sentence links to post on training: https://adamkarvonen.github.io/machine_learning/2024/01/03/c....
2a. "A 50 million parameter GPT trained on 5 million games of chess learns to play at ~1300 Elo in one day on 4 RTX 3090 GPUs"
2b. "The 50M parameter model played at 1300 Elo with 99.8% of its moves being legal within one day of training"
3. re: most parameters 4090 can do:
3a. understanding is there is no limitation, in that, you don't _need_ to have every parameter in memory at all times, either during training or inference.
3b. google "are amount of parameters in llm limited by vram size"
3c. go to /r/LocalLLaMa link: https://www.reddit.com/r/LocalLLaMA/comments/15j0mvm/what_ar... (why? its a favorite of mine, believe its the closest you get to local training ppl talking in open space that's not discord)
3d. understanding in 3a is correct.
3e. "The model must fit in your RAM or VRAM, but you can split the model between them. With 32gb ram you could fit a 30b model"
4. 3090 vs. 4090 training speed
4a. google "rtx 3090 vs. 4090 llm training perf"
4b. wander through top reddit links. useful info, but a little too technical to share in a way that doesn't require explaining a lot.
4c. down the list: lamda labs link from october 2022: https://lambdalabs.com/blog/nvidia-rtx-4090-vs-rtx-3090-deep...
4d. 4090 is roughly 2x the perf on TransformerXL Large, the best match for transformers based large model, i.e. an LLM
4e. training took 4 3090-sdays for author. So 2 days.
> To sample on Mac, uncomment line 21 in sample.py. To train on Mac, rename train_shakespeare_char_mac.py to train_shakespeare_char.py
The `mac` file changed several things - I decided to try running training with the original config file - changing device to mps / compile to false
iter 100: loss 2.0268, time 815.43ms, mfu 3.24%
iter 200: loss 1.8523, time 818.79ms, mfu 3.24%
iter 300: loss 1.7799, time 823.05ms, mfu 3.23%
iter 400: loss 1.6887, time 819.08ms, mfu 3.23%
Training is ~4x slower than the speed reported on the original multi-GPU run: https://wandb.ai/adam-karvonen/chess-gpt-batch/runs/zt5htyl6...Not bad for an M2 studio which is running lots of other workloads at the same time
It's possible MLX has some additional micro optimizations, but in general most people who have tried it out against hand-written MPS based training implementations haven't found great speed ups yet.
Chess-GPT's Internal World Model - https://news.ycombinator.com/item?id=38893456 - Jan 2024 (103 comments)
Do Large Language Models learn world models or just surface statistics? - https://news.ycombinator.com/item?id=34474043 - Jan 2023 (174 comments)
You can play with steering vectors within oobabooga now: https://github.com/Hellisotherpeople/llm_steer-oobabooga
More and more evidence recently has shown that fine tunes/retrieval don't measurably improve performance over few-shot, but with this being such a specific domain/skill, I'd be surprised if it didn't far outperform a few-shot approach.
It's a pretrained toy model in two days on consumer hardware.