First place in Tetris 99 using computer vision, classical AI, a lot of free time
bpinzone.github.io
bpinzone.github.io
Classic RL reward-hacking! You'll often see 'suicidal' agents when the reward-shaping goes wrong, or they converge prematurely after inadequate exploration. (Tom7's Mario agent famously pauses the game to avoid losing.)
> I’m not sure if there are good ways around this. Rigorous testing, maybe, but it’s hard to improve your algorithm by observation once it’s not making obvious mistakes. This is, perhaps, one of the promises of using reinforcement learning without human play training: if your algorithm is able to achieve human-level performance, it’s probably correct enough to continue well into superhuman performance.
In this case, because the random number generator's decisions are unaffected by your actions, you can use the clairvoyant oracle trick, I think. Just record games, and then do your DFS planning over the entire game with the known sequence of blocks; this then represents perfect play and is a benchmark for your actually feasible algorithms.
How much unnecessary frustration was caused by testing this out in prod? I know they said they only got a handful of wins, but how many people did it knock out over the course of testing it?
It seems like tech has gotten so dang self-absorbed.
They don't do T-spins. They don't stack large attacks for KO's, or even seems to alter targets under any condition at all even. The targeting reticule is hard stuck on attackers with zero input put into it. So it has no overarching strategy besides survival, but even then, it doesn't try to play the badge stealth game either - backstabbing a leading player is the leading method of making comebacks, and necessary ones too. Which is why it reliably gets in the top 15 and then loses, you can not outfire someone without badges, there ARE one-shots in Tetris 99, the bot probably only won when it lucked out early on in being the leader of the pack and then fed on a crowd of people with targetting set to "badges" (there's a firepower multiplier when targeted by many people).
Which is... kinda pathetic reading into when I'm someone who reliably got top 3 outside of the Invictus mode. I think the creator might have 0 idea of what defines Tetris 99 at all. Like, straight up no awareness whatsoever. They're bad at this game, and thus, made a bot equally bad and only wins with luck.
It's just playing normal Tetris. It literally has no reason to do this against humans, does nothing with the human element. So it'll beat other clueless humans and then lose to the actually good humans. Between those two human groups, they're just one more in a sea of typical indistinguishable threats and fodder.
Good on the creator for stopping before making something more competent.
Initially, we chose Tetris 99 because that was the context in which we had the idea. I won't deny, though, that being able to compare Jeff's performance with human players and see his interactions and victories in the Tetris 99 environment made the project more interesting for us; however, it was wrong to do so in a public setting where the players didn't know that our algorithm was involved and didn't opt in to playing against it.
We're no longer testing or developing Jeff, and if I do any similar future projects I'll choose singleplayer games or private environments where opponents voluntarily test their skills against the bot.
After you place first it unlocks a level where you battle against only other users who have also placed first. I don't see any mention of this second level in the article (though admittedly I skimmed it), which makes me suspicious of whether they actually got their bot to place first. Obviously the next step would be to try it out on that harder level.
Also, more often than not, there are not 99 player characters who you are battling against. The field has always been filled out with other bots, generated by the game designers themselves. I don't think they would ban a bot player from the game.
Plenty of games have both official bots and anti-cheat and player botting rules, the presence of one does not define what the presence of another will or won't be.
My only gotcha is with calling DFS 'classical AI'. I get they have a fancy hand-tuned cost function, but if that is AI, then I guess all approximate optimization algorithms are AI? Maybe in 60s that kinda of classification would have been fine, but not now, right?
Either way, that is a minor nerd-gripe with the messaging and does not, at all, detract from an otherwise rad project.
Thus classical, as a reference to the idea that:
As Soon As It Works, No One Calls It AI Anymore: https://quoteinvestigator.com/2024/06/20/not-ai/
It’s kind of surprising to me we still call feedforward neural networks AI.
I guess...hrm... if someone asked me what classical AI was today... I would answer with something like A*, decision trees, or Markov model-style algorithms (but never DFS). All of these have 'worked' for a long time but still feel like classical (ie no back-prop) AI.
To your second point, I'd be the first to admit I'm no expert (I don't do any ML professionally). In my limited understanding, it feels odd to separate MLPs from 'modern' AI. It's all still back-prop, yeah? I get that universal approximators and (idealized) Turing machine approximators like LSTMs are old hat. But Attention-based (key, value) query models just seem like the next step on the same ladder of "increasingly complex computing paradigms we've figured out a way to back-prop through." Not something uniquely separate. First, we could do functions, then we could do Turing machines, and now we can do data queries.
But again, I'm probably just showing how green I am. I would, time permitting, honestly appreciate clarification/correction.
It is the future of cheating, completely undetectable
And that would still be in the realm of "robotic" behaviours. We're beyond that point. Hearthstone is a game where bots will do everything in their power to fake humanity, even display intentionally toxic behaviors, like using certain lines in certain scenarios specifically known to mean to aggravate you (hovering cards that will get lethal next turn back and forth followed by Priest's Hello) or even straight up roping.
It looks like the bot just beat humans because they react faster
For example, at the beginning of the game, it shows you 6 pieces, but you can determine the 7th piece because it's the one that's not shown. The 8th piece could be any of the 7 pieces, but when you see the 13th piece, you know the 14th piece (and when you see the 12th piece, you know there are only two possibilities for the 13th piece).
Additional look ahead every so often might be helpful, given that they found
> When looking only 3 moves ahead, Jeff achieves a Tetris percent of 8.9%. When looking 6 moves ahead, he improves to 9.7%—nearly perfect!
I'd also be interested in seeing optimization for combos. Or t-spins; t-spins send 2x lines of garbage for each line cleared; which I hate, because I grew up before t-spins were encouraged. :P
Wait, is this true?
I always assumed each piece was selected perfectly randomly, making it possible (though rare) to get the same piece 3 times in a row.
If what you're saying is true, then that mean there should never be more than 12 pieces between I pieces, and if you get two in a row, then it'll be a minimum of 6 before you see another.
Which...all seems within the realm of possibility. Tetris always seemed really good at having a very even spread of pieces without "streaks".
https://harddrop.com/wiki/Tetris_(Game_Boy)#Randomizer
L: 10.7 %
J, I, Z: 13.7 %
O, S, T: 16.1 %> mention graph search for vision correction? https://github.com/bpinzone/TetrisAI/blob/master/tetris_ai_e... do we still use this?
Bad players are so desperate to cheat in order to win. They don't belive at fair competitions, but instead will use any dirty trick to win.
And this is also why most online games are full of invasive anti-cheat crap.
My friends and I have quit a number of multiplayer computer games when they've become infested with cheaters and developers couldn't or wouldn't keep up with the arms race of banning them. This is really not that different from someone writing and using an aimbot in a public lobby for a shooter.
If you're going to be writing and testing cheats or AI or bots do it in single player games or private multiplayer lobbies so normal players aren't affected.