That said, even without that database a modern AI will completely topple the best human at every common chess variant. Humans cannot defeat modern AIs in chess like games.
I'm sure some of those games are actually stockfish v stockfish or something similar. Its pretty easy to run stockfish or lichess locally and copy the moves from each engine back and forth.
Sure, some people are cheaters. Some are not. There is no personal win in cheating against Stockfish. Usually strong players do it for training purposes, or to entertain their watchers when they stream. I actually remember having seen one who did that, and he drew. That was a party.
Evidence given: "There exist some small number of games on lichess.org played against stockfish where the user won."
My counter argument is that games on lichess against stockfish don't imply a human beat stockfish. It could just be that stockfish (or other bots) can sometimes beat stockfish. And some humans surely use bots to play on their behalf in order to cheat in online games.
I don't know if any humans can beat stockfish. But I don't consider that to be strong evidence.
To assume that a human can beat that is just delusional.
Especially pre-NNUE, chess engines were often not fully well-rounded, and therefore a human with specific knowledge of the chess engine's weaknesses could take it down with enough attempts.
Also, Lichess' Stockfish runs in the browser (with all the slowdown that entails), plus is limited to one second of thinking time even on the highest level. It also has no tablebases and AFAIK no opening book. Even if you _can_ consistently beat Lichess Stockfish level 8, there's still a very long way from there to saying you can beat Stockfish at its maximum strength, which is generally what people would assume the best humans would be up against in such a duel.
People generally don't play unencumbered engines anymore because the result isn't interesting.
Add to that 24 years of hardware development, and you can imagine why no human player is particularly interested in playing full-strength engines in a non-odds match anymore. Even more so in FRC/Chess960 where you have absolutely zero chance of leading the game into some sort of super-drawish opening to try to hold on to half a point now and then.
I have achieved these results around 2015, sitting at home, relaxed. I was not in a match situation observed by millions. Such a situation can knowingly lead to blunders like Kramniks overlook of mate in 2.
I also sometimes "cheated" by aborting the game when I was tired and continuing it the next day (if at all). That's what the player in a match can not do.
I also sometimes restarted a game at a specific position. Can also not be done in a match. Finally, they used better hardware in these matches. I had eight threads on my old Laptop and I used four of them. The Laptop itself was bought around 2005. Between 2000 and approximately 2020 I trained every day and I was on my peak. I am still around 2400 on Lichess today, without training.
So, I hope it does not sound that extraordinary any more. It isn't. Maybe it is now, but not then.
Based on what data I can find, it's estimated that the difference between the 2025 stockfish (stockfish 6) and today's stockfish (stockfish 18) is nearly 400 points.
That's the difference between Magnus Carlson at his peak and someone who doesn't even have enough rating to qualify for the grandmaster title.
So yes, the fact that you beat stockfish in 2015 doesn't sound extraordinary, because AI today is vastly stronger than it was when you achieved those results. What sounds extraordinary to people is your belief that you could repeat those results against today's top chess engines.
In fact, there is only one game I could find in all of Chess history (Anand vs Touzane, 2001) where a super GM (rating >2700) dropped a classical game to someone more than 350 points below theirs (gap: 402 points). (it's estimated that there are between 2000 and 3000 classical games in history played between Super GMs and players >350 points below them) And it could easily be that Anand was ill, or suffering some other human condition which made his play significantly worse than his typical play for that game - which you would not see from a computer engine.
In other words, the Stockfish that you beat in 2015 would itself be expected to get 3-5 points (that is, 6-10 draws and 0 wins) in 500 matches against the best chess engine of today. The delta in strength is immense, and it is reasonable for everyone else in this comment thread to assert that you would have zero chance at all of picking up a draw against Stockfish 18 in a fair game of any time control, regardless of how many matches you played.
P.S: You should not take this bet. You will lose. You are mistaken if you think you beat stockfish.
There are some games of knight odds Leela playing superGM's. For example, Hikaru Nakamura went 1 win, 2 draws, and 13 losses against LeelaKnightOdds at 3 minutes + 2 sec increment: https://www.youtube.com/watch?v=pYO9w3tQU4Q So that's a score of 2 out of 16. Which is apparently actually very good. I know Fabi played a lot of games too, and also lost almost all of them.
And that is with knight odds lol. And stockfish is ever better than Leela, but generally less aggressive and more methodical.
You clarified in another post that you had won back in 2015. I have no clue the strength of engines back then (I imagine still very strong of course), but a decade of growth is a lot. They're completely insane nowadays.
You would lose every time, not even close.
* 100 games, to have some statistical relevance.
* One move per day, so that being tired is no disadvantage (engine can ponder all day).
* Human has access to endgame tablebases and opening databases, like the engine.
* Human can make notes and has a software like Chess Position Trainer, which can min max, like the engine.
If the human is a GM with Elo 2700+ I predict 25 draws and 5 wins for the human. The engine wins 70 games.> Rating Stockfish against a human scale, such as FIDE Elo, has become virtually impossible. The gap in strength is so large that a human player cannot secure the necessary draws or wins for an accurate Elo measurement.
There are many examples of top players playing Leela Knight Odds. And none of them even got remotely close to a decent record. Usually a few draws, and maybe a win. But almost entirely losses.
And that is with knight odds. Without that, zero chance.