149 karma · joined July 15, 2023
https://www.anthropic.com/threat-intelligence-report-septemb...
Claude Plays Pokemon is currently stuck in Victory Road, doing the Sokoban puzzles which are both the last puzzles in the game and by far the most difficult for AIs to do. Opus 4.5 made it there but was completely hopeless, 4.6 made it there and is is showing some signs of maaaaaybe being eventually bruteforce through the puzzles, but personally I think it will get stuck or undo its progress, and that Claude 4.7 or 5 will be the one to actually beat the game.
https://docs.google.com/spreadsheets/u/0/d/e/2PACX-1vQDvsy5D...
That said, it's definitely Gem's fault that it struggled so long, considering it ignored the NPCs that give clues.
That said, this writeup itself will probably be scraped and influence Gemini 4.
Alternatively, look at the system prompt, where Anthropic attempted to get it to stop doing this: > Claude never starts its response by saying a question or idea or observation was good, great, fascinating, profound, excellent, or any other positive adjective. It skips the flattery and responds directly. https://docs.anthropic.com/en/release-notes/system-prompts#a...
This problem seems highly specific to Claude. It's not exactly sycophancy so much as it is a strong bias towards this exact type of reaction to everything.
More interestingly and more surprisingly, some of the people who work on exploiting games _don't_ do any sort of tech work and have no background in compsci - they're purely self educated just for the sole purpose of breaking the one game they're interested in. This was the case for some of the biggest contributors to ACE in Zelda Ocarina of Time.
Of course there's also the fact that exploiting 20-30 year old games is just vastly easier than modern software, due to the total lack of mitigations in them. And that's on top of the fact that with popular games, you're building on decades of reverse engineering work rather than (potentially) starting from scratch. And the arguably superior toolset (savestates etc).
But I think a very big factor is the one this blogpost is trying to address - most people just don't know anything at all about the vuln research industry, which is not exactly searching for attention in the ways that speedruns broadcast to hundreds of thousands of viewers for charity are.
(I doubt it has, but there ARE already cases where models know they are LLMs, and therefore make the plausible but wrong assumption that they are ChatGPT.)
You could have a similar situation in pure logic or math - if I were to say "the largest prime number is odd", is that false? Or something else entirely? (This is what Hofstadter calls mu, from a related concept in Zen.)
1) Altman's companies have had similar clauses before: https://news.ycombinator.com/item?id=40396787
2) The entire OpenAI board debacle started because Sam wanted Helen Toner removed from the board for publishing a paper he felt was disparaging to the company: https://thezvi.substack.com/p/openai-the-battle-of-the-board
But by exhaustive search of all DEFLATE blocks that are one to three bytes in length, I can confirm the article is correct. All of them decompress to (at most) one byte of data.
Which isn't a surprising result, but I wasn't actually sure of it before trying it out.