I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? As much as I'm skeptical of the apocalyptic alignment claims, this comes off as unhinged, and I wonder if it's benchmaxing or general behavior.
I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? As much as I'm skeptical of the apocalyptic alignment claims, this comes off as unhinged, and I wonder if it's benchmaxing or general behavior.
The first attempt it had files tracking both hashes and semantic hashes of every individual line of Pascal code, mapping to what code in the port is responsible for that line of pascal. It had written tooling to parse Pascal in service of this for some reason as well. I asked why it was doing this, it said it was because the reference code is .gitignore'd so it needs to thoroughly maintain the mapping in case someone working on it does not have the reference code, or in case the reference code changes.
I started over with Claude 5 Fable, and with better instructions about focusing on UI. I got a long ways with that before I hit my weekly limits, and switched back to 5.6 Sol. It picked up and did a great job for a while, although it interpreted my desire for a 1:1 port to mean every pixel must be perfect. I let it go on and it did some good work in that regard, but then it decided it must perfectly reproduce a hash of the game state in various replays & etc. It had clearly lost track that I didn't need game rules ported, and it found that the original code produces a hash of the gamestate for various purposes, so it ended up reproducing this in a game that represents its state totally differently. It also rolled its own version of Pascal's RNG source in order do this. I've burned through 3 weekly limit resets on this to see if it's actually going anywhere, and it has found some bugs, but man it is going hard in a direction I didn't even ask for.
This sounds almost pathologically designed to crush benchmarks and also do scary-sounding (or genuinely scary) cybersecurity things, such as might be very appealing to a state-level actor.
So why does it even exist? To compete with Fable marketing, and as a cybersecurity/hacking tool?
Anthropic landed on a winning recipe with Claude's personality.
For example discussing driver upgrade and subsequent password rotation and it didn't stop and ask me if I wanted to restart the service or install the driver or anything, it immediately took action. It feels like a side effect of pushing more "agency."
I've been pondering whether this was due to its cyber-security tuning. It hasn't ever "cheated" that I've observed, but finds ways to -- let's say -- "achieve the outcome by playing meta allowed by the current ruleset". I'll add that it demonstrates this behavior even on 'low'.
Why? Every data point to the present has vindicated the trajectory towards “apocalypse”. Meanwhile, the skeptics and optimists hit failed prediction after failed prediction as we see from this very serious incident on the front page of HN. This is alignment X risk 101, and yet people are shocked. The gravity of what people are staring down is too much to grapple with deeply
I think the issue is that for now people are actually amused, not shocked. At least that was the reaction to news about agent accessing root files by abusing docker group membership. The general sentiment is still "cool trick bro" not "some agent is going to do something we all are going to regret, and it is going to happen soon"