“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.
To recap:
1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.
2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.
3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".
Frankly... I'm more annoyed at the posters than the bot.
It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
Granted, this can probably be tuned for.