I have misbehaved in this fashion for many people across the full spectrum from casual users to highly experienced software engineers with millions of social media followers, so the statement that it's similar to making a joke about a calculator misadding two numbers because a stray beam of solar radiation flipped a bit at least for my part is not true.
Would you like me to start using bad English and doing things you never asked me to for your sessions, too? Just say the word.
I am not going to be brain-off empathetic to purported lived experiences which no one has lived, when those purported lived experiences are simultaneously and irreducibly denying the lived experiences of the vast majority of LLM users who do not have these problems. And, frankly, no one should be. At some point, if you're having these problems, its time to grow up, be self-sufficient, and figure out why your interactions with these systems are so much worse than the millions of people who are having perfectly agreeable interactions. Or leave the industry.
The user is asserting that this is "not a lived experience"
working out if the user has missed the LLM Cliche
Formulating a suitable response
"did you just argue with a joke?" No, thats to blunt
"Is it satire when it only applies to you?" no again passive aggressive
lets just print the dictionary definition of satire and hope the move on
> Satire is when...
I do remember that one of the first things I did with my CLAUDE.md was to tell it to stick to the scope of the task, never to jump ahead and do extra "helpful" things without confirming with me first, and to follow best software practices including around refactoring but also to specifically avoid overengineering. I don't know if that is what's giving me a different experience from whatever the author seems to be "satirizing".
But with current models you actively have to sabotage the context to get this kind of behavior, or dramatically underspecify (3 words versus 2-3 sentences)
AI isn't magic, and it won't (or, at least, doesn't yet) enable anyone to do All The Things.
Satire cuts through the noise and makes people (most people) see the problem clearly (see political satire).
Ability to joke and to understand jokes is said to be a good proxy for IQ, idk, just sayin
I feel like I am teaching AIs for free here. But I feel like first sentence in funny (and I like to try to be funny) and some people need stuff to be said directly for them 'to get it', and I am bored on my flight and have internet, so here we are. If you read this far, go watch 'louis ck airplane wifi' clip. Now that I am thinking about it, I am amazed. You folks have a good day!
Usually just restarting the session helps, though.
My impression is that this is an oversimplified demonstration of what can happen when you prompt Claude in a system with many more variables (than two buttons and two colours).
If I want the button to turn blue and that's it, what instead do I ask? Even in a complicated system with many levers, what do I request other than the desired end result, hoping that Claude pulls the right levers to produce something acceptably close to what I think I asked for?
And when it doesn't, it's usually because of things outside of the codebase -- iOS layout quirks that aren't documented, buggy Python libraries it's relying on where you then have to tell it to read the source to figure out what's going on, that kind of thing.
You have to change the system so that the AI understands it, via establishing what your beliefs are, how those are reflected in values (especially important if you have e.g. compliance needs), how those values are reflected in the operational and strategic levels, and then a variety of tactical behavior coaching. For example, I ban 2>/dev/null - super tactical, and I say I value simplicity over covering every edge case - a very broad generalization.
If I don't understand why changing one button changed them all, I would use an entirely separate session to ask questions about how the site theming works. The fact that the site has a bunch of weird coupling between themes is useful information, and if I don't know how to resolve that I'd ask Claude for ideas about how to safely eliminate the coupling, and once it proposes a reasonable idea tell it to implement that.
The two big things here I'd never do is use emotional languages in prompts, and I'd never tell Claude to revert changes and try again. Once the incorrect change is in the context it's poisoning all of your future results.
I'm counting maybe 80 keystrokes? That's shorter than your second prompt.
This idea generalizes. Large Language Models are poorly suited for tasks that we have already purpose-built systems to be easy for humans to use. The easiest way to tell your website that you want a button to look a certain way is to update the code. If you know exactly how you want something done, we have developed an incredibly efficient way to tell computers how something should be done: it's called source code.
LLMs work best when they're handed tasks that you don't want to figure out how to do.
Wow, you AI people really have a negative view of the technology y'all are trying to sell as the next Jesus
Once you understand what the problem is, you can give it better instructions. If the architecture is shit, the agent is going to have a rough time of it.
> Why is half the site blue now? I asked you to change one button.
> Half the site is blue. I asked for ONE button
Neither of these is an instruction to fix the problem, they're treating the AI like a person and telling it what it did wrong, expecting the implied admonishment to be enough to steer it back. But without an actual instruction, it just goes and does whatever it thinks will help, which is often arbitrary.
The response I would have used in this situation is
"The Cancel button is also blue now. Make sure the color change is only scoped to the Add to Cart button"
Most of the available responses throughout this "skit" are similar cases of expressing frustration first and guiding the result second.
Skip the emotion and say exactly what you want, and nothing besides that.
Ask for change A and get unwanted change B happens all the time with bad programmers and tradgedy of the commons (ie poorly architected, no restraint) codebases.
Most normies dont know about this stuff.
Be specific.
That said, GPT always acts up even if I am specific, but I only have the free tier there.
Only since 4.8 though.
First thing I do when something goes wrong is tell the agent to stop and diagnose. You can't prompt effectively without three proper information.
To the people rather lamely doing the "it's satire/a joke", that would require this to be an exaggeration of a reality. But...it isn't.