For what it's worth, I think about the "engineering" in "Prompt/Context Engineering" almost more akin to how it's used in "Social Engineering". You are influencing the model to produce a desirable result.
For what it's worth, I think about the "engineering" in "Prompt/Context Engineering" almost more akin to how it's used in "Social Engineering". You are influencing the model to produce a desirable result.
And here you're assuming I'm (the author) not doing that.
Just the fact that the US wakes up and the available compute goes down affects model output significantly more than any magical prompting.
I've literally had the same prompt on the same code produce exactly two different results (proper magic and complete broken hallucination).
I've had one-line "do it"s one-shot complex problems, and I've had detailed precise instructions completely ignored producing horrendous code (and vice versa).
Until you have a way to measure your "should be obviously", it's nothing but wishful thinking.
For example Simon Willison (@simonw) has them draw a SVG of a pelican on a bicycle[0], something that has never ever been photographed ever, nor is it something the LLM can have a reference on, so it has to figure it out.
I have similar things, but not as visual. I know what the result should be and what it should look like. Then I compare and contrast.
Mistral.ai, for example, is by far the fastest (hardware clearly overprovisioned compared to load), but it also produces utter bullshit and hallucinations. With ChatGPT and Claude I can kinda feel the results slowing down or getting worse when USAians are awake and hogging the resources (They're either throttling or just plain using a shittier model unrer load). Deepseek even has a cheaper API price for off-hours.