And here you're assuming I'm (the author) not doing that.
Just the fact that the US wakes up and the available compute goes down affects model output significantly more than any magical prompting.
I've literally had the same prompt on the same code produce exactly two different results (proper magic and complete broken hallucination).
I've had one-line "do it"s one-shot complex problems, and I've had detailed precise instructions completely ignored producing horrendous code (and vice versa).
Until you have a way to measure your "should be obviously", it's nothing but wishful thinking.