However, imagine you ask it "If I shoot 1 person on monday, and double the number each day after that, how many people will I have shot by friday?".
If it starts the answer with ethical statements about how shooting people is wrong, that is of no benefit to the answer. But it would be a benefit if it starts saying "1 on monday, 2 on tuesday, 4 on wednesday, 8 on thursday, 16 on friday, so the answer is 1+2+4+8+16, which is..."
The AI does not think. It does not work like us, and so the causal chains you want to follow are not necessarily meaningful to it.
Ignoring caches+optimisations, a transformer model takes as input a string of words and generates one more word. No other internal state is stored or used for the next word apart from the previous words.
This is a rather contrived example, but the "mind" of an AI is different our own. We think inside of our brains and express that in words. We can substitute words without substituting the intent behind them. The AI can't. The words are the literal computation. Different words, different intent.
It also now has a lot of useless cruft I have to scan to get to what I want.
That's not really a great assumption. Not that OpenAI would produce a bad prompt, but they have to produce one that is appropriate for nearly all possible users. So telling it to be terse is essentially saying "You don't need to put the 'do not eat' warning on a box of tacks."
Also, a lot of these comments are not just about terseness, e.g. many request step-by-step, chain-of-thought style reasoning. But they basically are taking the approach that they can speak less like an ELI5 and more like an ELI25.
The main benefit of asking for terseness in your preferences is that it significantly reduces pleasantries etc. (Not that I want it completely dry and robotic, but it just waffles too much out of the box.)
Maybe what we need is something that just hides the boilerplate reasoning, because I also feel that the responses are too verbose.
The big problem I had earlier on, especially when doing code related chats, would be be it printing out all source code in every message and almost instantly forgetting what the original topic was.
What if I just ask it for a terse summary at the end? Maybe I’ll get the best of both worlds.
We tried the alternative, and it's less productive.
At some point, there is the theory and practice.
Since LLM output are anything but an exact science from the users perspective, trials and errors are what's up.
You can state all day long how it works internally and how people should use it, but people I've not waited for you, they used it intensively, for million of hours.
And they know.