The prompt would basically be (but with less words. Around 10 is the best for consistency from my experiments. Much higher than that and it's hit or miss).
Rearrange (if necessary) the following words to form a sensible sentence. Don’t modify the words, or use other words.
The words are:
access
capabilities
doesn’t
done
exploring
general
GPT-4
have
have
in
interesting
its
it’s
of
public
really
researchers
see
since
terms
the
to
to
what
A successful completion would be.
Since the general public doesn't have access to GPT-4, it's really interesting to see what researchers have done in terms of exploring its capabilities
The number of permutations of the 24 words in the pre-scrambled sentence without taking into consideration duplicate words is 24 * 23 * 22 * ... * 3 * 2 * 1 = ~ 6.2e+23 = ~ 620,000,000,000,000,000,000,000. Taking into account duplicate words involves dividing that number by (2 * 2) = 4. It's possible that there are other permutations of those 24 words that are sensible sentences.
For a language model to consistently make sensible predictions, it quite simply has to be able to "look ahead".
When the probabilities for the candidate tokens for the first generated token were calculated, it seems likely that GPT-4 had calculated an internal representation of the entire sensible sentence, and elevated the probability of the first token based on that internal representation.
I don't know what else to call looking backward, looking forward and then producing output anything other than a state.
If you want to see what this kind of completion would look like without much or any regard for a sensible completion of the sentence then just ask 3.5.