I tend to do this with GPT-4 even on the context window in default ChatGPT (or more often I bookend it with instructions). I find it pays off at even 1000 tokens.
I tend to do this with GPT-4 even on the context window in default ChatGPT (or more often I bookend it with instructions). I find it pays off at even 1000 tokens.
then I ask to generate. it's very powerful, as it removes the preamble and other chitchat from the response, and empower the system message over what's in the user message.
example: https://i.imgur.com/7fF0CZm.png?maxwidth=123456789&fidelity=... here the first agent message is the one conditioning the answer beginning, and I only generate the second agent.
(sorry mobile user imgur may return a low res unreadable image idk what's the alternative in 2023)
From what you're saying, it sounds like there is some kind of recency bias in these models.
Transformers predict the next token.
If your question is at the end of the prompt, the start of an answer is a more likely next token than if the question is at the beginning of the prompt followed by a ton of other relevant, but non-question-forming tokens.
Still, if you had to put the question at the beginning of your prompt, a transformer is more likely to give an answer than an RNN.
Genuine question; it would be interesting if some other mechanism was at play here.
I don't know about the super large contexts but you can also just make the text data clearly delimited instead of putting the query at the end, so that "predict the next token" isn't fighting the instruction-following training
DATA DATA
DATA DATA ...
-- boundary --
QUERY
or the arguably equivalent: QUOTED TEXT
-- boundary --
REPLY / COMMENTARY
The inverse shape is also common: INSTRUCTIONS
-- boundary --
DATA / TEXT ON WHICH TO WORK
For example, most exercise lists and test books are written like that.The somewhat less frequent patterns are more random mix of:
WHAT
-- boundary --
ON WHAT
-- boundary --
WHAT ELSE
-- boundary --
ON WHAT ELSE
-- boundary --
(...)
CLOSING REMARKS
Most of my HN comments are structured like that, for example. Including this one.Boundary here can take many forms. Extra newlines, --, ``` blocks ```, > - prefixed text, and lists (both OL and UL) are all common methods used to structure text, and are seen both in training data and in inference. We know LLM picks up on those structure markets at high-frequency level (e.g. using extra newlines or -- lines to separate distinct blocks seems effective). But I imagine it also picks up on the low-frequency patterns, which is why payload followed by, or bracketed with, instructions is something it "knows" how to process, whereas if you use less common structuring patterns, you're more likely to confuse the LLM.