GPT-4-turbo produces shorter completions when it "thinks" its December vs. May
twitter.com
twitter.com
Over and over again people create some statistically significant way of measuring a difference in the length of responses.
Why the obsession over that metric? People in the twitter replies going on about how that shows it's getting worse.
For me, the shorter responses are the better ones. It's specifically in my personal instructions to reduce the length of responses. In saying that, sometimes I want longer. Then I specify it so.
It's "cool" that it has different response lengths seemingly based on its perceived date. It doesn't make it "better" or "worse"
Inspectable evaluation flow in ChainForge: https://chainforge.ai/play/?f=2yvqkpe1vpus8
The current date is part of the system prompt as far as I can tell. At least on ChatGPT you can ask for current date, not sure about API (EDIT: indeed, this is explained further in the tweet that OP links to[0], go there since that's where the actual info is).
Maybe the model learnt that humans in general tend to give shorter/lazier responses around this date (being trained on texts with dates in metadata, think forum posts, blogs, tweets, etc.) so it could very well be imitating trends in seasonal human behavior.
[0] https://nitter.net/RobLynch99/status/1734278713762549970
The underlying data is embedded with artifacts of behaviors/affects that reflect the underlying state of the person who wrote them.
Stanford famously studied this for reddit (notably a huge part of the commoncrawl dataset) and reported that it had non-trivial percentages of anti-social sentiments embedded in the text [1]
You can't filter out the biases embedded in the data because it is a functional artifact of what the people creating the data were communicating. Best you can do is put guardrails and censor it, but that's a neverending game you can't win.
Junk in Junk out still applies to LLMs
[1]https://hci.stanford.edu/publications/2022/Park_ContentModAu...
I wonder how we would test it
Testing “friendliness” or “holiday cheer” could be done via some sortof proxy…
you could prompt it to reply to generic pleasantries while role playing “as someone busy who is late” and see if the word choice changes. Or maybe ask it to be a judge for a crime of desperation (stealing toys for orphans?) and determine sentencing durations as a proxy for sympathy? I suspect that is too far from training data though.
So maybe a better approach would be to geolocate the prompt in Australia, for example: it's December in Melbourne, Australia...
Does it happen with April and June? Does the response length trend downward as the months approach December? Does it change its behavior if it believes its in a hemisphere where May is snowy?
If it were performing better on standardized testing then I would find the results compelling, but generating longer responses isn't even necessarily desirable.
Surely? Whence the certainty? If humans act this way, and the training data is mostly from humans, I would be surprised if this wasn't a real effect.
But yea, the data could be better (but the hypothesis is very plausible, considering the training data).
Alternately, "It is a bright sunny summer day, with just enough of a cool breeze to feel invigorating. You woke from a wonderful night's sleep and started your day with coffee, eggs, modafinil and adderall. As you feel your IQ, motivation and creative spirits reaching peak flow, please ..."
That's going to trigger outright LARPing, and you should not trust information integrity at that point. Following it up with cocaine and antidepressants just negates it, so you might as well leave both conditions out and spare the tokens and confusion-- it's TMI.
I don't see the value in trying to correct this, at any rate. If people normally send and receive short emails in December, forcing it to write an amphetamine-fueled novella is not going to be natural or well-received.
I guess I need to be careful about what I might "prescribe" to a mental model mostly reflecting "typicals"! What strange things LLM's have become.
Wild result. gpt-4-turbo over the API produces (statisticall
significant) shorter completions when it "thinks" its December
vs. when it thinks its May (as determined by the date in the
system prompt).
I took the same exact prompt over the API (a code completion
task asking to implement a machine learning task without
libraries).
I created two system prompts, one that told the API it was May
and another that it was December and then compared the
distributions.
For the May system prompt, mean = 4298
For the December system prompt, mean = 4086
N = 477 completions in each sample from May and December
t-test p < 2.28e-07
To reproduce this you can just vary the date number in the
system message. Would love to see if this reproduces for others.
This is part of the ongoing discussion about whether or not ChatGPT has got "lazier".OpenAI claimed they had made no model changes and were puzzled as to what was going on: https://twitter.com/ChatGPTapp/status/1732979491071549792
we've heard all your feedback about GPT4 getting
lazier! we haven't updated the model since Nov 11th,
and this certainly isn't intentional. model behavior
can be unpredictable, and we're looking into fixing
it
The theory was that maybe the system prompt they inject telling it the current date might be influencing it, because its training data showed people worked less hard in December.New evidence suggests that theory might actually hold up!
Imagine being an sociologist with the ability to train LLMs on very specific hyperplanes of data.
This feels like the original author is over anthropomorphizing LLMs, and expecting them to interpret prompts the way humans would, but it seems obvious that changing the prompt results in a slightly different context window, which results in a slightly different response distribution? Similarly, if you changed whether the bit about time of year was at the beginning or the end of the prompt, I would expect a statistically different distribution of response lengths.
Another perspective is one of phenomena that comes and goes based on the state of the environment and the progression of some underlying process, e.g. a plant that flowers and then goes dormant.
I know that personally, I experience very different levels of productivity and mental states based on the time of year. Spring brings a feeling of newness and possibility. Summer a desire to socialize and have fun in the sun. Winter is more contemplative and occasionally depressive.
I could definitely see my output as someone who writes things taking on different characteristics as the conditions around me ebb and flow.
As others have pointed out, I don’t think this variance is limited to a single notion of better/worse, but it definitely modulates my output.
> time to build. take a deep breath and just go.
It's people all the way down...