Also, it isn’t comparing one 10-day period once and calling it done. The window rolls forward daily, and a change has to clear the pre-registered 99% threshold in two consecutive windows before it’s flagged.
Ten days isn’t sacred, though. Once there’s enough longitudinal data, one of the things I want to evaluate is whether that window length is actually well calibrated or should be changed in a future version.
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.
I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the temporary thrill of finding something marginally better than what you have.
_pushes model further_
"Ugh why is this model failing now even though I'm asking it to do harder things and also got sloppier with my prompts because I got used to it being able to figure stuff out"