GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.
We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.
(I work at OpenAI.)
Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.
Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.
https://github.com/BiggerPockets/.github/blob/main/.github/w...
If it's a single task where it dropped from 50 to 25, it could be random variation (not saying it is, but it could be). If it's the mean over hundreds of tasks, that suggests a problem with either the eval code/harness or our API.
(it matters if they are independent or dependent)
Serial testing over time is much less reliable than side by side testing, and even when I do side by side testing, I try to look at multiple attempts per prompt. Seeing multiple per prompt helps me realize how much intrinsic variation there is. My brain always wants to see patterns even when there isn’t enough data to prove them.
Subscription plans may be subject to other regime, e.g. lowering the thinking budget when the API is under heavy load, etc.
Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.