1,106 karma · joined August 27, 2020
https://www.primeintellect.ai/blog/measuring-autonomous-rese...
Basically they do 8 runs trying to optimize to under 3.28 loss in the fewest training steps possible under time/token constraint. I dunno why 18 * 8 != 153 (it's 144)
https://github.com/derac/WeatherTray
I use Linux now, so you're on your own if there are issues. It might require some windows library to be installed but I don't recall. I ran it for a long while on win11.
Old servers may have been grandfathered, but mine was set to 1/1 and couldn't be reshaped.
I am taking all the measurements from the GPT-5.6 announcement blog and presenting them in a tabular format with relative scale against a chosen model/effort.
The json is on the page, it's all static, so feel free to take the data and do what you want with it.