https://lifearchitect.ai/models-table/
Love those GPQA scores hovering around 5% when chance (on 4-way multi-choice) would have got them 25%!
Love those GPQA scores hovering around 5% when chance (on 4-way multi-choice) would have got them 25%!
If the clock is running faster than regular time, it will at point catch up to regular time and thus be correct for a split second. If the clock is slower than regular time, regular time will catch up to the clock and the clock will be right for a split second.
Somehow this thing manages to accumulate an error of ~15 minutes in a month.
Could be right within 15 min accuracy in the appropriate timezone. And such a mechanism can be corrected for in the postprocessing step.
Procedural error in testing perhaps? I'm not familiar with the methodology for GPQA.