METR's time-horizon of coding tasks does not mean what you think it means
killerstorm.github.io
killerstorm.github.io
METR considers this "raw baseline" largely irrelevant as it might be affected by people getting bored / not paid enough, etc. But they admit this introduces a bias which makes reported numbers less relevant for human-vs-AI comparison.