Big error bars and METR people are saying the longer end of the benchmark are less accurate right now. I think they mean this is a lower bound!
METR currently simply runs out of tasks at 10-20h, and as a result you have a small N and lots of uncertainty there. (They fit a logistic to the discrete 0/1 results to get the thresholds you see in the graph.) They need new tasks, then we'll know better.