1) Any particular reasoning behind estimating OpenAI’s margins are 60%?
2) How much does human preference diverge from benchmark scores in your experience?
3) Do woodpeckers stop attacking houses when it’s winter in Alberta?
2) How much does human preference diverge from benchmark scores in your experience?
3) Do woodpeckers stop attacking houses when it’s winter in Alberta?
2) It generally tracks pretty well unless the model is gaming the metric (training on the test set, overfit to the specific source of data, etc). The relative rankings will typically match in both.
3) alas, not with the mild winter North America’s having. They only stop below -5C or so. I am lucky though. The woodpecker stopped attacking my house and started attacking my neighbor’s. Even worse, it used to be a downy woodpecker,and it’s now been replaced by a pileated one (think: Woody).