ParentFull threadWanderPanda·Why don‘t these benchmarks judge the likelihood of the example answer? Just taking the MAP predictions seems like a waste of informationView on HN