HNHacker News
TopNewBestAskShowJobs

doktor_pepsi

1 karma · joined October 8, 2026

submissionscomments
doktor_pepsi··on Claude Haiku 5.5
Appreciate providing this benchmark. I am curious, do you think that as new models (from the same company) are released and you repeat the benchmark, they adapt/extend their training data to include your dataset, thus polluting the benchmark results? Would you be doing anything to combat this?