How about XLNet which cost something like $30k-60k to train [1]? GPT-2 may have been around the same [2] is estimated around the same, while thankfully BERT only costs about $7k[3], unless of course you're going to do any new hyperparameter tuning on their models which you of course will do on your own model. Who cares about apples-to-apples comparisons?
We're not talking about spending an extra couple hours and a little money on updated replication. We're talking about an immediate overhead of tens to hundreds of thousands of dollars per new paper.
Tasks are updated over time already to take issues into account, but not continuously as far as I know.
[0] https://www.wired.com/story/deepminds-losses-future-artifici...
[1] https://twitter.com/jekbradbury/status/1143397614093651969
[2] https://news.ycombinator.com/item?id=19402666
[3] https://syncedreview.com/2019/06/27/the-staggering-cost-of-t...