Like you’d train the model to give you the most accurate response based on your current problem space, but when they changes, you’d have to retrain on what you’re working on but by that stage it’s already out of date ?
I doubt that.
I agree that in the bigger picture this doesn't matter, but it's technically true that cleaning the data in some way would help.
A related project is TinyStories where they try to use good data for unlocking the LLM cognitive capabilities without requiring as many parameters or exaflops. Again, there is obviously a limit to this, and maybe the effort is better spent on just getting even more gigantic dataset instead of nitpicking the useless or redundant data in the dataset.
Yes, the correct thing to do is get more data. Much more.
It matters a little bit, in a quantitative but not qualitative way. Probably with good data cleaning you could get as high quality result with only one pebibyte of data if it normally needs two pebibytes. If training time is proportional to dataset size then maybe it takes three months instead of six months to train. Maybe it would save hundreds of millions or a billion dollars which I guess would matter to someone. It probably wouldn't matter qualitatively though.
Or maybe clarification of people publicly saying they are going to ignore that restriction.
More data will fix variance problems, but not bias.