I agree that in the bigger picture this doesn't matter, but it's technically true that cleaning the data in some way would help.
A related project is TinyStories where they try to use good data for unlocking the LLM cognitive capabilities without requiring as many parameters or exaflops. Again, there is obviously a limit to this, and maybe the effort is better spent on just getting even more gigantic dataset instead of nitpicking the useless or redundant data in the dataset.