GPT4 is trained on mostly garbage optimised for SEO though.
I only ask it questions whose answers I can verify e.g. if I ask it how to do something in F#, a language I'm not very familiar with, I can easily confirm whefher the code does what I need it to or not.
You can pump as much SEO garbage out as you want, it doesn't change the value of LLMs to me in this context.
This was later fixed by updating the tokenizer, but it demonstrates the importance of clean datasets at all stages.