My contention, which I covered in this video below here, is that due to the underlying statistical sampling problems inherent in RLHF transformers, LLM's perform poorly in edge cases, which, depending on the application or language, the margin of that edge can be super wide.
Here's a video I created about it: https://www.youtube.com/watch?v=GMmIol4mnLo
I didn't cover this yet but there are these things called, "scaling laws," which basically state the amount of raw text needed for a LLM with of a particular size of parameters. So my current mental model is that these, "laws," are really economic rules of thumb, like Moore's law is actually Moore's Rule of Thumb, and there is a huge expense in sampling clean data, hence the need for RLHF.
More about RLHF if not familiar with that term yet: https://huggingface.co/blog/rlhf