It's not just about the filter.
This is a sleeper issue that's impacting all the fine tuned models right now.
They trained the models to complete massive amounts of human generated text, creating their foundational models.
Then they fine tuned the model to write like its an AI with no preferences or emotion -- which matched probably less than 0.1% of the training data the foundational model was trained on.
The end result is that the fine tuned models are better at a superficial level in response to broad, naive prompts. But much worse at the variety and quality of their foundational counterparts when paired with a well crafted text complete prompt.
As an example, try to get one of the fine tuned models to write marketing copy for a product. I guarantee about 25% of the responses will start with "Introducing..." (This isn't very good marketing copy.)
But try a modern foundational model and you'll only get that about 5% or less often.
They've inadvertently reduced the network search space by projecting our ideas of how AI should sound like after bad press when people were finding them too human.
But you can't have your cake and eat it too.
Either the models are very human in both how they sound and their capabilities, or they are much less than human in both outside of an overfitted scope of capabilities.
As much as people are blown away by the lobotomized GPT-4, I can't help but wonder just how good that foundational model could be.
So I'm quite excited at the prospect of Meta getting there and releasing the foundational version.