Ensuring a model never outputs copyrighted content is unimportant and tangential. It's irrelevant. You don't look for a way to make humans output no copyrighted content, you address each time they do case by case.
A model training being rendered fair use doesn't mean any of its output can be used for whatever regardless.