Is this what happens when you let large-scale language models into the wild? It seems like as soon as you start modelling your responses to queries on real-world text data you're going to run into this quality control problem. What happens when bad actors start polluting public textual data sets to make these kinds of responses more common?