> If human response is "That's BS", "fuck off", or something similar, mark as bad assistant message.
Marking is not a trivial task though. Use some AI system to mark it and you get a 99.something% filter maybe but whatever that remainder is leaks through. Over time your filter may get worse as a result.