Is the lack of training data the only thing preventing this approach from being applied to both positive and negative prompts together?
What size data set is actually needed? Does it need to be machine generated or can you get away with something smaller, perhaps crowdsourced?