mm, no, not unless you're doing some LIKE-specific optimization (and even then, I think you'd want "positive%").
So like, at the end of all the decoders, the model gives you an output vector; you multiply this by your embeddings to get your token probabilities, then you sample from them to choose a token.
Instead of sampling, you could just look at the probabilities for the tokens "positive" and "negative" and return whichever of those two is highest.