Probably not the point, but shouldn't you be able to choose to sample only the tokens for "positive" and "negative" (they're both one token!) instead of (or in addition to) needing to put a request for model to restrict its responses in the context?
I guess this is the SQL query you have in mind that uses the LIKE operator:
SELECT ChatGPT("Respond to the review with a solution to address the reviewer's concern", review)
FROM postgres_data.review_table
WHERE ChatGPT("Is the review positive or negative?", review) LIKE "%positive%"
AND location = “waffle house”;
From a query processing standpoint, both queries should have equivalent performance -- unless we build an index over the output of the ChatGPT query in EvaDB, in which case the former query would be faster than this one.So like, at the end of all the decoders, the model gives you an output vector; you multiply this by your embeddings to get your token probabilities, then you sample from them to choose a token.
Instead of sampling, you could just look at the probabilities for the tokens "positive" and "negative" and return whichever of those two is highest.
Doesn't this token sampling optimization require using a locally-running model like Llama?
I am presuming that OpenAI doesn't provide direct access to token probabilities in its API.