Surprised it took them so long — llama.cpp got this feature 1.5 years ago (actually an even more general version of it that allows the user to provide any context free grammar, not just JSON schema)
Surprised it took them so long — llama.cpp got this feature 1.5 years ago (actually an even more general version of it that allows the user to provide any context free grammar, not just JSON schema)
> The model can fail to follow the schema if the model chooses to refuse an unsafe request. If it chooses to refuse, the return message will have the refusal boolean set to true to indicate this.
I'm not sure how they implemented that, maybe they've figured out a way to give the grammar a token or set of tokens that are always valid mid generation and indicate the model would rather not continue generating.
Right now JSON generation is one of the most reliable ways to get around refusals, and they managed not to introduce that weakness into their model
Is this just a schema validation layer on their end to avoid the round trip (and cost) of repeating the call?
The simplest algorithm for getting good quality output is to just always pick the highest probability token.
If you want more creativity, maybe you pick randomly among the top 5 highest probability tokens or something. There are a lot of methods.
All that grammar-constrained decoding does is zero out the probability of any token that would violate the grammar.
Does it keep validating the predicted tokens and backtrack when it’s not valid?
There are contrived grammars you can give it that will make it use exponential memory, but in practice most real-world grammars aren't like this.