How does llama.cpp’s grammar adherence work?
Does it keep validating the predicted tokens and backtrack when it’s not valid?
Does it keep validating the predicted tokens and backtrack when it’s not valid?
There are contrived grammars you can give it that will make it use exponential memory, but in practice most real-world grammars aren't like this.