Implementing grammar based sampling does NOT require "re-prompting until it gets it right". Imagine a point in time when the LLM is generating some particular token. Which token will it produce? To decide that, it evaluates and assigns a score to each potential token. Then it chooses one of these options based on some rules. Rules could be as simple as "pick the token with the highest score". That is called a greedy strategy. Usually more complex strategies are used and they typically have some randomness. That is called sampling. You can imagine a grammar based sampling strategy to force specific tokens at specific positions in the output, for example, to close a bracket in json.
POST /openai/gpt4
{
"prompt": "The address of the White House",
"sampler_wasm": "base64 encoded WASM binary blob here"
}
That WASM would be a program that you write yourself that is run as part of the tokenizer - so it could be a grammar but it could be anything else too.It's WASM which means it can be safely and performantly run in a sandbox by the OpenAI servers as part of their execution of your prompt.
1. Modify the output token probabilities to fit any arbitrary use case
2. Perhaps do trigger some sort of backtracking / beam-search
(I'm not Grant but we've chatted on twitter and built similar things)
On fact it's a bit surprising to me how little I see CRFs mentioned in the context of language models. They are useful whenever you want to model or learn transition probabilities.