ReLLM: Exact Structure for Large Language Model Completions
matt-rickard.com
matt-rickard.com
In the ideal case, such as when you ask for a json array, the first "[" token you enforce will have a relatively high probability and forcing it to go down that path will give you good results.
In the dangerous case the model doesn't have a good structure-compliant completion to your prompt and the regex you supply forces the model into a path of extremely low probability and you get trash results.
It seems like there's two clear paths:
- Allow the model to complete whatever it wants and then anneal the structure into compliance - Force the model into a compliant structure and then anneal the quality
I think both options can make sense in different cases.
One case I'm thinking about where the second option feels simpler is when you want to implement a boolean function using a language model. I'm imagining a probability distribution that looks like:
- (40%) the answer is true - (39%) True - (21%) False
In this case it seems significantly more straightforward to force the model into completing T or F. I guess you then run into the "dangerous case" where you have
- (40%) the answer is false - (39%) True - (21%) False
For now, this is a good way to make sure that I can parse the output reliably in the minimal amount of completions (instead of looping until conformant).
https://github.com/hellisotherpeople/constrained-text-genera...
I think LMQL is the best example I've seen of the "forcing the LLM to walk a certain path" technique. It's a DSL written in Python for Python though, so it kinda constrains it's utility.
A regex is a better fit for a different class of problems. You might implement a JSONformer/clownfish with this instead.
fn complete(input, masker):
completion = input
while (not done yet):
mask = masker(completion)
completion = complete(complete, mask)
end
return completion
end
is that accurate?You have to go one token at a time, otherwise the masking becomes combinatoric rather than linear (two tokens at a time -- need to generate all two token pairs, etc.).
But otherwise, that's what the code does! https://github.com/r2d4/rellm/blob/main/rellm/rellm.py#L21
In general, it seems like building a production-ready LLM-based application requires _a lot_ of these little tricks and methods to hammer results into shape before sending them downstream.