Guidance: A guidance language for controlling large language models
github.com
github.com
I dug into this the other day, and just about figured out how the old text-davinci-003 version works.
When it runs against a text completion model (like text-davinci-003) the trick seems to be that it breaks your overall Mustache-templated program up into a sequence of prompts.
These are executed one at a time. Some of them will be open ended, but some of them will include restrictions based on the rules that you laid out.
So you might have a completion prompt that asks for a maximum of 1 token and uses the logit_bias argument to ensure that the returned value can only come from a specific set of tokens. That's how you would answer a piece in the program that says "next should be just the sequence 'true' or 'false'" for example.
What I don't yet understand is how it works against non-completion models. There are open issues complaining about broken examples using it with gpt-3.5-turbo for example.
And how does it work with models other than the OpenAI ones?
This is how we implemented it anyhow, with some more parameters to control how that all works (and the LLM params) at each "pause" point. The _neat_ part for us was that a template helper could make use of the partially generated content. Hadn't thought about that before for a templating engine, but was trivial to implement in the end
Each expression becomes a new inference request, so it's not a single inference pass. Because each subsequent pass includes the previously inferenced text, the LLM ends up doing a lot of prefill and less decode. You only decode as much as you actually inference, the repeated passes only end up costing more in prefill (which tend to be much faster tok/s).
To work with chat tuned instruction models, you can basically still treat it as a completion model. I provide the previously completed inference text as a partially completed assistant response, e.g. with llama 2 it goes after [/INST]. You can add a bit of instruction for each inference expression which gets added to the [INST]. This approach lets you start off the inference with `{ "someField": "` for example to guarantee (at least the start of) a json response and allow you to add a little bit of instruction or context just for that field.
I didn't even try with openai api's since afaict you can't provide a partial assistant response for it to continue from. Even if you were to request a single token at a time and use logit_bias for biased sampling, I don't see how you can get it to continue a partially completed inference.
I find guidance to be fantastic for doing complicated prompting. I haven't used the 'controlling' the output feature as much as used it for chain prompting. Ask to come up with answers to a prompt N times, then discuss pros and cons of each answer, then make a new answer based on the best parts of the output. Stuff like that.
https://blog.simonfarshid.com/native-json-output-from-gpt-4
(it works perfectly with GPT-3.5 as well)
I would not expect it to make a difference in your current applications. Getting JSON is all about the model, training, and prompt, in that order
If you are looking for low-hanging fruit to improve your JSON responses from LLMs, fine-tuning will likely get you the most bang for your buck. Start from a coding model like codellama, code-bison, or starcoder
everyone is doing this, it's just part of the pipeline, certainly nothing innovative on that front in guidance
The mostly widely accessible form of this is probably BNF grammar biasing in llama.cpp: https://github.com/ggerganov/llama.cpp/blob/master/grammars/...
Is that needed in the LLM space yet? I'm just not convinced the abstraction pays for itself in reduced cognitive load, or at least not yet, but very happy to be convinced otherwise.
It's obviously extremely valuable if you're doing anything with the LLM output other than displaying it as a block of text to the user, or if you care about the output format at all.
LangChain: I found having a framework useful to ramp up people without prior LLM exposure, in an open-ended experimental space. The library covers many usecases and gets people thinking. But honestly their documentation is somewhat lacking for that purpose (stale text, shallow examples). Personally, coming from a search background I was able to DIY semantic RAG in the time it took to figure out how to do the same thing in LangChain.
You can get the same thing with Go text/templates by adding chat function(s) as custom a helper: https://github.com/hofstadter-io/hof/blob/_dev/lib/templates...
As a developer of these things, I don't get why they want to put so much effort into the mundane parts rather than focusing on the interesting parts. These things are mostly just the same as any other workflow or API call: https://github.com/hofstadter-io/hof/blob/_dev/flow/chat/cmd... (unless you get into the python and (i.e.) start messing with the logits or token probabilities)
I look forward to when we have something that can run any LLM without compatibility issues, can expose APIs etc and has a robust plugin or augmentation system.
There are many projects like these I'm tracking, but they all kinda cool off after the initial prototype and have thus many quirks and limitations
So far the only one that I could reliably use was llamacpp grammars, and those are fairly slow
How often does a project need to release to not be considered dead? It's only been 10 weeks, in the summer, at the peak of vacation time
Look at the most recent commits, they are setting up new governance, which likely took more than 1o weeks to work through the bureaucracy of Mircosoft
I thought guidance was smart, but LMQL seems brilliant as it merges pythonic constructions with LLMs (I think it may be an outright superset of python with LLM functionalities?)
It's predicated off a paper as well : https://arxiv.org/pdf/2212.06094
I wonder what other folks are building on this sort of workflow? I've been playing around with it and trying to figure out interesting applications that weren't possible before.
Somehow, the VCs and investors made us think it was cool to be working for them rather than our users
I haven't taken time to compare between the different releases but if anyone is having the same type of issues, I recommend downgrading even if it might mean less features.
https://github.com/microsoft/guidance redirects to https://github.com/guidance-ai/guidance now.
Do you have similar capabilities in your LLM project?
Notes on grammars here: https://til.simonwillison.net/llms/llama-cpp-python-grammars
Seeing how far prompt engineering can get you on this from (specifying the grammar as a one/few shot) does pretty well, it seems like something that could be better handled at training/refining time? My general feeling is that it is inserting a specific and complicated cog, and probably has to be tweaked for each model (as they all have their little quirks)
Curious what you think having been working directly on these things?
From my experience, you can get an LLM to follow a "grammar" for pretty much anything, without it being an actual grammar spec in one of the many formats. You can pretty much make it up. Here's an example of us getting CUE out of a model by giving "tricking" the LLM to generate JSON with less syntax (a subset of both CUE and JSON). Bonus, fewer quotes and commas meant fewer tokens. We turn this into JSON afterwards, works surprisingly well
https://github.com/hofstadter-io/hof/blob/_dev/flow/chat/pro...
I really like how grammars offer a realistic path to getting completely dependable formatted output from these models.
I haven't pushed on codellama2 much yet, but my initial experiments, it did not really output anything extra, and my prompt became a one-liner compared to the really long instructions I had to give chatgpt for controlling output. Shows how far you can get with a purpose trained model
fine-tuning is important to getting more consistent output, none of the smaller models (open-source sized) are going to get there with just few-shot. Sounds like the grammar logit influencer is a low-cost/effort way to constrain output without the fine-tuning cycle. I can imagine they might be better together, but my hunch is that fine-tuning will still dominate the improvements and consistency. If you don't have the training data, that is a very good reason to use this technique too