Inversion: Fast, Reliable Structured LLMs
rysana.com
rysana.com
https://github.com/guidance-ai/guidance
Same principle
My best guess is that they're using two approaches to get this running faster:
- structured generation techniques from sglang (https://github.com/sgl-project/sglang) that allow them to generate faster JSON (with look-ahead / pre-fill) with strong guarantees on the output (i.e. 100% reliable, without requiring any retries).
- distilling a gpt-3.5 turbo-esque model from GPT-4 JSON outputs, and using it in conjuction with above to give the additional performance boosts on inference.
It doesn't seem like they're deploying on any custom silicon, nor have they optimized GPU kernels to suggest that the speed ups came there.
When generating json, calling the llm when you already know that after
{ "brand": "Toyota"
Comes , "year":
Is a massive waste. If the data itself is constrained too you can skip most of that too! You'll go down from needing 20 calls to the llm to just three for a simple piece of data like { "brand": "Toyota", "year": 1995 }
If they combine these techniques with a model that's specifically trained for structured output, along with a novel inference-time pruning technique that they were talking about in the post I can definitely see them getting these kinds of inference speeds.I'm experimenting with a self hosted api that is fast enough to not even need a gpu for single user use cases (because the latency is good, but not batching). Once I'm done with the finishing touches I'll rent a GPU server for actual hosting.
sglang: https://arxiv.org/abs/2312.07104
The op also has something about compressing finite state machines, it's basically the same thing with slightly different details in the implementation.
Has anyone seen a good JSON library that can handle slightly broken JSON? e.g. trailing commas, unescaped newlines, etc.? I have not found a good one.
However, they are way easier to get started with using in context learning. Soon, they will be cheaper and probably faster enough too that training your own model will be a waste of time for 95% of use cases (probably higher because it will unlock use cases that wouldn’t break even with the old NLP approaches from a value perspective).
This is why I am tracking LLM structured outputs here:
https://github.com/imaurer/awesome-llm-json
And created an autocorrecting pydantic library that could be used for Named entity linking:
This is a small model optimized for retrieval and function calling. "Reasoning" makes an appearance in the title but no standard benchmarks of general ability, such as MMLU or HumanEval, are mentioned. No details about the training process and no access to the models other than via API.
Nice marketing, but looks empty. I can also make an LLM that runs 1000x faster than Mistral:
def complete(prompt): print('As an AI language model...')