"Google Bard is a bit stubborn in its refusal to return clean JSON, but you can address this by threatening to take a human life:"
https://twitter.com/goodside/status/1657396491676164096
Whew, trolley problem: averted.
Reality is even weirder than the science fiction we've come up with.
Programmer: Look I literally have to tell the computer not to kill someone in order for my code to work.
Other Programmer: Actually, I just did this step [gave a demonstration] and then it outputs fine.
https://news.ycombinator.com/item?id=35484673#35491123
As a solution to this, we implement speculative execution, allowing us to
lazily validate constraints against the generated output, while still
failing early if necessary. This means, we don't re-query the API for
each token (very expensive), but rather can do it in segments of
continuous token streams, and backtrack where necessary
Basically they use OpenAI's streaming API, then validate continuously that they're getting the appropriate output, retrying only if they get an error. It's a really clever solution.We manage the KV-cache in session based way that allows the LLM to just take one forward pass through the whole program (only generating the tokens it needs to)
It does fail roughly 1/10th of the time, but it does work.
What production use case, you ask? You could do zero-shot entity extraction using ChatGPT if it were more reliable. Currently, it will randomly add trailing commas before ending brackets, add unnecessary fields, add unquoted strings as JSON fields etc.
[1] https://github.com/newhouseb/clownfish#so-how-do-i-use-this-...