Looks interesting! How would you say it compares to Microsoft's TypeChat (beyond the obvious Python/TypeScript difference)?
https://microsoft.github.io/TypeChat/blog/introducing-typech...
https://microsoft.github.io/TypeChat/blog/introducing-typech...
Our method on the other guarantees that the output will follow the specs of the JSON schema. No need to call the LLM several times.
Guidance (and this project?): Let's not even bother with trying to convince the model; instead, we'll only sample from the set of tokens that are guaranteed to be correct for the grammar we want to emit.