Can anyone comment on how well this does at coercing json output vs OpenAI function calling?
I would not expect it to make a difference in your current applications. Getting JSON is all about the model, training, and prompt, in that order
If you are looking for low-hanging fruit to improve your JSON responses from LLMs, fine-tuning will likely get you the most bang for your buck. Start from a coding model like codellama, code-bison, or starcoder
everyone is doing this, it's just part of the pipeline, certainly nothing innovative on that front in guidance
The mostly widely accessible form of this is probably BNF grammar biasing in llama.cpp: https://github.com/ggerganov/llama.cpp/blob/master/grammars/...