On the topic of JSON output from these models, someone has added context free grammars to llama.cpp. This enforces that the output matches the grammar, effectively zeroing the probability of the next token not conforming to it.
https://twitter.com/GrantSlatton/status/1657559506069463040
https://github.com/grantslatton/llama.cpp/commit/007e26a99d4...
It's so obvious, it's genius.