At least with OpenAI, wouldn't it be better if under the hood it was using the new function call feature?
I imagine closing the loop (using the TS compiler to restrict token output weights) is in the works, though it's probably not totally trivial. You'd need:
* An incremental TS compiler that could report "valid" or "valid prefix" (ie, valid as long as the next token is not EOF)
* The ability to backtrack the model
Idk how hard either one piece is.
Eg, given even the type:
{"aLongerKey": "value"}
The generation prefix: {"a
would by your algorithm produce the following invalid output: {"a}