This is a cute idea and it looks like it should work, but I could see this getting expensive with larger models and input prompts. Probably not a fix for all scenarios.
This is a cute idea and it looks like it should work, but I could see this getting expensive with larger models and input prompts. Probably not a fix for all scenarios.
I imagine closing the loop (using the TS compiler to restrict token output weights) is in the works, though it's probably not totally trivial. You'd need:
* An incremental TS compiler that could report "valid" or "valid prefix" (ie, valid as long as the next token is not EOF)
* The ability to backtrack the model
Idk how hard either one piece is.
Eg, given even the type:
{"aLongerKey": "value"}
The generation prefix: {"a
would by your algorithm produce the following invalid output: {"a}Here’s an example of one of my implementations of logit bias.
https://github.com/ShelbyJenkins/shelby-as-a-service/blob/74...
There's also a good assumption that models will be improving structured output as the market is demanding it.