You work around it with post-processing and retries. But it’s still a bit brittle given how much stuff happens downstream without supervision.
You work around it with post-processing and retries. But it’s still a bit brittle given how much stuff happens downstream without supervision.
Out of curiosity- do those orgs not find the loss of generality that comes from custom models to be an issue? e.g. vs using Llama or Mistral or some other open model?
Might not need JSON but whatever format it outputs, it needs to be reliable.
I was parsing a document recently, 10-ish questions for 1 document, would make things expensive.
Might be what’s needed but not ideal.
Seems like it would be universally useful.
If you're using anything less you should have a grammar that enforces exactly what tokens are allowed to be output. Fine Tuning can help too in case you're worried about the effects of constraining the generation, but in my experience it's not really a thing