https://developers.openai.com/api/docs/guides/structured-out...
Hallucinations in tool call results _do still exist_, but basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.
That all goes out the window when a string field has a “hidden” schema in that only particular strings are valid, but that restriction isn’t in the JSON schema. I have had failures when I want a field to be specifically formatted Markdown or something.
We’ll probably see something that handles this better in the future like jsonnet or Cue or dhall that has some execution capabilities so you can write a custom validator beyond what JSON schema supports.
Basically no one is doing that since 2024, have you been living under the rock? Read about constrained decoding.
> I have had failures when I want a field to be specifically formatted Markdown or something.
And Jev has an advantage here because it can’t generate anything?
Well, that's essential what happens, but on the "backend" side still, so it's significantly more efficient. Also, model providers can play with e.g. how likely are they to accept a valid/invalid character, so for example the first token may only be [, { or ".
As for what's not in the schema, that's an orthogonal question.
I should also add that the source comes from one of about 400 possible places and in a variety of messed up formats, it's the raw feed from a news scraper...