My hypothesis here is that due to RLFH, there's likely some implicit learning that tangentially related content is better than no content.
Given that, you'd likely still get better results with your schema being:
"string | null" so the LLM can output a null instead of "" since there is probably not as much training data that gives "" high log prob values.
But we're looking forward to evaluating the functions call, and seeing what the metrics show!
https://letscooktime.com/Blog/ai,/machine/learning,/chatgpt,...
Hopefully this saves you some time!
Also looking to integrate the new function feature and now already got some learnings out of the post without even starting to code.
I had a schema with a string enum property to categorise some inputs. One of the category names was "media/other" or something to that effect. Sometimes the output would stop at just media even though it wasn't a valid option in the schema.