If you tell it "Return what this says about sharks or nothing if it does not mention them", it will mess up.
User text: "Blah blah ... Sharks ... Surfing ..." Instruction: Return an JSON object containing an array of all sentences in the user text which mention sharks directly or by implication. Response: {"list_of_shark_related_sentences": [
Stop token: ']}'
It'll try to complete the JSON response and it'll try to end it by closing the array and object as shown in the stop token. This severely limits rambling, and if it does add a spurious field it'll (usually) still be valid JSON and you can usually just ignore the unwanted field.
wrt OpenAI, text-davinci-003 handles this well, the other models not so much.
Funnily enough, there is a certain propensity for it to output round numbers (50, 100, etc.) so I have to ask it not to do this and provide examples ("like 27, 63, or 4"). Now that I think about it I should probably randomize those.
My hypothesis here is that due to RLFH, there's likely some implicit learning that tangentially related content is better than no content.
Given that, you'd likely still get better results with your schema being:
"string | null" so the LLM can output a null instead of "" since there is probably not as much training data that gives "" high log prob values.
But we're looking forward to evaluating the functions call, and seeing what the metrics show!
https://letscooktime.com/Blog/ai,/machine/learning,/chatgpt,...
Hopefully this saves you some time!
Also looking to integrate the new function feature and now already got some learnings out of the post without even starting to code.
I had a schema with a string enum property to categorise some inputs. One of the category names was "media/other" or something to that effect. Sometimes the output would stop at just media even though it wasn't a valid option in the schema.
Basically, give the LLM a schema that is loose enough for the LLM to expand where it feels expansion is needed. Saying always "return a number" is super limiting if the LLM has figured out you need a range instead. Saying "always populate this field" is silly because sometimes the field doesn't need to be populated.