Having read the pdf I don't think these models were post-trained, so how do we explain the questions in B)?
And if indeed there's no post-training and authors expected exploration of recipes to come from the context window.... I think that's way too short for RL-style improvement.
In short, I don't understand how they could've tested those models with post training, and without post training they all did unbelievably well.
If the authors read this: can you give us an idea how many API query and API pairs fit within the context window, on average? Follow up, do you get better results if you abbreviate the API call names, so that more response pairs fit within one context window?