It would be great if we could have detailed QA evaluation to show this, but of course then the open source people would train their models on it as a fine-tuning datasaet.
I mean… the horror.
There's also already tools for conversation flows (which just means you prepend the conversation history to the prompt).
I'm not saying the performance is nearly as good, but the actual workflow does already exist and is massively improving. The interesting part to me is that this finetuning can be done in a couple few hours on a consumer gpu (4090).
langchain agents are a good starting implementation.
you can build your own prompt and get the ai to work by iself hallucinating tools, which may be cheaper to test out than going back and forth with an agent manager. not as accurate, but you can still extract useful work, i.e. https://i.imgur.com/AE4R3dR.png (gpt-35-turbo is traditionally failing this task completely, prompt get it to work at it)
these prompt all require the model to work off data within the prompt within the first shoot. model require a degree to introspection for that to work.