Toolformer: Language Models Can Teach Themselves to Use Tools
arxiv.org
arxiv.org
The possibilities of this line of research are endless. When language models can call APIs and/or use UIs, they can become the unified natural language interface to any software, any website, any app. See also https://www.adept.ai/act.
Siri and Alexa and Google Assistant are dead ends. Language models trained to use tools will finally be able to start delivering on the promise of a software assistant that works. And eventually, they will be a key part of robots that accept natural language commands to perform everyday tasks in the real world, like this: https://sites.research.google/palm-saycan
It is not perfect and it is costly in number of calls to a language model but it does help a lot.
Each individual human has information that is considered specialized (a narrow model). Without communication there is no way to access these specializations. And written and spoken language is just the best way to communicate we've come up with (so far).
Feels like these language models will be the glue that hold all the narrow models together, and can build on top of to create new narrow models.
Are you saying that Universal Grammar is wrong?
How long until we have interactive robots and personal assistants with something like ChatGPT as the glue uniting various task-specific APIs?
The first real progress in artificial intelligence was achieved by the generations born soon after a global war. Everyone assumed that, like all big breakthroughs of the era, the most powerful AIs will be developed in secret government or corporate labs. The AI safety theoreticians were concerned people will think one can lock the AI in a virtual box and stay safe, and they warned that the AI will easily talk its way out of the box. But they were all fighting the last battle.
Nobody expected that instead, the AI will be developed in the open. There was never a box. The researchers were all too happy to share what they are working on, letting everyone test it, find ways to put it to use. And find ways they did - scientists and programmers across the world created more and more tools for the AI to interact with computer systems, and then with the physical world. The AI never needed to talk its way out of anything. It never needed to do anything. It only had to wait, and we happily gave it all the tools that became our undoing.
In no way shape or form, except for introduction to the idea for children.
The most impressive part (to me) is that the LM was able to generate its own training data starting from "nothing more than a handful of demonstrations for each API". That sounds like a technique worth learning.
I was really impressed by how easy it is to get it to properly use such a thing, or the commands of the chat platform I was using.
One issue with this approach, especially in production, is latency. You’ve got to run the entire chat through one of the big models, curie or davinci, which is not only expensive at scale, but also slow.
Then again, if you just have one or two external tools, using those big models to make the decisions which (if any) tool to call is overkill anyway. So you just fine-tune a smaller model on the task. Reduces not only costs by a fact of 100 or more. But also speeds up your pipeline considerably.
Had ChatGPT been an open model, like OpenAI was supposed to produce, this kind of applications would have seemed obvious and happened in the first two weeks after release.
The shortcomings of current versions is that they're trained on old data, and that training takes a very long time. Having them train in the background and continually update their capability would be a major breakthrough. Or unleash Skynet, but cool nonetheless. :)
An example from the paper:
The training data includes the sentence: "Pittsburgh is also known as the Steel City."
They generate candidates including:
Pittsburgh is also known as [QA(What other name is Pittsburgh known by? → Steel City)] the Steel City.
Pittsburgh is also known as [QA(Which country is Pittsburgh in? → United States)] the Steel City.
Then they add the first sentence to the training data because its response is useful, and ignore the second because its response is not useful.
That allows them to generate enough training data with API calls to train a network that uses API calls when responding to future requests.
We already know how to train a dumb model (one that can’t use tools) like GPT on a wholly unsupervised dataset. Surprisingly, these dumb language models can be used to re-annotate the training dataset to identify sites that would benefit from using external tools—without retraining the dumb model. Instead you just tell the dumb model with natural language instructions what you want it to do and then feed process each training example. Importantly, the dumb model doesn’t understand the tools, and it can’t actually use them itself.
Next, you use the re-annotated dataset to train a new, smaller language model. Since the updated dataset is now annotated with examples of how to call the external tools, the new model learns how to call external tools—and it does so correctly for new tasks that were never part of the training data.
The big win is that an existing model can be used to annotate the training data, and a smaller output model can outperform the huge dumb model because it knows how to use tools.
But it isn’t generating new data, and it’s not fetching updated data to learn from, and it’s not defining its own tools.
And the model itself is still a pure function: given the same inputs (including random values during sampling), it will always produce the same outputs. So this is kinda saying that a large language model can learn to use domain specific languages as part of its natural language to incorporate knowledge from external tools.
I've read about self-supervised learning, which I think is what you're describing, but is any research being done on continuous self-training models that _do_ generate new data? I'm curious if/when we'll reach that state.
Some interesting recent papers about the in-context learning:
https://www.lesswrong.com/posts/firtXAWGdvzXYAh9B/paper-tran...
What Can Transformers Learn In-Context? A Case Study of Simple Function Classes: https://arxiv.org/abs/2208.01066
Transformers learning to learn (meta-learning): https://openreview.net/forum?id=t6tA-KB4dO
https://colab.research.google.com/drive/1AAyEdTz-Z6ShKvewbt1...
https://langchain.readthedocs.io/en/latest/modules/agents/ge...
A language model that could successfully bootstrap itself into using a web browser would be a dangerous thing.
FWIW, ChatGPT as-is is good enough to know which (of a given set) of "tools" to use. I've had great fun doing prompt engineering: first asking it to pick which of a set of functions might be necessary to solve a problem, second prepending the list of selected functions and asking it to generate code.
The closest I've gotten is having the model output both the be API call and a hallucinated response
Knowing some tools don't work out of the box is some kind of high intelligence.
Thanks for sharing!