Gorilla: Large Language Model Connected with APIs
shishirpatil.github.io
shishirpatil.github.io
There is a finite number of public APIs (100K?) which keeps the problem manageable. IMO, adding support for custom/private APIs (something like OpenAI functions) will make this a very powerful tool.
Minor nitpick, but we still don't have a clear legal answer on whether this would be binding to people who didn't sign that agreement, because we still don't have a clear legal answer on whether model weights are covered by copryight.
That being said, it is good for projects to point out that there's uncertainty over whether Llama can be used commercially; so I agree with the overall point.
I have a few questions:
* you say (page 4): "We then perform standard instruction finetuning on the base LLaMA-7B model" Could you perhaps provide a reference to the _exact_ finetuning approach you used? I'm afraid different groups of people have a different notion of "standart" (see for example pages 131-155 from https://arxiv.org/abs/2302.08575 for various fine-tuning approaches) and without knowing exactly how fine-tuning was carried out, it can be very difficult reproduce your research and results exactly.
* the idea of using AST Sub-Tree Matching is nice. Could you please let me know which function in which file from your GitHub repository this is implemented in?
Again, great job on publishing this paper!
---
Best regards,
friederrr.org
Thank you! Yes, the code can be found here: https://github.com/ShishirPatil/gorilla/tree/main/eval/eval-...
Hope this helps. Let me know if you have any follow-ups!
I'm still not sure though about some nitpicky things: - do you change all the weights, or just the ones from the last layer when fine-tuning? - do you just train on the _code_ field from the JSON file with the self-instruct data, or do you also use the other fields to train (or do you use the other fields just for downstream evaluation purposes)?
I think it could be a major selling point of your paper if on Github (or in an appendix to your preprint, if you update it on arxiv), you had a section where you document the training process in detail
It seems like building anything on top of these runs into either a big GPU cost for yourself or a big compute cost if you scale for others.
4090 can run the same, very slightly faster for much more money.
# Query Gorilla server
def get_gorilla_response(prompt="I would like to translate from English to French.", model="gorilla-7b-hf-v0"):
try:
completion = openai.ChatCompletion.create(
model=model,
messages=[{"role": "user", "content": prompt}]
)
return completion.choices[0].message.content
except Exception as e:
raise_issue(e, model, prompt)[0] https://news.ycombinator.com/item?id=36326525 [1] https://credalai.notion.site/Drop-In-APIs-3a45d32405c347e8bf...
Is this really just tested against 95 API calls, and I'm guessing largely from just a small number of libriaries like pytorch?
More importantly, if anywhere near true, is there any reason to (so far) use this for use cases like OpenAI's around calling generic OpenAPI style libs (zapier scenario), known specific tools, or random python libs not in that dataset?
I'm really thinking 3 scenarios for our users:
-- python libraries we know they'll want to use ahead of time, like pandas and pygraphistry
-- Same for CLI, like AWS and az, and OpenAPI from and index
-- Long-tail that we don't expect, esp in python + js, so on the fly, with limited time budget for inspecting GitHub/Google/etc
So far, we generally find auto approaches too unreliable for non-hobbyists, and have to tune a bunch for each tool and database we teach louie. This line of research is def interesting to us...
> Gorilla is a LLM that can provide appropriate API calls. It is trained on three massive machine learning hub datasets: Torch Hub, TensorFlow Hub and HuggingFace. We are rapidly adding new domains, including Kubernetes, GCP, AWS, OpenAPI, and more. Zero-shot Gorilla outperforms GPT-4, Chat-GPT and Claude. Gorilla is extremely reliable, and significantly reduces hallucination errors.
My reading of that abstract is that it's an LLM that outputs API calls instead of natural language (or maybe it still outputs natural language, but it can use API calls during inference? I didn't read very far), whereas LangChain is simply a software library. In theory, you could probably get Gorilla to output LangChain "API" (function) calls...
Input: I would like to translate from English to Chinese.
Output:
<<<domain>>>: Natural Language Processing Text2Text Generation
<<<api_call>>>: M2M100ForConditionalGeneration.from_pretrained('facebook/m2m100_1.2B')
<<<api_provider>>>: Hugging Face Transformers
<<<explanation>>>: 1. Import M2M100ForConditionalGeneration and M2M100Tokenizer from the transformers library.
2. Load the pre-trained M2M100 model and tokenizer using the from_pretrained() method. The model is trained to translate text from English to Chinese, among other languages.
3. Encode the input text in English using the tokenizer.
4. Generate the translation using the model.generate() method.
5. Decode the output tokens using the tokenizer to obtain the translated text in Chinese.
6. Print the translated text.
Sounds like it's another LLaMA variant specifically fine tuned for API calls.