Home /paper

Gorilla Large Language Model Connected with Massive APIs

Notes on Gorilla: Large Language Model Connected with Massive APIs by Shishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. Gonzalez.

Title and authors of Gorilla: Large Language Model Connected with Massive APIs.

This research paper, first released in May 2023, introduces Gorilla, a language model designed to write accurate API calls.

Part of the Tool Use paradigm within Agentic Reasoning.

Gorilla outperforms GPT-4 in generating API calls, especially in combination with a document retriever, which allows it to adapt to changes in test-time documentation.

To assess Gorilla's capabilities, the researchers created APIBench, a comprehensive dataset of HuggingFace, TorchHub and TensorFlow Hub APIs.

The paper explores the use of Self-Instruct fine-tuning and retrieval to enable LLMs to accurately select from a large, overlapping, and changing set of tools expressed using their APIs and API documentation.

Gorilla outperforms other LLMs in terms of API functionality accuracy and reduces hallucination errors, highlighting its potential for making LLMs more reliable and applicable.

API calls generated by GPT-4, Claude and Gorilla for a speech-to-text prompt. GPT-4 hallucinates a model, Claude picks the wrong library, and Gorilla produces a correct Torch Hub call.

Four scatter plots of accuracy against hallucination for Gorilla, GPT-3.5, GPT-4, Claude and LLaMA, in zero-shot and with BM25, GPT and oracle retrievers. Gorilla is highest and furthest left in the zero-shot setting.

Gorilla system diagram. Training: 1,645 API calls from Torch Hub, TensorFlow Hub and HuggingFace are turned into 16,450 instruction and API pairs with Self-Instruct to train Gorilla-7B. Inference: a user request, optionally with retrieved API documentation, produces an API call that is executed.

Retriever-Aware Training

During training, they add retrieved API documentation to the prompt, so the model learns to use whatever documentation it's given at test time. This is what lets it adapt when an API changes.

Three examples of Gorilla adapting to changed API documentation at test time: its default response, a response using an updated model, and a response using a different model repository.

APIBench dataset curation: 1,645 API calls, with 94 from Torch Hub, 626 from TensorFlow Hub v2 and 925 from HuggingFace, converted to JSON.

Bar charts of accuracy with a GPT retriever on Torch Hub, HuggingFace and TensorFlow Hub. Gorilla scores 61.82, 47.45 and 64.96, the highest on the first two and just behind GPT-3.5 on TensorFlow Hub.

Section 4.3 of the paper, API Call with Constraints, with Table 3 comparing GPT-3.5, GPT-4, Gorilla, LLaMA and Claude on picking APIs that meet an accuracy constraint.

Section 4.2 of the paper, Test-Time Documentation Change, explaining that retriever-aware training lets Gorilla adapt when API documentation changes.

Conclusion of the paper, summarising that Gorilla, a fine-tuned LLM for calling APIs, beats prompting GPT-4, reduces hallucination, adapts to documentation changes and can satisfy constraints.