API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
CoRR(2024)
Abstract
We introduce API Pack, a massive multi-programming language dataset
containing more than 1 million instruction-API call pairs to improve the API
call generation capabilities of large language models. By fine-tuning
CodeLlama-13B on 20,000 Python instances from API Pack, we achieved around 10
and 5
generating unseen API calls. Fine-tuning on API Pack enables cross-programming
language generalization by leveraging a large amount of data in one language
and small amounts of data from other languages. Scaling the training data to 1
million instances further improves the model's generalization to new APIs not
encountered during training. We open-source the API Pack dataset, trained
models, and associated source code at https://github.com/zguo0525/API-Pack to
facilitate further research.
MoreTranslated text
AI Read Science
Must-Reading Tree
Example
Generate MRT to find the research sequence of this paper
Chat Paper
Summary is being generated by the instructions you defined