BigCode Project Releases StarCoder: A 15B Code LLM
huggingface.co
huggingface.co
Release thread: https://twitter.com/BigCodeProject/status/165417494197606811...
Model: https://huggingface.co/bigcode/starcoder
Paper: https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0k...
Also available in HuggingChat: https://huggingface.co/chat
The performance of StarCoder seems superior on HumanEval pass@1:
• replit-code-v1-3b: 21.9%
• replit-finetuned-v1-3b: 30.5%
• StarCoder: 33.6%
• StarCoder prompted: 40.8%
Replit's model seems to have focused on being cheap to train and run. StarCoder seems to be vastly better on quality.
OK for some commercial uses but I wonder how many other models have restriction lists like this
To me it would be even more responsible if they would define a default output license compatible to training input of the model. I guess they could simply refer to the input data set in there to fulfill attribution clauses of most permissive licenses. Or extract some really long list of copyright notices from the sources.
> (m) To provide medical advice or medical results interpretation that is intended to be a substitute for professional medical advice, diagnosis, or treatment;
This seems particularly overly broad -- there is a lot of benefit for people with no access to a doctor of any kind that is left on the ground by excluding that use case.
The problem of bad advice is hard, but it's one that will likely become increasingly solved as the tech evolves, and this seems to exclude this model from that technological progression.
Evaluation results are impressive!
Can someone explain to me the "Model Formats" section at the end of that page. It makes sense to me as a description of how to format the training data, but it also says "Use these templates to explore the model’s capacities"
How would you make use of this prefix format in a prompt?
<reponame>REPONAME<filename>FILENAME<gh_stars>STARS
code<|endoftext|>Page 30 of the TR has a few examples:
https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0k...
Note that they were produced with StarCoderBase. E.g., here is one:
Input:
<commit_before>def fibonacci(n):<commit_msg>add type hints to function<commit_after>def
Output: fibonacci(n: int) -> int:
(If you try this in the playground, dial down the repetition penalty to say 1.0 in Advanced Settings.)The <reponame> token specifies the name of the repository, and the same goes for the filename. A high gh_stars count can make it mimick popular open-source libraries. Admittedly, there is a bit of information that can be retrieved from those tokens, but not as much as the others.
The more interesting ones are the commit tokens, which you can use to ask it as an assistant to implement a commit whose high-level description you give in plain English, and the issues token, to make it solve an issue.
The Fill-in-the-middle tokens are most useful for autocompletion: it will allow it to gain information from the code that is present after your cursor, instead of just the code before your cursor.
Of course, in practice, those tokens are meant for code editor plugin writers. Normal users won’t know about them.
Just curious what that is vs the main https://huggingface.co/bigcode/starcoder model
It could have other uses, but its not what you want code generation. For code generation, use StarCoder or StarCoderBase.
But, I would hesitate to say that BigCode supports 80 PLs. Any LLM that claims to support 80 PLs is not presenting evidence that it does.
Quantized versions here:
https://huggingface.co/mayank31398/starcoder-GPTQ
they will be benchmarked on humaneval and released soon—maybe tomorrow?