SantaCoder: A new 1.1B code model for generation and infilling
huggingface.co
huggingface.co
def all_prime_elements(sequence):
"""Returns every prime element of the sequence."""
return [i for i in sequence if is_prime(i)]
def is_prime(i):
"""Returns True if i is prime, False otherwise."""
if i == 2:
return True
if i == 3:
return True
if i == 5:
return True
if i == 7:
return True
if i == 11:
return True
if i == 13:
return True
if i == 17:
return True
if i == 19:
return True
if i == 23:
return True
if i == 29:
return True
if i == 31:https://i.imgur.com/RGLFim0_d.webp?maxwidth=2560&fidelity=hi...
def all_prime_elements(sequence):
"""Returns every prime element of the sequence."""
return [x for x in sequence if is_prime(x)]
def all_prime_factors(sequence):
"""Returns every prime factor of the sequence."""
return [x for x in sequence if is_prime(x) and x!= 1]
def all_prime_factors_with_limit(sequence, limit):
"""Returns every prime factor of the sequence up to the given limit."""
return [x for x in sequence if is_prime(x) and x!= 1]
def all_prime_factors_with_limit_and_step(sequence, limit, step):
"""Returns every prime factor of the sequence up to the given limit."""
return [x for x in sequence if is_prime(x) and x!= 1]
def all_prime_factors_with_limit_and_step_and_start(sequence, limit, step, start):
"""Returns every prime factor of the sequence up to the given limit."""
return [x for x in sequence if is_prime(x) and x!= 1]
def all_prime_factors_with_limit_and_step_and_start_and_end(sequence, limit, step, start, end):
"""Returns every prime factor of the sequence up to the given limit."""
return [x for x in sequence if is_prime(x) and x!= 1]
def all_prime_factors_with_limit_and_step_and_start_and_end_and_step(sequence, limit, step, start, end, step):
"""Returns every prime factor of the sequence up to the given limit."""
return [x for x in sequence if is_prime(x) and x!= 1]
(and so on) def is_prime(element):
"""Returns whether a number is prime."""
if element < 2:
return False
if element == 2:
return True
if element % 2 == 0:
return False
for i in range(3, int(math.sqrt(element)) + 1, 2):
if element % i == 0:
return False
return True
Of course, if you ask it to keep generating past that there's the usual slow descent into madness. def is_prime(n):
""" Use sieve of erasthotenes to check if n is prime. """
if n < 2:
return False
if n == 2:
return True
if n % 2 == 0:
return False
for i in range(3, int(n\*0.5)+1, 2):
if n % i == 0:
return False
return TruePaper: https://hf.co/datasets/bigcode/admin/resolve/main/BigCode_Sa...
Dataset search: https://huggingface.co/spaces/bigcode/santacoder-search
Model weights: https://huggingface.co/bigcode/santacoder
The SantaCoder paper does have some benchmarks on MultiPL-E though, so you could compare them to the Codex results on that benchmark reported here (but keep in mind that code-davinci-002 is probably even larger than the model used by Copilot): https://arxiv.org/abs/2208.08227
[1] https://thakkarparth007.github.io/copilot-explorer/posts/cop...
With a fuller context and just a handful of tries, it's unlikely that 6.7B version of incoder will be outperformed by SantaCoder.
It's true they haven't actually trained a model on the stack, and this is...not copilot. But I like what they're doing and I think it should be appreciated. Honestly, I may even say they're doing with code what stability.ai is doing with images.
What do you mean? SantaCoder is trained on The Stack:
> Dataset
> The base training dataset for the experiments in this paper contains 268 GB of Python, Java and JavaScript files from The Stack v1.1 (Kocetkov et al., 2022) after removing data from opt-out requests, near-deduplication, PII-redaction (see Section 4), and filtering based on line-length and percentage of alphanumeric characters. This dataset was also decontaminated by removing files that contained test-samples from the following benchmarks: HumanEval (Chen et al., 2021), APPS (Hendrycks et al., 2021), MBPP (Austin et al., 2021) and MultiPL-E (Cassano et al., 2022).
It's definitely not on par with Copilot yet, but SantaCoder is a trial run for a larger & better model that they're planning to train in 2023. Stay tuned! :)
def all_elements_in_range_excluding_and_including_and_excluding_and_including_and_excluding(sequence, start, end):
def all_odd_prime_elements(sequence):
"""Returns every odd prime element of the sequence."""
return [x for x in sequence if x % 2 == 1]
def all_even_prime_elements(sequence):
"""Returns every even prime element of the sequence."""
return [x for x inWe wrote up some of our learnings so far in @swyx's blog recently: https://lspace.swyx.io/p/what-building-copilot-for-x-really
I don't use anticomplete at all. What I would like is something that can take my current, bad code and style transfer it into proper, modern code. best case, take code as I write it naturally and confirm it to the style guide of my organization.
These kinds of models are particularly good at repetitive, boring work like refactoring legacy code and completing framework migrations. Unlike Copilot, we've specialized specifically in these areas and completing them end-to-end (instead of just sitting in the IDE, we open already-verified PRs).
IMO tough question of who can do codegen as a scalable standalone startup, but that's ok. Pretty darn easy & useful for many productivity platforms like ours where it's just a super nice feature as part of delivering a broader magical experience.
Related: we are hiring a k8s/pydata person, ideally who has need a user & builder of investigation platforms, as we are working w co's like Nvidia to bring this kind of thing to some pretty major enterprise & gov teams. See gdoc linked on our careers page.
So, we're not that far off from basically pair programming with an AI that will do most of the boring/tedious work we currently do manually. Something like chat gpt integrated into an IDE could be useful right now.
I do wish that the demo was a little more interactive (not needing to click buttons to create a generation) since it makes it hard to see the full power of the model.
One of the things we tried at Codeium for our playground on browser was to make it super clear how well the model performs by making the experience interactive - https://www.codeium.com/playground
> e investigate the impact of 4 preprocessing methods on the training data: filtering files from repositories with 5+ GitHub stars, filtering files with a high comments-to- code ratio, more aggressive filtering of near-duplicates, and filtering files with a low character-to-token ratio. We observe modest impact of the new filters except for the stars filter, which deteriorates performance on text2code benchmarks significantly. This is an interesting result given that previous work has explicitly filtered for GitHub Stars as a proxy for data quality
EDIT: just tried it and it didn't seem to go past ~6gb
https://github.com/arjunguha/BigCode-demos/blob/main/bigcode...
A GPU will help, but I found it passable on a CPU as well.
Its ability to generate what are essentially highly specialized tutorials that match exactly your use cases is also a really big deal.
Overall it's really extended what I'm capable of doing. Not because I couldn't do the things before but because I can skip over the boring part in the beginning and save my emotional energy for the part that actually matters.
>do I basically have to describe an algorithm in English or can I go up to a higher level and talk about features and what the software ought to do?
It understands any "well known" algorithm, api, paradigm or pattern that was written about before 2022. Even pretty obscure stuff. One thing I tried was copy and pasting some of my code into it and having it generate unit tests.
Only works somewhat well for very simple tasks and well known tasks. Any mildly more complex and it fails. It also seems to have no understanding of imports. It's barfing out a dozen of oneline-functions for common tasks, which all are just a call to some lib-function, of which half are not even in the standard-python-library.
Also kinda strange that at some point it drifts away from the demanded task, or just ends on unfinished code if the token-number is too small. For example, I asked some code relating xml-parsing and handling, and after some xml-functions, it's moving to json and yaml.
I guess with some optimization and integration, there might be some benefit for this, to replace the common stack overflow-copy n'paste. But I don't see this yet at adding significant value to actual work.
- def closest_point_on_curve(target,curve):
"""Returns the 3D point on the curve closest to target 3D point"""
Response:- return closest_point_on_curve_3d(target,curve)
def point_line_projection(line,point): """Returns the perpendicular projection of the point on the line.""" return line.point_projection(point)
def line_intersection(line1,line2): """Returns the intersection point of two lines.""" return line1.intersection(line2)
def line_intersection_point(line1,line2): """Returns the intersection point of two lines.""" return line1.intersection_point(line2)
def line_intersection_point_line(line1,line2): """Returns the intersection point of two lines.""" return line1.intersection_point_line(line2)
def line_intersection_point_line_parallel(line1,line2): """Returns the intersection point of two lines.""" return line1.intersection_point_line_parallel(line2)
def line_intersection_point_line_parallel_point(line1,line2,point): """Returns the intersection point of two lines.""" return line1.intersection_point_line_parallel
...
Prompt:
> let x = <FILL-HERE>; > assert(x + 50 === 200);
Output:
> let x = 100; > assert(x + 50 === 200);
Not yet :/
Have there been examples of novel code, i.e. code that was not in the input set?
No.
How would a model get trained on that? You'd have to pass in the entire repository for each sample. It's prohibitively difficult to create that sort of model.
If you want that, you'll have to build tooling on top of a text model (ie. an application that calls a model repeatedly), that takes a prompt and breaks it up into per-file prompts, then incrementally generates the files passing the context of previous files, and the 'context' would be too large, so you'd get large scale consistency errors.
Broadly speaking the number of tokens = the size of the text it can generate.
With small models, the number is trivial (code fragment), so generally speaking 'generate an entirely application' one-step models currently don't exist.
That said, stable diffusion has proved that you can iterate in latent space and use a VAE to upscale to larger sizes to reduce the over all model size while still having output that is ~order of magnitude larger than the latent space.
...so it's not totally out of the question that's coming.
However, right now? no.
This already exists. People have created full apps in ChatGPT by doing this. Many examples online.
What you’ve seen is applications running on top of chatGPT and using it iteratively to generate multiple code segments.
And again it does generate a full application, just not all of it at once.
Just see https://news.ycombinator.com/item?id=33854638
Sure you can't prompt it once and download a zip, but you still get a consistent full app at the end that can be used as a base from this prompting.