StarCoder and StarCoderBase: 15.5B parameter models with 8K context length
arxiv.org
arxiv.org
Interesting to note that The Stack is 6TB - the whole of the RedPajama LLM training set (a lot more than just code) is only 2.6TB.
To get an idea what that training data looks like, I grabbed the first 300MB SQL file from https://huggingface.co/datasets/bigcode/the-stack/tree/main/... and then dumped the first 1,000 rows from that into JSON and loaded it into Datasette Lite:
https://lite.datasette.io/?json=https://gist.github.com/simo...
Here's a query that shows a random row - hit the blue "Run SQL" button to see another one: https://lite.datasette.io/?json=https://gist.github.com/simo...
(If we're comparing you to the model, is the model starting at "baby" or "teenager"?)
Leaving the details aside, the fact that a human is not starting from scratch is not in dispute. But the whole point of the discussion it seems to me is the question of exactly how humans are not starting from scratch, i.e why do we learn so much faster, and how could we apply the answer to current techniques in machine learning?
Those are still interesting questions whether or not humans and randomly wired neural nets are both starting from scratch.
It's not well understood sure but the brain is evidently playing a crucial process. Children learn to speak languages at about the same time with the same milestones occurring at roughly the same ages. Not to mention the fact that despite wildly different cultures and situations (some cultures don't attempt correct their children ever, some cultures don't speak to babies), children learn language just fine. Controversy on exactly how much aside, we're obviously predisposed to it.
>But the whole point of the discussion it seems to me is the question of exactly how humans are not starting from scratch, i.e why do we learn so much faster, and how could we apply the answer to current techniques in machine learning?
The closest biological equivalent to a parameter in an ann is a synapse. Well humans have about 100 trillion synapses. We already know that the higher the parameter count, the lower the training data required. a 50 billion parameter model will far outperform a 5 billion one trained on the same data. and a 500b one would far outperform that 50 billion one.
Economics limits how far we can go and i'm not making any declarative statements but who's to say that's not the issue ?
ann and whatever the brain does diverged in details a long time ago. It's cool to speculate and all but any special insight on the brain would have little implications on the future of deep learning. That's just not what drives architectural advances.
we could have expert level machines in a couple years but any approach trying to copy the brain is decades if not centuries away. That's how little we understand. and how little impact that actually has on the DL of today.
Current LLMS are actually nowhere near the scale of the human brain, either in parameters/neurons or training data (all the text we've ever trained an LLM on would be dwarfed by all the data humans perceive). as well as not having the headstart the human brain has. It's kind of a bogus comparison when you think about. You could easily make the case that LLMs are far more effective.
The weak spot of my argument is that it took me 20 years of on and off training, maybe 4-12 hours per day most days to get to this state. By comparison AI gets to maybe my experience level after a few years or so in months. So maybe it doesn’t actually take that much time comparatively (despite having a much lower ceiling).
The part that I’m not quite sold on though is the comparison on number of neurons. We don’t actually have a good handle on how many neurons are equivalent and a non-trivial amount of a brain’s neural net is responsible for real time signals processing of high fidelity audio and video, propiecption, motor controls etc, running your body, filtering and converting inputs into long term storage + combining it all with higher order executive functions and thought that can override a good chunk of it. It doesn’t feel like the strongest argument to make to say all that complexity is needed to create human-level intelligence in terms of comparing neurons (there may be reasons those are things are needed, but creating an LLM with the same number of neurons probably won’t work).
The compelling part for me is to continue the analogy of the brain which motivated this line of AI research. We know that the brain has all sorts of different structures and they map pretty closely to different functions and it’s not just one giant language center. Wouldn’t it make sense that we’d need different kinds of AI models to build a fully functional AI? Not least of which because specialization can be computationally more efficient (eg various computational imagery tasks are doing extraordinary things and they’re not just throwing large and larger LLMs at the problem)
A pretty obvious difference is that these models are still nowhere near as large or complex as a human brain. This network has 15 billion parameters, whereas a human brain is estimated to have 60 trillion neuronal connections. Additionally each neuron, of which a human brain has around 90 billion, can fulfill many more roles than a "neuron" in a language model.
Apples to oranges, but there's a pretty obvious complexity gap.
I'm sure there are (evolutionary?) NN models that try to do things like this but I have no idea how successful they've been.
And still, human experts in those fields don’t need as much data, even with our slow brains, the convergence rate is astounding, compared to machine learning.
The point you're missing is that those games have been designed by humans, for humans, so even if the natural selection didn't give us any advantage in playing chess per se, it conditioned our brain in a way that made us invent chess in the first place.
That being said, the original argument of comparing NN training data and natural selection is stupid anyway.
any Code LLM will be learning language and code and everything else, with absolutely no predisposition to either.
Your brain is baked with millions of years of evolution with specific areas already predisposed to certain types of processing before you ever utter a word.
Still, we can rewrite the parent's argument as:
If we train an AI on the amount of non-code-related dataset (writing read/speech heard) I've consumed, and then add to it all the amount of code-related writing/speech (coding book, coding lessons taken, code and manual pages read, man pages, etc.) I've consumed, would it even remotely as good as coding as me? Or even as good as itself is now?
I'd guess no. It's less effiecient, and thus needs way more coding dataset to get the point of coding than a human. Which brings us to:
>Your brain is baked with millions of years of evolution with specific areas already predisposed to certain types of processing before you ever utter a word.
Isn't that the whole point the parent is making?
That our advantage isn't about dataset-volume, but architecture.
Current LLMs are actually nowhere near the scale of the human brain, either in parameters/neurons or training data (all the text we've ever trained an LLM on would be dwarfed by all the sense data humans perceive), as well as not having the headstart the human brain has.
It's a bogus comparison when you really think about it. You could easily make the case that LLMs are far more efficient.
https://www.youtube.com/watch?v=HB5TrK7A4pI is a recently posted video to HN Frontpage which was summarized as such:
> Though we have been building and programming computing machines for about 60 years and have learned a great deal about composition and abstraction, we have just begun to scratch the surface.
> A mammalian neuron takes about ten milliseconds to respond to a stimulus. A driver can respond to a visual stimulus in a few hundred milliseconds, and decide an action, such as making a turn. So the computational depth of this behavior is only a few tens of steps. We don't know how to make such a machine, and we wouldn't know how to program it.
> The human genome -- the information required to build a human from a single, undifferentiated eukariotic cell -- is about 1GB. The instructions to build a mammal are written in very dense code, and the program is extremely flexible. Only small patches to the human genome are required to build a cow or a dog rather than a human. Bigger patches result in a frog or a snake. We don't have any idea how to make a description of such a complex machine that is both dense and flexible.
> New design principles and new linguistic support are needed. I will address this issue and show some ideas that can perhaps get us to the next phase of engineering design.
> Gerald Sussman Massachusetts Institute of Technology
For example, the cerebellum is 50-80% of what people keep quoting here (Number of neurons in the brain) and is not activated much in language processing.
Wernicke's area spans just a few percent of the cortical neurons. The amount of pre processing we do by providing text is actually quite enormous, so that already removes a remarkable amount of complexity from the model. So, despite the differences between biology and ANNs, it's not unreasonable what were seeing right now.
We do things in a massively parallel way and that is why and how we can do things quickly and efficiently!
Object recognition leads to abstraction. Motion perception to causality. I wouldn’t be surprised if proprioception is key to human self-awareness.
These are key logical concepts that are used in language, they are not isolated.
There are a few cases where overlaps in sensory cortex above visual, audio, and linguistic processing (the main systems every decent AI already has as inputs, which are a very small fraction of the brain) would be very helpful, but clearly not absolutely necessary, in improving the capability of a world model - for example, know that a metal container half full water will slosh differently than a full or empty one. That requires proprioception, motor skills, as well as visual inputs etc. So cases such as this will be slightly less performant, but they're typically not relevant for tasks we are interested in automating.
I'm not so sure. I'm pretty sure there are diminishing returns at play after some point.
Plus haven't we already seen models with much less billions of parameters perform the same or very close to ChatGPT with had a much higher count (Llama and its siblings)?
We can speculate about just how far this scaling can go or how far is even necessary but all i've said there is true. We have models trained and evaluated on all those sizes.
>Plus haven't we already seen models with much less billions of parameters perform the same or very close to ChatGPT with had a much higher count (Llama and its siblings)?
Only by training on far more data. Llama 13b has to be trained on over 3x more data just to reach the original GPT-3 model from 2020 (not 3.5).
The part about "far outperforming", which is the main claim, is wrong though. We saw models much smaller being developed that fare quite well, and are even competitive, with the larger ones.
You already said "only by training on far more data", which is different than "more parameters" being the only option.
I never said more parameters was the only way to increase performance. I said the training data required to reach any arbitrary performance x reduces with parameter size.
It's literally right there in what I wrote.
>a 50 billion parameter model will far outperform a 5 billion one TRAINED ON THE SAME DATA.
..because? Do you have some data to support your assertion?
I have a point of view, based on a general understanding of the universe and past inventions, and limits. Let's call it my training set.
Neural networks in literally every other field always repeat the exact same pattern. You can get from 0-80 without breaking a sweat. 80-90 is dramatically harder but you finally get there. So everybody imagines getting from 90-100 will be little more than a matter of a bit more compute and a bit more massaging of the model. But it turns out that each fraction of a percent progress you make starts becoming exponentially more difficult - and you eventually run into an asymptote that's nowhere near what you are aiming for.
A prediction based on the typical history of neural nets would be that OpenAI will be able to continue to make progress on extremely specific metrics, like scoring well on some test or another, largely by hardcoding case-specific workarounds and tweaks. But in terms of general model usage, we're unlikely to see any real revolutionary leaps in the foreseeable future.
If we see model accuracy increase I'd expect it to be thanks not to model improvement, but instead by doing something like adding a second layer where the software cross references the generated output against a 'fact database' and regenerates its answer when some correlation factor is insufficiently high. Of course that'd completely cripple the model's ability to ever move 'beyond' its training. It'd be like if mankind was forced to double check that any response on astronomy we made confirmed that the Earth is indeed the center of the universe, with no ability to ever change that ourselves.
[1] - https://www.wired.com/story/openai-ceo-sam-altman-the-age-of...
The tremendous progress over the last year makes me vary of your statement that progress will stop coming from model size improvements.
As if competitors, say Google, will take a competitor at his words and say "damn, let's scrap the expansion plans, then"?
That argument sounds highly implausible.
>The tremendous progress over the last year makes me vary of your statement that progress will stop coming from model size improvements.
Isn't "tremendous progress" before the dead-end always the case with diminishing returns and low hanging fruits?
Also not everyone can bring 500m and more to the table to train a big model in the first place.
> tremendous progress
There are things which just seem to scale and others which don't. So far it seems that adding more data and more compute don't seem to flatten out that much.
At least we should give it another year to see where it leads us.
He never said anything about technical diminishing returns. He's saying we're hitting a wall economically.
The Chief Scientist at Open AI thinks there's plenty of ability left to squeeze out.
Economics was not hinted or implied in any way. Diminishing returns on model size doesn't mean there's nothing left to squeeze out, it just means that what gains are made are going to be in model refinement, rather than going the NVidia vision of a quadrillion weight system and expecting large, or even linear, gains from that hop up in model size.
Do you have any reference to back this claim, because it sounds is very curious to me. My understanding was pretty much the opposite, that current LLM technology require a bigger training set as you increase the parameter count. I'm no NN expert in any way though.
https://arxiv.org/abs/2204.02311
If you're increasing parameter size, it's a no brainer to increase data too as that will still also increase performance.
The point is that for any arbitrary performance x, the data required to reach it reduces with size.
It also increases the cost by a lot, so it's not a no-brainer at all.
If they could beat the state of the art with only a fraction of the training cost, I suspect that they'd do so…
> The point is that for any arbitrary performance x, the data required to reach it reduces with size.
This is the claim you're making, but it's not substantiated.
Okay?.. Parameter size increases also increase cost a lot. Far more than more training data. Costs that stay well beyond training. Training on 1T tokens vs 500b won't change how resources it takes to run. Not the cases with parameter sizes.
>If they could beat the state of the art with only a fraction of the training cost, I suspect that they'd do so…
Not sure what this has to do with anything lol
>This is the claim you're making, but it's not substantiated.
I'm sorry but can you perhaps just read the paper sent ?
Google trained 3 differently sized models of the same architecture (8b, 62b, 540b) on the same dataset of 780b tokens and evaluated all 3 on various tasks.
That's quite a small sample to argue the generic point that "for any arbitrary performance x, the data required to reach it reduces with size".
Key part being: "for any arbitrary performance".
Llama 13b was better than 7b and Llama 66b was better than 33b.
If you're bothered with how general a statement I'm making then Ok, point is that all training so far has pointed towards that.
Yup, and that's why lots of work goes into smaller model trained beyond the Chinchilla-optimality. But increasing the model size alone doesn't seem to make sense to anyone for some reason.
> I'm sorry but can you perhaps just read the paper sent?
I did skim it, and it's not making the claim you are.
> Google trained 3 differently sized models of the same architecture (8b, 62b, 540b) on the same dataset of 780b tokens and evaluated all 3 on various tasks.
This has nothing to do with your claim that “We already know that the higher the parameter count, the lower the training data required”. To back such a claim we'd need a 540b model trained on 10b token beating / rivaling with a 8b parameters trained on 400b. I'm not aware of anything like this existing today.
That a big model trained with enough data can beat a smaller model on the same data isn't the same claim at all.
It's not Economically viable or efficient to just scale model size.
>This has nothing to do with your claim that “We already know that the higher the parameter count, the lower the training data required”. To back such a claim we'd need a 540b model trained on 10b token beating / rivaling with a 8b parameters trained on 400b. I'm not aware of anything like this existing today.
Literally this is what I said
>a 50 billion parameter model will far outperform a 5 billion one TRAINED ON THE SAME DATA.
A 400b dataset is not the same training data as a 10b dataset
You also literally said that:
> We already know that the higher the parameter count, the lower the training data required
And if you scroll up a bit, you'll see that this was the assertion that I've been questioning since the beginning.
Also, even this other assertion
> a 50 billion parameter model will far outperform a 5 billion one TRAINED ON THE SAME DATA.
is unsupported in the general case: will it be the case if both were trained on 10b Token? They'll both be fairly under-trained, but I suspect the performance of the biggest model would suffer more than the small one.
AFAIK, there's no reason to believe that the current architecture of LLM scaled to 100 trillions of parameters would be able to be trained efficiently on just a few millions of token like humans, and the paper you quoted sure isn't backing this original argument of yours.
> We already know that the higher the parameter count, the lower the training data required
>And if you scroll up a bit, you'll see that this was the assertion that I've been questioning since the beginning.
They follow each other. If you have a target in mind, it's the same thing in different words.
>AFAIK, there's no reason to believe that the current architecture of LLM scaled to 100 trillions of parameters would be able to be trained efficiently on just a few millions of token like humans
I didn't say it was a given. And in my original comment , I say as much.
Also Object recognition leads to abstraction. Motion perception to causality. Proprioception is a big part of human reasoning. We're not trained on only millions of tokens. And our objective function(s) are different.
Humans would not in fact outperform Language models on what they are actually trained to do. https://arxiv.org/abs/2212.11281
That's why you see it do "nonsensical" things like destroying the "unbuildable plates" at the bottom of its natural ramp. This 100% wouldn't happen if it had only learned through self-play.
Even your statement about programming skills is debatable, it depends on how you measure programming skill. They certainly are faster at it, and they know more computer languages than most people have even heard of. In fact, human programming strength seems to be more about general logic and planning skills over programming-specific skills, both things where the bulk of training happened evolutionarily and more generally over the course of a life.
The truth is, the two are not directly comparable. They are completely different architectures, at completely different scales, with entirely different strengths and weaknesses.
Ideally this includes a bunch of text books. That should give the LLM time to grok language before it starts training on more difficult texts.
Exactly, and not only that, we are agents from birth, so we enjoy the 5 E's: "Embodied, Embedded, Enacted and Extended into the Environment"
LLM's can't even write a script to see if it works, they got no feedback and can't make any meaningful choice with consequences. They are trained with "teacher forcing" and self-supervised methods, no deviation allowed.
On the other hand we got THE WORLD which is infinitely more rich than any simulation, and human society which is made of super-GPT agents, and search based access to information.
Remember, most LLMs work closed book and train on a dry, static dataset. Don't directly compare them with humans. Humans can't write great code top-down without computers either. We are trial-and-feedback monkeys, without feedback we're just no good.
It's very hard for me to ignore the intuition that there's some higher level cognitive process going on that can pick out abstract concepts and use them to control and focus the lower level "training" that might look more similar to what we're doing in ML these days.
Board games are many orders of magnitude simpler than real life, so it should be a lot easier for a computer to outperform a human with equivalent experience.
The first computers to beat (and completely surpass) the best human beings at chess were not trained on anything. Just efficient search techniques and human feedback through heuristics/opening books.
But overall, yes, the current state of machine learning relies on huge brute force compared to animal learning.
Experts Chess players need to play many games to acquire sufficient intuitive knowledge, but they converge orders of magnitude faster than current algorithms.
This weakness might be relevant later, for very dynamic and adaptable systems.
IMO it’s a meaningless comparison
Imagine Alice is abducted by aliens and given reams and reams of unfamiliar symbols and trained to predict which one came next given a long long prefix. Alice held in a cell alone with just symbol sequences for 15 years, and by the end of that period she's gotten pretty good at predicting which symbol comes next. Bob's experience is exactly the same. Neither has any way to understand what any of the symbols mean. Finally, Alice and Bob are let out of their cells for a break, and meet Krang. Krang explains that Alice has been doing a sometimes acceptable job of producing computer code for a kind of computer she's never been able to directly interact with! She might have gotten really good by the end of year 1 if anyone had explained that she was writing programs, or given her access to a REPL, or a debugger, or a manual. But she's been trained with exactly the same procedure as Bob, who has been pumping out advertising copy.
Current code LLMs are only doing next token prediction, and critically they don't have access to a model of formal semantics for each language, an interpreter or debugger or compiler, etc. This is a shame, because program generation is arguably one of relatively few areas in which we could give our models a "complete" view of the domain. An appropriately structured model could generate the program, predict and observe the AST, predict and observe the IR graph, predict and observe generated bytecode, predict and observe program traces from execution, etc, etc. But it doesn't do any of that. It doesn't have an explicit model of what the program will do during execution. It doesn't have an ability to check that an invariant is maintained at each iteration of a loop. It doesn't get to check that what it wrote behaves as intended.
Yesterday, one of the chat models which also can generate code gave me a Kotlin example which used a language feature that Kotlin doesn't actually have (basically scala-style pattern matching), and of course was totally unaware that the generated code was not even valid Kotlin because it never attempted to call any part of the toolchain.
In RPG terms, the models put everything into intelligence and no points into wisdom.
- multimodal
- feedforward
- proximal zone of developpement (you don’t start reading with shakespeare)
The programs that beat humans at chess and go have added structure to be able to plan ahead; they use a Monte Carlo search to play out the moves that "intuitively" look better, with another "intuition" check to see how good the position looks in the end. Similarly, AlphaCode [1] generates a large set of potential solutions and uses additional logic to verify that the code compiles, runs, and passes tests.
[1] https://www.deepmind.com/blog/competitive-programming-with-a...
So, as you say, if we had as much training as those LLMs, we'd be similarly good at coding by gut feel, with barely a conscious thought - and that's across pretty much any domain and technology that existed today. Compare with generic LLMs: a typical adult will be quite adept at saying somewhat coherent things on autopilot when prompted (!), which is reasonable given nearly two decades of constant exposure to natural language as written and spoken - but that same adult will be nowhere as good at this as GPT-4, and definitely not across so many different domains.
Strongly disagree.
LLM traps you inside an intellectual bell curve.
My wife sometimes views me as long-form autocomplete, and sometimes as a spell and grammar checker. Hell, my reply to your comment here is indistinguishable from a "long-form autocomplete".
Point being, that autocomplete has to work in some way. Our LLM autocompletes have been getting better and better at zero-shot completion to arbitrary long-form text, including arbitrary simulated conversations with a simulated human, without commensurate increase in complexity or resource utilization. This means they're getting better and better at compressing their training data - but in the limit, what is the difference between compression and understanding? I can't prove it formally, but I rather strongly believe they are, fundamentally, the same thing.
Also: if it walks like a duck, quacks like a duck, swims like a duck, ducks like a duck, and is indistinguishable from a duck on any possible test you can think of or apply to it, then maybe your artificial faux-duck effectively turned into a real duck?
I'm not sure this is true in general. I feel as if I understand something when I grasp it in its entirety, not when I've been able to summarize it concisely. And conceptually I can compress something without understanding it by manually implementing compression algorithms and following their instructions by rote.
I think understanding and compression are plausibly related; one test of whether I understand something is whether I can explain it to a layperson. But I don't see how they're equivalent even asymptotically.
> then maybe your artificial faux-duck effectively turned into a real duck?
I can't really get behind this sentiment. If a language model behaves like a duck in every readily observable particular then we can substitute language models for ducks, sure. But that does not imply that a language model is a duck, and whether it even could be a duck remains an interesting and important question. I'm sympathetic to the argument that it doesn't really matter in day-to-day practice, but that shouldn't stop us from raising the question.
You wrote:
> I feel as if I understand something when I grasp it in its entirety, not when I've been able to summarize it concisely.
But what does it mean to "grasp it in its entirety"? To me, it means you learned the patterns that predict the thing and its behavior. That understanding lets you say, "it is ${so-and-so}, because ${reason}", and also "it will do ${specific thing} when ${specific condition} happens, because ${reason}", and have such predictions reliably turn true.
To me, replacing a lot of memorized observations with more general principles - more general understanding - is compression.
A simplified model: you observe pairs of numbers in some specific context. You see (1, 2) and (3, 6), then (9, 18), then (27, 54), and then some more numbers you quickly notice all follow a pattern:
Pair_n = (x, y), where:
- y = 2*x
- x = 3^n
A thousand of such pairs pass you by, before they finally stop. Do you remember them all? It's not a big deal ever since you figured out the pattern - you don't need to remember all the number pairs, you only need to remember the formula above, and that n started at 0 and ended at 999.This is what I mean by understanding being fundamentally equivalent to compression: each pattern or concept you learn lets you replace memorizing some facts with a smaller formula (program) you can use to re-derive those facts. It's exactly how compression algorithms work.
And yes, in this sense, we are lossy compressors.
Markov chains (and a lot of caching) were a good high-level working model, but quite inadequate in power when inspected in detail. Deep language models I initially ignored, as they felt more like doubling down on caching alone and building convoluted lookup tables. But, to my surprise, LLMs turned not only to be a better high-level analogy - the way they work in practice feels so close to my experience with my own "inner voice", that I can't believe this is just a coincidence.
What I mean here is, in short: whenever I read articles and comments about strengths and weaknesses of current LLMs (especially GPT-4), I find that they might just as well be talking about my own "inner voice" / gut-level, intuition-driven thinking - it has the same strengths and the same failure modes.
I would offer the counter to that: had anyone written that much code, they may be able to code in their sleep, but my 5 years of college tell me that just reading calculus books does nothing in the world for making me able to use calculus
I don't recall the 10,000 hour hypothesis being that one watches 10,000 hours of tv on a subject, rather that one has practiced something for 10,000 hours
That's a bold claim. What makes you think so?
The transformer architecture isn't capable of general/unstructured thought like we are, but I'd love for someone to build an ideation model that feeds into a transformer when it needs to get input/output from the outside world. Abstract concepts will give us stronger AI, not 6TB of text; that should be reserved exclusively for communication purposes.
It's like when you're thinking of something, but you can't remember the word for it. Then you remember the word. Transformers only work with the latter, what we need is something that can operate on the "something" without having to know the word for it (which is only relevant when talking to a human anyway).
I'm not sure whether the AI learning that it can just write "#TODO" is a sign our jobs are safe or a sign our jobs are truly in danger.
I have seen approaches with merging context across multiple levels. But that can only do so much. Is it viable to fine-train a model to a specific code-base so it has knowledge across all files? Does anyone have more info on this kind of problem space?
Can you please share more about the merging context across levels? This sounds interesting!
1: "Language Models are Few-Shot Learners" Brown et al. https://arxiv.org/pdf/2005.14165.pdf
I too have a job where almost every question is about structural understanding and improvement of a large existing codebase. I'd love to have AI help, but I think it's going to take another iteration or three of model architecture to get there.
So far the SourceGraph product ("Cody") is rather underwhelming, it doesn't seem to make deep use of the project context, where the blog post seemed to say that SourceGraph's special sauce would make that possible
The results seem quite similar with Copilot Chat - in both cases they seem to basically stuff the currently focused file as context to your prompt, and the results are no better than if you did the same with ChatGPT, and looks worse because it's cramped into a VS Code sidebar.
From what I saw on the Discord after getting my invite, people were making requests in the chat for specific github repos to get indexed... is that how it works? So for projects with dependencies which have been indexed I might get better results? And I need to get my own project indexed too?
Not to mention legal being happy with handing over their codebase to an external vendor for reasons other than source control.
And all this despite the "pass@k" evaluation metric, which is very misleading: it's clearly selected to make a code generating model look its absolute best. For example, the "pass@1" metric is _estimated_ not by choosing a single solution generated by a model, for a given programming task, and checking whether the solution completes the programming task correctly, but by generating a single solution multiple times (200 or 20, depending on model) and then averaging over them. So while it's called "pass-at-one" the "one" is actually a bunch of randomly drawn samples and not a single solution. Like I say, very misleading. See Section 6.1.1 in the paper.
Also, investing serious time here on cleaner processes and better documentation of their specific implementation seems like a waste of time. I don't think much of any of this will still be in use next year, everybody will have moved on to more advanced projects.
But I agree, once things stabilize, the most popular models should invest in clean up a fair bit.
I feel quite similar personally, I've worked hard on open source and I'll never have the same permissive license again after this.
A big THANK YOU to everyone who made it possible.
I'm looking forward to playing with it -- and also, eventually, inevitably, running a quantized, super-efficient version on my laptop.
When 4-bit quantization comes out, I would expect a GPU with 12GB VRAM to be able to run it.
Disclaimer: I work at Hugging Face
I don't think StarCoderBase is instruction-tuned off the bat but would serve as a good starting point for a new technique.
RLHF is fine for things that are hard to measure and evaluate, but code is runnable and testable.
I propose we try Reinforcement Learning Machine Feedback or RLMF.
Prompts and responses are evaluated by how accurate the response evaluated to. We can then train a reward model to help refine StarCoder.
[1] https://open.spotify.com/episode/2a8Rtm4mhjzennOoAByFKx around 15:10
>"The Stack is a collection of source code from repositories with various licenses. Any use of all or part of the code gathered in The Stack must abide by the terms of the original licenses, including attribution clauses when relevant."
Does it have a view of what licenses can mix, or is it simply disallowed from crossing that boundary and only offer answers sourced entirely within the confines of this or that specific license? The latter poses some interesting scenarios and questions.
There are 193 licenses in total. v1.0 of The Stack included MPL/EPL/LGPL whereas v1.1+ doesn't include them.
This also includes humans. We "hallucinate" in very similar ways. For example mistaking localhost:8080 for localhost:8008 in a large config file. Attempting to use methods that were deprecated and no longer exist, etc.
IMO there's two ways to prevent this is - one is to make better performing models (architecture/training data/training amount/etc)
The other is the exact same as humans. Compile time tools that let it know immediately if it hallucinated, types, linting, tests, etc.
You just do it as a loop the exact same as a human. You write code, the compiler tells you that method doesn't exist, you adjust your code/consult the documents (also doable with agents).
What hardware are you using? (CPU,RAM,GPU,VRAM)
Have you considered using llama.cpp for a mixed CPU+GPU use (if you have enough RAM)
2x 24G - for dual GPU ~ 28B model
1x 24G ~ 14B model
etc.
Personally I'm concerned about how model hosting has been concentrated in one company, and was previously very unhappy that they required accounts, but I think that's past. Let me know if it's still the case for some things.
When I go to https://huggingface.co/bigcode/starcoder it says "You need to agree to share your contact informations to access this model" and "[Log in] or [Sign Up] to review the conditions and access this model content."
It's enough of a deal breaker for me to not bother using the model. Especially when it's developed by a company that (I assume) wants to harvest your contact info - unless there's some other explanation for the login requirement.
(I tried git lfs clone and got asked for hf login credentials)
Also, if you click through to the link posted up thread and go to "files" to try and download them, it explicitly says
You need to agree to share your contact informations to access this model
That may be boilerplate but it certainly implies they're taking your information, not just asking you to agree to a license.Well OK but then I don't like their license either, if they have make me do things I don't like as a requirement to make their license work.
fn convert_ogg_to_w av(input: Path) -> Result<"We perform the most comprehensive evaluation of Code LLMs to date and show that StarCoderBase outperforms every open Code LLM that supports multiple programming languages and matches or outperforms the OpenAI code-cushman-001 model."
So I'd assume not up to par with gpt4 or copilot. Can't wait to see it evolve from here!