Replit's new Code LLM: Open Source, 77% smaller than Codex, trained in 1 week
latent.space
latent.space
- Repo: https://github.com/replit/ReplitLM/tree/main/replit-code-v1-...
- HuggingFace: https://huggingface.co/replit/replit-code-v1-3b
- Demo: https://huggingface.co/spaces/replit/replit-code-v1-3b-demo
- Early benchmark results: https://twitter.com/amasad/status/1651019556423598081
A lot about this project was surprising. We knew it was going to be good, but didn't expect to be this good -- especially surprising was the finetuned performance boost, and the fact that the model is decent at language tasks and reasoning (in some cases much better than much larger general-purpose models).
It feels like there is a lot more to do with this model, and I have a suspicion you can even make a half-decent chatbot (at least one focused on code) by finetuning it on conversation (and/or instruction) datasets.
Will follow up with a more comprehensive technical report and the UL2R version (fill-in-the-middle support).
1 - Why did you choose Markdown? It seems an odd choice for training a model like this.
2 - Have you tried to train only one single PL and then benchmark it against this more general version?
2- I like how portable it is being a single small model doing a lot of languages. Single code models are an approach that models like Salesforce/Codegen did that, but I believe we beat (or get very close) to their mono models on benchmarks.
I created this as the basis for my origami folding descriptive language. I tried to find something similar, requirements being both well structured and English-like but couldn't find any, so I created it.
The origami folding app will hopefully be out in 2 weeks, so you can see how it's used.
Many of the most-represented "languages" on GitHub are actually things like JSON, XML, HTML, CSV, text, markdown, YAML, and SVG.
More details from them here: https://blog.replit.com/llm-training
> We implement near-deduplication in our pre-processing pipeline on top of exact deduplication. We first split the files into words/tokens based on non-alphanumeric characters and remove files with fewer than 10 tokens. Next, we compute the MinHash with 256 permutations of all documents, and use Locality Sensitive Hashing to find clusters of duplicates. We further reduce these clusters by ensuring that each file in the original cluster is similar to at least one other file in the reduced cluster. We consider two files similar when their Jaccard similarity exceeds 0.85.
Near-duplicates are still difficult to measure. So we should expect duplication, and it should be proportional to the number of samples we have (even if the same variance, but I'd wager higher variance with larger duplications).
[0] https://github.com/openai/code-align-evals-data/tree/97446d9...
> It is important for these tasks to be hand-written, since our models are trained on a large fraction of GitHub, which already contains solutions to problems from a variety of sources.
So to answer your question, yes, the evaluation dataset is spoiled. You can find such unique and never before seen docstrings like
> For a given list of input numbers calculate the Mean Absolute Deviation around the mean of this dataset. Mean Absolute Deviation is the absolute difference between each element and a centerpoint (mean in this case)[1]
And here's a repo I found that is 8 years old[2]. But how about a more recent one that is even closer?[3] There's plenty more examples[4] (does anyone know how actually limit the date to prior to 2021? `pushed:<2021` doesn't work nor does using the `created` keyword. Date searching doesn't seem to work well).
In essence, we can still use this evaluation method to determine how good our model is at doing fuzzy searching. Which, mind you, is still a useful thing. But I would be careful in concluding that this means the model is good at generalizing arbitrary descriptions of code or novel pieces of code. That said, one may be able to argue that not many lines of code are actually that novel. Still, we need to be careful about our conclusions and understand the limitations of our metrics (something I am currently deeply troubled by)
[0] https://arxiv.org/abs/2107.03374
[1] https://github.com/openai/code-align-evals-data/blob/97446d9...
[2] https://github.com/bertomartin/stat4701/blob/ec2b64f629cbbf6...
[3] https://github.com/danielwatson6/hate-speech-project/blob/64...
[4] https://github.com/search?q=abs%28x+-+mean%29+for+language%3...
I wanted to demonstrate what I said above so I came up with some examples of things I think a human would have an easy time implementing but might be hard to implement. BUT a key part is that I expect these to be in the dataset! I just don't expect these to be in hundreds or thousands of githubs because they will be uncommon (but not rare). Also, we'll pretty much ask for few-liners to give the model the biggest advantage we can (errors will compound).
Prompt:
from torch import nn
class LipSwish(nn.Module):
""""
The Swish activation function is defined by a gated linear unit,
where the gate is defined by a sigmoid function and multiplies the input with
a learnable parameter, beta. Beta is initialized as 0.5.
The Lipswish function normalizes the output by the upper bound of 1.1.
""""
def __init__(self:
super().__init__()
Result: Mostly correct but missing the division by 1.1. The forward is `return x * F.sigmoid(self.beta * x)`, which is Swish (it also assumes we had "import torch" and applied type hinting). It did properly set the parameter (this is just a 3 liner)Discussion: The Swish function should be in the dataset and is a well known activation function (though beta is not in the pytorch version). Despite LipSwish being in the dataset (introduced in 2019 from Residual Flows[0]) it is not common. I could get the code to generate the swish function (initializing beta, and performing the gate) but could not get the code to divide the output by 1.1. I would not expect a human to have difficulties with this.
Okay, so let's try something else that might be a bit more common and older. The same paper uses a concatenated activation function, and those aren't "uncommon". CReLU was introduced in 2016[1] and there's plenty of concatenated activations around since then. The pytorch documentation even uses it as an example[2]. There's far more examples of CReLU (3k python results for "class CReLU" vs 58 for "class LipSwish. Use these numbers as weak hints because search sucks and isn't always accurate).
Prompt:
from torch import nn
from torch.nn import functional as F
class CReLU(nn.Module):
""""
Concatenated version of ReLU. The activation is applied to both the positive and
negative of our input and the result is concatenated.
""""
def __init__(self):
super().__init__()
def forward(self, x):
Result: `return torch.cat([x.clamp(min=0), -x.clamp(min=0)], 1)`. This is correct but not the expected one-liner result.Discussion: This was a bit surprising, it didn't use functional as we might expect (or hinted). But interestingly it will if we change the class name to "ConcatenatedReLU". I found exact copies on GitHub with the full name (memorization) but the fist page of instances for CReLU I found used functional (I did find one that was exactly the above code, when adding "clamp" to the search, but missing the minus sign. There were plenty of errors in CReLU implementations). Interesting side note: CReLU continues and defines a function CReLU6 with uses the same docstring but clamps with a max of 6 on the positive input whereas Concatenated starts to define a convolutional block (Conv + BatchNorm + ReLU) called Conv2d.
So we have kinda mixed results, and in both cases these are rather odd and probably not what we wanted. We can clearly see that there are issues where a human would not have too much trouble. There's a big issue in these types of problems: we need to memorize a lot of information (otherwise we can't even write code or know library calls) but too much memorization prevents creativity. There is a lot of gray area between the _pure_ "Stochastic Parrot"/"Fancy copy machine" vs a generalized intelligence (with a broad and flexible definition of intelligence). I'd still call them stochastic parrots because to me the evidence suggests that we're closer to the memorization side than the creation side. But that doesn't mean these frameworks aren't useful. We all know a lot of code is boiler plate (otherwise we wouldn't have the joke "copy paste from SO") and these tools can be very useful for that. But I think the utility is highly going to depend on what you are coding for and how you code. If you're doing standard stuff, this probably has high utility to you and can save you a lot of time. The same way writing macros does, but this is FAR more powerful. It can also help novices a lot. Also, if your main errors are reading mistakes (e.g. you're dyslexic) -- this is my largest problem -- then this might make things difficult as you have a tendency to gloss over text and miss minor errors. I also don't think these tools would help if you're a researcher or writing optimized or specialized code. These differences are probably why we see such differences in people's reactions. But it may also be a hint into what people do and how they work when we see who raves and who rants about these.
[0] https://arxiv.org/abs/1906.02735
[1] https://arxiv.org/abs/1603.05201
[2] https://pytorch.org/docs/stable/generated/torch.nn.ReLU.html
Edit: We can also check if code is in the stack[3]. We see that [0] is indeed in the dataset so we know there is information leakage. Interestingly the exact copy I found in the previous comment[4] isn't! (The repo, though the user is)
[3] https://huggingface.co/spaces/bigcode/in-the-stack
[4] https://github.com/bertomartin/stat4701/blob/ec2b64f629cbbf6...
I'd be very interested to hear about the choice/evaluation of the ALiBi approach for positional embedding (perhaps in the technical report).
My intuition suggests that while this allows for better generalizability for longer sequence lengths, it penalizes scenarios where an LLM might need to check for things like a function signature far away from where the next token is generated. My initial testing of this model tracks with this intuition but that's by no means a rigorous evaluation.
While intuitively it does seem like ALiBi would make it hard for the model to attend to things that are far away, in many scenarios we've tested with different models trained on different datasets, ALiBi always performs better than sinusoidal, rotary, and other embedding types, even when we're not using it to extrapolate to longer sequence lengths.
These findings have been confirmed by others, including by the BLOOM open source LM project.
Thanks for the link (which I've now skimmed beyond the abstract). What wasn't obvious to me from the abstract is that different attention heads have different penalty strengths, so if some prediction task requires long range dependencies you might expect one of the less-penalized heads to end up specializing. I wonder what would happen if the penalty for one head is zero? (The paper suggests this might've been tried and just made things worse, but unclear)
I must admit that this is a wonderfully elegant (and interpretable) way to do this... much more intuitive (to me at least, a wannabe practitioner) than all of the trig-based embeddings.
> I wonder what would happen if the penalty for one head is zero? (The paper suggests this might've been tried and just made things worse, but unclear) Yup, this is something we tried. Making one of the heads zero doesn't improve or degrade performance.
>I must admit that this is a wonderfully elegant (and interpretable) way to do this... much more intuitive (to me at least, a wannabe practitioner) than all of the trig-based embeddings.
Thanks so much!!
Could you, say, fine-tune the model every week with the latest merges? Every hour?
The base model checkpoint is licensed under the Creative Commons license (CC BY-SA-4.0). Under the license, you must give credit to Replit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests that Replit endorses you or your use.
Have you considered using Google's sparse "scaling transformer" architecture as the base? Even at 3B scale it can generate 3-4x more tokens per FLOP while being competitive at perplexity with a dense transformer. I think OpenAI uses a variant of it in their ChatGPT-3.5-Turbo product.
Here is the paper https://arxiv.org/abs/2111.12763 and the implementation https://github.com/google/trax/blob/master/trax/models/resea... if you are interested.
Hope you get to look into this!
Like why did we even get excited? This? Great work.
>You are free to:
>Share — copy and redistribute the material in any medium or format
>Adapt — remix, transform, and build upon the material
>for any purpose, even commercially.
Compare this to the latest release from StabilityAI lab DeepFloyd, "IF", which in addition to various restrictive clauses strictly prohibits commercial use: https://github.com/deep-floyd/IF/blob/develop/LICENSE-MODEL
Repl.it's release is as open as it gets these days, in my book.
is that a guess or is there a source? im curious to read more
I have an expanded list of foundational research that is likely to serve as basis for gpt4 here in my blog: https://kir-gadjello.github.io/posts/gpt4-some-technical-hyp...
Hope it helps!
Reference: How Replit used legal threats to kill my open-source project https://intuitiveexplanations.com/tech/replit/
4000+ upvotes on HN.
For instance, even this simple snippet generates wrong inline completions:
// Only return even numbers bigger than 10 from the array
const arrayFilter = (array) =>
Replit-code-v1: // Only return even numbers bigger than 10 from the array
const arrayFilter = (array) => {
return array.filter((item) => item > 10);
};
Gets it wrong, returns odd numbers.Codeium:
// Only return even numbers bigger than 10 from the array
const arrayFilter = (array) => {
return array.filter((num) => num > 10 && num % 2 === 0);
};
ChatGPT (GPT-3.5 Turbo) - Code-only, without the rest of the completion since it's instruction-tuned: const arrayFilter = (array) => {
return array.filter(num => num % 2 === 0 && num > 10);
}
Not comparable at all. For reference if anyone wants to test I ran this through the HuggingFace space using the default parameters, ChatGPT through chat.openai.com, and Codeium through the VSCodium extension on an empty JavaScript file. // return even numbers that are also more than 10
const arrayFilter = (array) =>
It would do the right thing. The fine-tuned version gets your prompt right so maybe it benefited from natural language data. Will look more into it.For example, if the instruction says "return person objects that are at least 20 years old", it might be more reasonable to generate:
array.filter(item => item.age >= 20)
as oppose to
array.filter(item => (item instanceof Person) && (item.age >= 20))
And then you dig in, and it's always far behind in some important way.
Not hating here, I love the pace of iteration, just not the hyperbole.
I have felt similar frustrations with statements that feel disingenuous too. Thanks for articulating this with such a beautifully hilarious metaphor.
On first look this seems to blow the current llama based models out of the water including the 30B ones.
Pasting what you want + url + example json with no other context and it "knows" what the url and the json is for, without even telling it.
I'm not even saying it's as good as chatGPT, but this is a tenth the size of the best llama models I've seen.
Hehe, yeah, imagine saying you made a new programming language with 77% less lines of code than Python.
It's a pretty well accepted fact now that bigger LLM = moar better without exceptions. I'm not sure why there's a race to the bottom of who'll make the most useless model that can run everywhere.
That's not true, the amount of training is a MAJOR factor.
See the Chinchilla paper - https://arxiv.org/abs/2203.15556
tl;dr - a "fully" trained small model can outperform a "undertrained" larger model. If you have a fixed amount of compute (budget), then you need to optimize for the largest model that you can fully train, and not simply up the parameter count.
EDIT: Also you can't necessarily compare the parameter count across model architectures*
This thing seems to outperform the finetuned 30B llama models I've seen.
But the problem is that these models don't exist in a vacuum, and have to go against slightly larger ones that are also compute optimal and use more data, which will definitely perform better.
Then again maybe there is a sweet spot for a model that's small enough to run effortlessly on regular machines while only serving as a control node in an autoGPT style setting, where it fetches the context it can't possibly have from a curated online database to make up for its shortcomings.
They don't have to go against those though. Most of these models are research models, either from academia or from companies experimenting to see what works. From my understanding, most of these are a - "We have X amount of USD for the next month or so, we'll try a few things, then whatever our best bet is we'll stick the time out on that".
Very few companies have the resources to train big models with as much compute as Google/OpenAI/Microsoft/Facebook.
These are also not being monetized as they're open source.
Going from their 2.7B model to 10B would be ~10X the compute (FLOPS) required for an optimal model. And this is likely their first open model and not their last, since Replit likely doesn't have the budget that openai does it makes sense they didn't want to blow their entire year's budget on their first open model.
2.7B would also be a really nice if anyone can get it working because it's more likely to be able to run in the IDE at that point instead of needed a massively scaled cloud (which might be valuable for replit).
my favorite learning is how they are pushing the state of the art - openai’s HumanEval is the industry standard benchmark for code LLMs, but Reza kindly went above and beyond to show how they use “AmjadEval” - using coder intuition to capture human preference on what output is more helpful to coders (see screenshots https://twitter.com/swyx/status/1653791019421569024?s=20)
please AMA!
Input:
below is a SQL statement:
SELECT
CAST(DATE_TRUNC('week', "t1"."TIMESTAMP") AS DATE) AS "WEEK_START",
COUNT(\*) AS "EVENT_COUNT"
FROM "ANALYTICS"."POSTHOG"."POSTHOG_EVENTS" AS "t1"
GROUP BY
"WEEK_START"
ORDER BY
"WEEK_START"
LIMIT 2000
Explain this SQL. Respond in JSON format with the following keys:
TITLE, DESCRIPTION, TABLES
JSON response:
output: {
"title": "Weekly Events Count",
"description": "Count of weekly events",
"tables": [
{
"name": "POSTHOG_EVENTS",
"columns": [
"WEEK_START",
"EVENT_COUNT"
]
}
]
}https://platform.openai.com/docs/model-index-for-researchers
https://help.openai.com/en/articles/6195637-getting-started-...
What's important is that they're preparing for the future by building all the tooling/UI/UX around coding copilots. This way, when costs and feasibility of building ChatGPT-quality LLM's drop and multiple open-source models are available, Replit has the ability to immediately drop them into their production environment. They'll also have the skills and systems to finetune any new models and wring extra performance out of them.
This is more important to users than it seems at first because current UX of things like GitHub Copilot don't allow me to use their AI against my codebase the way that I want to (the way I use ChatGPT). Right now GitHub Copilot is a glorified auto-complete, but I want it to do widespread scaffolding, refactoring, and analysis across my whole codebase. Microsoft has access to LLM's that can do this through their control of OpenAI -- but Microsoft lacks the tooling/UI/UX to bring the power of ChatGPT to me as a user of VSCode/IntelliJ/PyCharm/Visual Studio.
So if Replit can find more innovative, boundary-pushing ways of integrating LLM's, they won't necessarily need the highest quality LLM's to produce a superior user experience. It's a strong signal that Replit is well-positioned for the future, when ChatGPT-like models are democratized.
Hopefully JetBrains is paying attention. They definitely have time to wait a bit more (1-2 years?), but not a lot of time. JetBrains shouldn't solely rely on Github Copilot plug-in to provide their users with LLM's, because it's not clear that the user experience of that plug-in will stay competitive with the user experience that GitHub Copilot will offer directly in VSCode. The IntelliJ/PyCharm plugin may remain "just a fancy auto-complete" while VSCode gets more interactive workflows.
Future IDE's with LLM integration require novel, smart, clever UX typically invented only by very creative people.
It's also worth noting that Replit is not just trying to be an IDE -- they're also building a marketplace to buy/sell coding work, and establishing a small foothold as a niche cloud computing provider.
Highly recommended.
Generally we have continued finding that the more "other"/general stuff an AI model is trained on, the better it performs on specific tasks. As in, an AI model trained to identify photos of all animals will perform better than an AI model that is only trained to identify breeds of dogs. Even at identifying breeds of dogs.
Taken to the extreme, we've found that training image models with "multi-modal" LLM capabilities improves their ability to identify dogs/etc. A lot of people don't realize that GPT-4 is actually multi-modal...while OpenAI has only allowed API access to use text input, the model itself can also accept image input.
Note that we've moved on from ImageNet-style tests "Choose the most appropriate label for this image from 200 possible labels" to much more advanced "Reasoning" tests[0]. PaLI[1] is potentially the SoTA here but BeIT-3[2] may be better example for my thesis. Notice that BeIT-3 is trained on not just images, but also trained like an LLM. Yet it outperforms purely image-trained models on pure-image tasks like Object Detection and Semantic Segmentation.
More importantly, it can understand human questioning like "What type of flowers are in the blue buckets of this image?" and respond intelligently.
0: https://paperswithcode.com/area/reasoning
1: https://arxiv.org/pdf/2209.06794v2.pdf
2: https://paperswithcode.com/paper/image-as-a-foreign-language...
3: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Not if that tool is censored, and you need an uncensored version to do your work. Or maybe you have privacy considerations, or your company policies forbid using something hosted remotely or owned by another company, etc...
Currently you don't really use LLMs for designing the structure, just completing the implementation, and I think that will be very doable locally.
I'm very excited about everyone doing work even when they're not beating ChatGPT right now, of course.
But how it compares to ChatGPT right now is extremely relevant to lots of people.
It's also become very common to vaguely reference OpenAI's offerings when announcing new models without saying how they actually compare, or only mentioning some small way in which it compares favorably.
(Though it seems to often be that some comment from the article comparing to OpenAI gets promoted to the title when posted on HN, like here.)
ChatGPT ought to be a non starter for many use cases where data cannot be shared with OpenAI or where the copyright situation of the generated output could become too vague.
Having the option of open source models that potentially could be self hosted could make those use cases viable.
OpenAI probably hasn't gone through all the SOC2/etc/etc/etc/etc audit certification that AWS/GCP/Azure have, but if you're using those, then this decision is just a matter of degree. Plus OpenAI is clearly aware of the concerns and beginning to address them in order to expand their addressable market.
For defense companies, yeah, this is a non-starter. But they often don't even have access to StackOverflow and cell signals are physically, purposefully blocked by the building construction materials. And they only recently even started using cloud computing and use a purpose-built cloud at Azure/GCP/AWS that's specifically walled off for DoD partners.
Saying that they're "working on it" is not useful IMO - at the end of the day, they'll be exactly as unethical as they can get away with. We live in a time where we can comfortably say that that is "very unethical".
How is it different from storing all the sourcecode on a private GitHub.com repo?
```
# Write a function that performs the forward algorithm of an HMM. Ideally it uses numpy vectorization,
# But if you cannot do that it is fine too. Just perform forward inference, given HMM parameters and Observations.
# def HMM_forward() :
# Write your code here
# Hint:
# - Remember that HMMs are stochastic, so you will need to multiply transition matrices by observation probabilities
# - Remember that HMMs are stochastic, so you will need to multiply emission matrices by observation probabilities
# - Remember that HMMs are stochastic, so you will need to multiply emission matrices by transition matrices
# - Remember that HMMs are stochastic, so you will need to multiply emission matrices by transition matrices
# - Remember that HMMs are stochastic, so
```
It gets stuck in an infinite loop.
ChatGPT gets it write, first shot:```
import numpy as np
def forward_algorithm(A, B, pi, obs):
"""
Performs the forward algorithm for an HMM.
Parameters:
A: numpy.ndarray, shape (N, N)
Transition matrix of the HMM, where N is the number of states.
B: numpy.ndarray, shape (N, M)
Emission matrix of the HMM, where M is the number of possible observations.
pi: numpy.ndarray, shape (N,)
Initial probability distribution over states.
obs: numpy.ndarray, shape (T,)
Sequence of T observations.
Returns:
alpha: numpy.ndarray, shape (T, N)
Forward probabilities for each state at each time step.
"""
T = obs.shape[0]
N = A.shape[0]
alpha = np.zeros((T, N))
alpha[0] = pi * B[:, obs[0]]
for t in range(1, T):
alpha[t] = np.dot(alpha[t-1], A) * B[:, obs[t]]
return alpha
```
OpenAI managed to do the important but extremely hard, they moved out of the DL benchmark frame and made something that is general purpose useful. Great effort and congrats to Replit team though, hopefully they can keep iterating on this and reach ChatGPT capabilities somedayStuff like this will make your outcomes worse for any model.
Same for questions where you ask “is X good or bad? And why?”. It answers good or bad before it comes up with the reasons. That’s very plausibly ok, but it’s different from how people imagine it works and thinks.
I tried a prompt of:
# python function that returns a random integer between min and max
And it produced: def random_int(min, max):
return random.randint(min, max)
# define the size of the grid
n = 5
It doesn't add the needed import statement, and I'm unclear why it's "defining the size of the grid". """python function that returns a random integer between min and max"""
return random.randint(min, max)
def gen_random_float(min, max):
"""python function that returns a random float between min and max"""
return random.uniform(
So, it assumed the triple-quote was a function's doc string, despite it not being indented. It then assumes I'll want a similar function for floats (I assume it was cut off by a token limit).As for your prompt, it's following your prompt a little too closely and generating just the function. You can however condition it that this is the start of the program it will do the import, e.g.
# python function that returns a random integer between min and max
import
This is in fact a suggestion from OpenAI on best practices for prompting called "leading words" https://help.openai.com/en/articles/6654000-best-practices-f... # python script that prints out an integer between min and max
And it did better. Included the import, didn't add unrelated code, but did still put the code inside a function.This is fantastic for the world, this means LLMs will not be controlled by a couple of companies with the associated rents.
Yes, private LLMs will likely be a couple of years ahead of 'free' alternatives, but that's OK, we want to incentivize for profit research so long as the services are low priced in time (and in this case in short order).
AMAZING WORK.
I thought I read that, is it based upon:
https://arxiv.org/abs/2211.15533 (The Stack) ?
The point is that there will be alternatives and that will reduce the price in time further increasing the impact of the technology.
There was a possible future where only MSFT and maybe GOOG and maybe one or two other companies had this technology and extracted massive rents.
And probably, yes. While it contains 358 programming languages, obviously there's a long tail after the 20 most-represented languages. Some people might not expect without thinking about it for a bit that many of the most-represented "languages" are actually things like JSON, XML, HTML, CSV, text, markdown, YAML, SVG.
Also note that it won't be able to parse natural language nearly as well without additionally being trained on something like the LAION dataset, so this version will be more of an autocomplete like Copilot rather than something which can manifest high level business logic from whole cloth like ChatGPT.
type point = { x: int; y : int }
let manhattan_distance (a: point) (b: point) : int =
which it completed to type point = { x: int; y : int }
let manhattan_distance (a: point) (b: point) : int =
abs (a.x - b.x) + abs (a.y - b.y)
...which is a valid and correct OCaml definition of this method:https://try.ocamlpro.com/#code/type'point'='$4'x:'int;'y':'i...
I believe the overparametrization largely helps with generalization and reducing overfitting, at 2 tokens/param there's much more degrees of freedom than structures that can be learned from what I can tell (the extra capacity just provides good breathing room for internal structures). But if your model has enough capacity, and you can find a good enough training method (and you have enough data to learn the task), then you should be able to succeed in arbitrary low tokens/param, which is good to keep in mind to make efficient models.
Prompt:
>def nth_prime(n):
Completion:
> if n == 1:
> return 2
> if n == 2:
> return 3
> if n == 3:
> return 5
> if n == 4
# a method that approximates the hyperbolic tangent (clamped tanh)
def rational_tanh(x):
return (x + 1) / (x - 1)
Even gave it the BIG hint of a "clamped" and "rational" tanh, but that ain't it, chief. Forget GPT-4, I would be embarrassed to even show this as a tech demo.Here's GPT4's response:
```
import math
def clamped_tanh(x, n_terms=10): """ Approximate the hyperbolic tangent (tanh) function using a Maclaurin series expansion.
Args:
x (float): The input value for which to compute the tanh.
n_terms (int, optional): The number of terms to use in the Maclaurin series. Default is 10.
Returns:
float: The approximated tanh value.
"""
tanh_approx = 0
for n in range(n_terms):
coef = ((-1) ** n) * (2 * n + 1)
term = coef * (x ** (2 * n + 1)) / math.factorial(2 * n + 1)
tanh_approx += term
# Clamping the tanh approximation to the range [-1, 1]
tanh_approx = max(-1, min(tanh_approx, 1))
return tanh_approx
# Example usage
x = 0.5
result = clamped_tanh(x)
print(f"clamped_tanh({x}) = {result}")```
Keep in mind that I'm not even an expert (I merely earned a minor in math). In fact, I only know this because I did some research years ago on rational approximations of hyperbolic functions. The right answer gets clipped to -1 or 1 beyond (-3, 3), but it should look like: `x * ( 27 + x * x ) / ( 27 + 9 * x * x )`—though the coefficients (and clip range) can vary.
Pretending I know my Pade from my Maclaurian, I'd follow up with: "use the better Pade approximation".
I haven't heard of any similar behavior since then, which is a good sign. But a reputation can be a hard thing to shake. The CEO should have considered that before doing what he did.
It's pretty obvious that lots of people will want to take a strong code completion model, then fine tune it on their docs + libraries and then make it available inside their docs/discord/slack as a support thing.
The software ecosystem is pretty immature, and there are numerous things that need to change before the core technologies are good enough to fine tune competitive LLMs.
I do think fine tuning moderate sized LLMs on your own (pretty expensive) hardware using consumer GPUs maybe possible this year.
Unfortunately all the evidence is that training (as opposed to inference) requires high-precision, and hence high memory. This is something that consumer GPUs for the most part lack. New techniques are likely to be required (eg better sharing of training on low memory GPUs) but it's hard to predict how they will develop.
See https://github.com/fauxpilot/fauxpilot/blob/main/documentati...
def sieve_eratosthenes(n):
##a function to sort 10 numbers
def bubble_sort(a):
##a function to sort 10 numbers def insertion_sort(a):
##a function to sort 10 numbers def quick_sort(a):And the reason I say this is because these tools are answering a question that we haven't asked yet: what common problems need to be solved in this programming language, and where do I get code to solve that problem?
These LLM modules are basically telling us how to duplicate code, and what we need is the opposite: how to stop reinventing the wheel for the 100th time.
Instead of writing code for me, tell me if I already have it. If I'm writing it, tell me there's a library for that. If I'm a library writer, give me suggestions for what libraries are missing from the toolkit.
All we've done so far is begun the process of automating the production of duplicate code. With absolutely no way to go back in time and correct bugs introduced in earlier iterations. We are likely, for instance, to see 0 day attacks that affect hundreds of applications, but with no simple way to describe which applications are affected. That's going to be a first rate trainwreck.
The idea of libraries may not have been a good one. It saved human time but no library is perfect because no abstraction is perfect and this causes unnecessary bloat. It seems tha Nature does not use libraries, it uses replication instead, and we can now have that too.
Nature uses replication, but it's also horrifically complex and we have no real idea about the specifics of how it all works, or what to do when many, many things go wrong.
Also, I think nature uses cloning, which I kind of think would be called a 'library' in this case, for single-celled organisms (archaea and bacteria). In addition many eukaryotic organisms can reproduce via cloning under special situations.
I don't know, I'm not really trying to argue one way or the other. I'm kinda' thinking out loud here... but I'd like to see LLMs used to create really great libraries, or some other abstractions, that are easy to use and also understandable. It might not happen soon, but I think that there is a lot of value in moving things that way.
Sounds a lot like neural networks.
I think libraries are a result of the human brain's limited working memory capacity , while organisms or neural networks aren't so limited in what they can focus on. Perhaps the transformer became successful in syntactic abstraction because it is limited on how many things it can focus on.
Computer libraries are equally a result of limited memory and disk. Javascript is a counterexample of what happens when computing is free (because it runs on someone else's computer)
But libraries and especially frameworks as they are these days are also a giant liability more often than not. APIs change for no reason, they can be removed from the package manager at any moment without warning, people may slip malicious code into them past LGTM reviews, have recursive dependencies upon dependencies that bloat and slow down your build process, etc.
Sometimes you don't need the to install the entire damn car manufacturing plant and dealership it comes with just to get that one wheel you needed. And an LLM can just write you the code for a very nicely customized wheel in a few seconds anyway.
That's what I really want. But that would also put me out of a job.
No Kotlin T_T
I wonder if fine-tuning to a new language would even make sense. AFAIK, it is the core knowledge within the model that really matters, finetuning is essentially specialising.
Code LLM right now is not responding how a Chat LLM would respond.
~~~~~ Hats off to the team on the impressive work!
I don't know how I can test the model, but it seem loading worked. When I run `nvidia-smi` on another terminal, I see `5188MiB / 8192MiB` in the memory-usage column.
I managed to run inference locally by installing the requirements and running app.py from the demo: https://huggingface.co/spaces/replit/replit-code-v1-3b-demo/...
It is very fast on my RTX 3070, VRAM usage goes to ~= 6.3GB during inference.
Great effort of course bla bla bla...
Open source really needs some benchmarking, and up their game quality-wise.
And yes I know they're expensive as shit to train... let's not keep wasting our money and actually work together, pool our resources, to make a GOOD model.
But oh no, everyone wants to put their stamp on it. "Replit did this! Look at us!"
I think it will take some time for it to be clear who is a leader in training open source models (maybe it will be the red pajama folks?) and I think they'll get more support after that.
What does "permissively licensed" mean?