NanoGPT
github.com
github.com
[0] I think https://www.youtube.com/watch?v=LxfUGhug-iQ
--
To the mod (dang): IMHO Andrej's comment should probably be at the top of the page, not my comment. UPDATE: Looks like that's done. Thank you :-)
Do you leave hyperparams (like learning rate, batch size) the same when switching from 8xA100 to fewer GPUs, or do these need to be adjusted?
Separately, when going from 8xA100 GPU to a single A100 GPU, in the worst case we can expect the same model performance after training 8x as long correct? (And likely a bit better because we get more gradient updates in with smaller batch size)
Submissions like yours and other projects like this one (recently featured here as well) -> https://github.com/ggerganov/whisper.cpp, makes it pretty clear to me that this intuition is correct.
There's a couple tools I created back then that could push things further towards this direction, unfortunately they're not mature enough to warrant a release but the ideas they portray are worth taking a look at (IMHO) and I'll be happy to share them. If there's interest on your side (or anyone reading this thread) I'd love to talk more about it.
[1] https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs...
(edit: self-promo: I'm currently working on a Typescript follow-through of this same series of video lectures, if you want to follow along with stronger types for explanation: https://github.com/Marviel/lab-grad)
Fast.AI is great, but it takes the top down, vs the bottom up, approach. It takes you from a production-level black box that you don't understand, down to the details. The benefit there is you get good high-level intuition of how it behaves at the "let me use this technology for a job" level.
Separately, the fast.ai library is also highly recommendable -- it comes with some state-of-the-art image recognition models, and its training wrappers are really helpful particularly for image-recognition dataset training.
Karpathy's "Neural Networks: Zero to Hero" video series starts at the level of individual neurons, and works you up to the final product. For some reason both this style, and Karpathy's conciseness appeal to me slightly more. I'm also super detail-oriented, though -- and any level of "hand waving" (even if further explanation comes later) always bothers me. He's also got some pretty high-profile industry experience which carries some weight with me.
But I'll say that both are really high-quality. -- ultimately, my recommendation would be to follow whichever one speaks most to you personally after the first 1hr or so.
EDIT: Per Jeremy's response below, if you want the bottom-up approach but like the fast.ai teaching style, you should check out "part 2" of the fast.ai set of tutorials, which is exactly that.
Maybe Karpathy's approach will speak to me more--thanks for the recommendation!
There will be a new version of the part 2 course out in a few weeks. It even covers stuff like random number generation from scratch, convolutions from scratch, etc. It gradually works all the way up to Stable Diffusion.
@karpathy's and the fast.ai lessons work well together. They cover similar topics from different angles.
(I'm the primary creator of the fast.ai courses.)
Thanks for your work on fast.ai!
> running on a single 8XA100 40GB node in 38 hours of training
This is a $40-80k machine. Not a diss, but I would love to see an advance that would allow anyone with a high end computer to be able to improve on this model. Before that happens this whole field is going to be owned by big corporations.
Edit: i just noticed lambda lab. It seems they ask $8.8 per hour for an instance of this caliber. That puts the total training cost around $334. I wonder how come it is that much cheaper.
If training a large model now costs the same as driving to visit grandma, that seems like a pretty good deal.
This could correspond to taking public transport in your analogy, and would bring this within reach of most students.
If you want to reduce cost, you need to reduce the model size, and you'll get worse results for less money.
I know there are memory channel bandwidth limits and whatnot but I really wish there was a card out there with a 3090 sized die but with 96GB of VRAM solely to make it easier to experiment with larger models. If it takes 8 days to train vs. 1, thats fine. having only two of them to get 192GB and still fit on a desk and draw normal power would be great.
And NVIDIA does make cards like you are asking for - the A100 is the fast memory offering, the A40 the bulk slower memory (though they added the 80GB A100 and did not double the A40 to 96GB so this is less true now than the P40 vs P100 gen).
Oddly, you can get close to what you are asking for with a M1 Mac Studio - 128GB of decently fast memory with a GPU that is ~0.5x a 3090 in training.
The closest would be if you used some form of software bug to actually cause physical damage, certainly not impossible, but extremely unlikely compared with actually physically damaging a car.
Edit: or about $340 if you get the 8xA100 instance from lambdalabs, in the realm of normal hobby spending
But if you are looking to get performance of ChatGPT or GPT-3 then don't waste your time, all GPT-3 like small LLM models (below at least 60B params) are useless for any real world use case, they are just toys.
Fortunately very few real world use cases need to be this general.
If you are training a LLM on a domain specific corpus or finetuning on specific downstream tasks even relatively tiny models at 330m params are definitely useful and not “toys” and can be used to accurately perform tasks such as semantic text search, document summarization and named entity recognition.
Yes, thanks, that's what I meant.
> If you are training a LLM on a domain specific corpus or finetuning on specific downstream tasks even relatively tiny models at 330m params are definitely useful and not “toys” and can be used to accurately perform tasks such as semantic text search, document summarization and named entity recognition.
Agree, BERT family is a good example here.
https://horovod.readthedocs.io/en/stable/elastic_include.htm...
"...Spot instances can be interrupted, causing jobs to take longer to start or finish. You can configure your managed spot training job to use checkpoints. SageMaker copies checkpoint data from a local path to Amazon S3. When the job is restarted, SageMaker copies the data from Amazon S3 back into the local path. The training job can then resume from the last checkpoint instead of restarting...."
https://docs.aws.amazon.com/sagemaker/latest/dg/model-manage...
(I assume. I've never worked with GPT, but have done similar work in other domains).
> This creates a much smaller Transformer (4 layers, 4 heads, 64 embedding size), runs only on CPU, does not torch.compile the model (torch seems to give an error if you try), only evaluates for one iteration so you can see the training loop at work immediately, and also makes sure the context length is much smaller (e.g. 64 tokens), and the batch size is reduced to 8. On my MacBook Air (M1) this takes about 400ms per iteration. The network is still pretty expensive because the current vocabulary is hard-coded to be the GPT-2 BPE encodings of vocab_size=50257. So the embeddings table and the last layer are still massive. In the future I may modify the code to support simple character-level encoding, in which case this would fly. (The required changes would actually be pretty minimal, TODO)
These machines are pretty much crunching numbers 24/7, and your project will get appended to a queue.
So if you buy two of these cards it will take 12-13 days instead of 38 hours but only require a $2500 PC.
James Betker, who created tortoise TTS, built his own $15k machine with 8x RTX 3090 and trained the models with it. He now works for OpenAI…
If 4 GPUs connected via PCIe 4.0x16 are enough you can choose among various sRTX4 boards for 3000 series AMD Threadripper CPUs.
[§] https://www.reddit.com/r/deeplearning/comments/tw0olq/commen...
Another useful URL: https://www.pugetsystems.com/labs/articles/Quad-GeForce-RTX-...
https://timdettmers.com/2023/01/16/which-gpu-for-deep-learni...
TL;DR: You probably don't need that expensive Threadripper because 2x PCIe 4.0 x16 will not be very beneficial. Go cheap, go 2x PCIe 4.0 x8.
Besides the rental options discussed below these nvidia boxen don't look too big so either used ones will be available for cheap relatively soon, or you could just locate and liberate one in Promethean fashion.
Alternatively, if you’re researching (with the caveat that you have to either publish, open source or share your results in a blog post) you can also get access to Google’s TPU research cloud which gives you a few v3-8s for 30 days (can’t do distributed training on devices but can run workloads in parallel). You can also ask nicely for a pod, I’ve been granted access to a v3-32 for 14 days pretty trivially which (if optimized) has more throughput than 8xA100 on transformer models.
TPUs and moreso pods are a bit harder to work with and TF performs far better than PyTorch on them.
https://www.deepspeed.ai/tutorials/zero-offload/
https://medium.com/analytics-vidhya/googles-tpu-research-clo...
All I see in the comments is praise for the author as a person, so just wondering what's unique about this that's not available elsewhere? 730 upvotes and counting, assuming I'm missing something...
That doesn't mean an ONNX-ed nanoGPT won't be better, but the field of optimized text generation isn't as new as people claim.
It's an order of magnitude easier to grok the basics from this repo than from going through (admittedly more ergonomic or performant or production-ready) huggingface repos.
You will also probably need to fine tune for a specific use case, so a common approach is downloading a pre-trained model and fine tuning.
I think including the “from scratch” tuning script is educational more than anything else.
Also you can perform "fine tuning" which means you start with a trained model and train it further on your own data, allowing you to customize the model for specific tasks.
The lower the cost of training, the more profitable any resultant business. You can even envision businesses that train the model regularly to bring in new knowledge. The cheaper this is, the more opportunities open up.
More seriously, the risk that a few companies become even more powerful thanks to their restricted access to such NN is very frightening. The worth is, without legal restrictions, there is nothing that we can do against it. And I doubt that legal restrictions come in the next months / years.
We now know that whatever AI Models succeed in the future, they'll be trained by a huge company and finetuned to a specific use case. Small companies should be working on use cases, and then just upgrade to the latest SOTA model.
That sounds a bit condescending. We are probably at a point from which the government should intervene and help establish level playing field. Otherwise we are going to see a deeper divide between multibillion businesses conquering multiple markets and sort of neofiefdom situation. This is not good.
It's quite reasonable to make use of models already trained for small players.
Governments already routinely do that for pharmaceutical research or for nuclear (fusion) research. In fact, almost all major impact research and development was funded by the government, mostly the military. Lasers, microwaves, silicon, interconnected computers - all funded by the US tax payer, back in the golden times when you'd get laughed out of the room if you dared think about "small government". And the sums involved were ridiculously larger than the worth of a house. We're talking of billions of dollars.
Nowadays, R&D funding is way WAY more complex. Some things like AI or mRNA vaccines are mostly funded by private venture capital, some are funded by large philanthropic donors (e.g. Gates Foundation), some by the inconceivably enormous university endowments, a lot by in-house researchers at large corporations, and a select few by government grants.
The result of that complexity:
- professors have to spend an absurd percentage of their time "chasing grants" (anecdata, up to 40% [1]) instead of actually doing research
- because grants are time-restricted, it's rare to have tenure track any more
- because of the time restriction and low grant amounts, it's very hard for the support staff as well. In Germany and Austria, for example, extremely low paid "chain contracts" are common - one contract after another, usually for a year, but sometimes as low as half a year. It's virtually impossible to have a social life if you have to up-root it for every contract because you have to take contracts wherever they are, and forget about starting a family because it's just so damn insecure. The only ones that can make it usually come from highly privileged environments: rich parents or, rarely, partners that can support you.
Everyone in academia outside of tenured professors struggles with surviving, and the system ruthlessly grinds people to their bones. It's a disgrace.
[1] https://www.johndcook.com/blog/2011/04/25/chasing-grants/
I've also read it at many places, that academic research funding is way too misaligned. It's a shame, really.
The breakthrough will be developing this equivalent in an accessible manner and us taking care to train the thing for a couple of decades but then it becomes our friend.
But my point, poorly explained, is that whatever ChatGPT is, it isn’t original or creative thought as a human would do it.
Chomsky’s example (which is based off Turing): Do submarines swim? Yes, they swim — if that’s what you mean by swimming.
But yes, scientists can look at your experiments and show that they don't have anything in common with human thought.
I’m a nobody that you’ve never heard of and I’ve arguably made meaningful contributions. If that’s true, don’t you think there could be way more people out there than you or sibling commenter imply?
Yes, brute forcing with hard AI can produce many thoughts. But the AI wouldn’t know they are correct. It couldn’t explain why. Any discovery would only be attributable to randomness. It wouldn’t be learning from itself and its priors.
Actually there are many indications that GPT understands the data, because its output mostly makes sense. The reason it can't assign meaning the way a human would is because a human can correlate words with other sensory data that GPT doesn't have access to. That's where GPT creates nonsense.
Think carefully about what "understanding" means in a mechanistic sense. It's a form of compression, and a few billion parameters encoding the contents of a large part of the internet seems like pretty good compression to me.
In any case, GPT could still understand non-abstract things just fine. People with low IQ also struggle with abstract reasoning, and IQ tests place GPT-3 at around 83.
I'm not joking, this is really something I think will/should happen.
Yes. From 2017: "Prediction 4: The simplest 2D text encodings for neural networks will be TLs. High level TLs will be found to translate machine written programs into understandable trees."
We have something coming out that is an OOM better than anything else out there right now.
Sure, a single researcher can't replicate this at their university, but even though OpenAI likes to publish it this way, we're not really talking about research here. Research was inventing the transformer architecture, this is just making it bigger by (very smart) engineering choices. It's something companies should do (and are doing), not researchers.
It is estimated that it cost around $5M in compute time to train GPT-3.
OpenAI has received billions in investment prior to launching GPT-3, including $1B from Microsoft in 2019.
[0]: https://blogs.microsoft.com/ai/openai-azure-supercomputer/
OpenAI had raised $1B from Microsoft in 2019 and used it to train a 175B param model. Now, they have raised $10B and are training GPT-4 with 1.5T params. GPUs are capital intensive and as long as there are returns to bigger models, that's exactly where things will go.
And then three years later GPT-11 will be required to run the latest games.
That said, GPT-AlephOne only makes sense if there's a preceding GPT-∞.
On the contrary, in this thread we are are mainly talking about that.
One idea I had was to not use one single model to learn all steps of the task, but to break it up. The human brain has dedicated grammar processing parts. It is unclear whether something like a universal grammar exists, but we have at least an innate sense for rhythm. Applied to NLP, you could heavily preprocess the input. Tokenize it, annotate parts of speech. Maybe add pronunciation, so the model doesn't have to think about weird english spelling rules, and so you can deal with audio more easily later. So I would build all these little expert-knowledge black boxes and offer them as input to my network.
But there is also some inherent resource cost in large language models. If you want to store and process the knowledge of the world, it is going to be expensive no matter what. Maybe we could split the problem into two parts: Understanding language, and world knowledge (with some messy middle ground). I believe you could replace the world knowledge with a huge graph database or triple store. Not just subject-verb-object, but with attribution and certainty numbers for every fact. The idea would be to query the database at inference time. I don't know how to use this in conjunction with a transformer network like GPT-3, so you'd likely need a very different architecture.
The big benefit of this would be that it is feasible to train the language part without the world knowledge part with much less resources. But you have other benefits, too. ChatGPT is trained to "win the language game". But as they say, winning the argument does not make you right. If you have a clean fact database, you can have it weigh statements from trustworthy sources higher. You then basically have a nice natural language frontend to a logical reasoning system that can respond with facts (or better: conclusions).
So yes, very different architecture.
Whatever humans have it is many orders of magnitude better…
A good example that is not, word randomised order and kombination with Mrs Spelling and fonetic spel-ing prevent ye knot that which I wrote you to komprehend.
(My apologies to non-native speakers of English; if someone did that to me in German I'd have no clue what was meant).
A better point is that GPT-3's training set is more tokens than the number of times an average human synapse fires in a lifetime, squeezed into a network with about 3 orders of magnitude fewer parameters than the human brain has synapses.
It's wrong to model AI as anything like natural intelligence, but if someone insists, my go-to comparison (with an equivalent for image generators) is this: "Imagine someone made a rat immortal, then made it browse the web for 50,000 years. It's still a rat, despite being very well-trained."
At least for me it's perfectly understandable (except the "Mrs" part). This reminds of those "did you know you can flip characters randomly and our brain can still understand the text" copypastas that can be found everywhere. I think it's probably quite similar for word order: As long as your sentence structure is not extremely complicated, you can probably get away with changing it any way you like. Just like nobody has issues understanding Yoda in Star Wars.
Although I think there are some limits to changing word order - I can imagine complicated legal documents might get impossible to decipher if you start randomizing word order.
That's the key difference. We use language to express conceptualizations. We have some kind of abstract model somewhere that we are translating.
Maybe it isn't a cohesive model either. All I can say for certain is that - whatever it is - we are expressing it.
GPT does not express. It parrots. There is no conceptualization.
But you are of course right with GPT, it has no inner life and only parrots. It completely lacks something like an inner state, an existence outside of the brief moment it is invoked, or anything like reflection. Reminds me of the novel "Blindsight" (which I actually haven't read yet, but heard good things about!) where there are beings that are intelligent, but not conscious.
We can take a concept and refactor it symbolically. GPT can't do that. All it does is find symbols that are semantically close to other symbols.
That's circular reasoning.
Here we go again. They must have something in common, because for about 90% of the tasks the language model agrees with humans, even on novel tasks.
> We, as humans, do not use language in a generative way
Oh, do you want to say we are only doing classification from a short list of classes and don't generate open ended language? Weird, I speak novel word combinations all the time.
As was said, a different architecture.
It might be connected to the world, of course. And it might even use toys such as simulators, code execution, math verification and fact checking to further ground itself. I was thinking about the second scenario.
However high-quality data is scarce. I would be willing to fund a proper effort to create high-quality data.
https://www.deepmind.com/publications/improving-language-mod...
"Chinchilla (70B) Greatly Outperforms GPT-3 (175B) and Gopher (280B)" - https://towardsdatascience.com/a-new-ai-trend-chinchilla-70b...
I claim instead that we are still hardly scratching the surface with how we evaluate NLP systems. Also, some fields have straight up trash evaluation schemes. Summarization and ROGUE scores are totally BS and I find the claim that they even correlate with high quality summaries suspect. I say this with publications in the that subfield, so I have personal experience with just how crummy many summarizes are.
Overfitting?
A larger discussion is that the scaling laws achieve loss-optimal compute time, but the pre-training loss only improves predictions on the corpus, which contains texts written by people that were wrong or whose prose was lacking. In a real system, what you want to optimize for is accuracy, composability, inventiveness.
[0]: https://github.com/karpathy/nanoGPT/blob/master/scaling_laws...
There is also Federated Learning which seemed to start taking off, but then interest rapidly declined.
> Could this be distributed? Put all those mining GPUs to work.
Nope. It's a strictly O(n) process. If it weren't for the foresight of George Patrick Turnbull in 1668, we would not be anywhere close to these amazing results today.
https://lambdalabs.com/blog/demystifying-gpt-3
"We are waiting for OpenAI to reveal more details about the training infrastructure and model implementation. But to put things into perspective, GPT-3 175B model required 3.14E23 FLOPS of computing for training. Even at theoretical 28 TFLOPS for V100 and lowest 3 year reserved cloud pricing we could find, this will take 355 GPU-years and cost $4.6M for a single training run. Similarly, a single RTX 8000, assuming 15 TFLOPS, would take 665 years to run."
What this tells you is that there is very little money in optimizing deep learning and that NVIDIA has made it very easy to just throw more hardware at then problem.
Oh - there are a lot of people working on optimizing AI. Amongst hobbyists, academia, and corporations alike.
The thing is, if you come up with a neat optimization that saves 30% of compute for the same results, typically instead of reducing your compute budget 30%, you instead increase your model/data size 30% and get better results.
https://ieeexplore.ieee.org/document/9635657
It is somewhere from 8x to 25x faster than doing dense machine learning. The speedup was higher on the original CPU implementation and the GPU paper mentions that if there isn't enough shared memory on the GPU it will have to switch to an algorithm that has more overhead.
By neurons I actually meant "nodes"
My comment is effectively a summary of this article: https://www.kdnuggets.com/2020/03/deep-learning-breakthrough...
Edit: There is a paper for sparse spiking gradient descent promising a 150x improvement. I am not sure how practical this is because spiking neural network hardware heavily limits your model size but here it is:
Yes, but you don't know which 0.5% depending on the input text.
The system are moving in the opposite direction (look at Dojo architecture or TensTorrent)
The silver lining is that the cost of training will fall substantially with those architecture that are not based in reusing gpu.
Disclaimer: I work on these projects, both are based on our research over the past three years
> In our experiments on the Pile, a standard language modeling benchmark, a 7.5 billion parameter RETRO model outperforms the 175 billion parameter Jurassic-1 on 10 out of 16 datasets and outperforms the 280B Gopher on 9 out of 16 datasets.
https://www.deepmind.com/blog/improving-language-models-by-r...
Though, there hasn't been much follow-up research on it (or DeepMind is not publishing it).
Annotated paper: https://github.com/labmlai/annotated_deep_learning_paper_imp...
RETRO did get press, but it was not the first retrieval model, and in fact was not SOTA when it got published; FiD was, which later evolved into Atlas[0], published a few months ago.
Let's pave the road for SkyNet hard lift-off :
-The first obvious one is use of external knowledge store, aka instead of having to store facts in the neural weights where they struggle, just store them in a database and teach your neural network to use it. (This is also similar to something like webgpt where you allow your network to search the web). This will allow you to have a network of 1G parameters (and external indexes of a few TB) that is as performant as a network of 100G parameters, and with better scaling property too. You can probably gain at least 2 orders of magnitude there.
-The second leap is better architecture of your neural networks, approximating transformer that are quadratic compute by something that is linear compute (linformer) or n log n compute (reformer) can get you an order of magnitude faster by simply reducing your iteration time. Similarly using some architectures based on sparsity can give you faster computation (although some of the gains are reduced by lesser efficiency of sparse memory access pattern). Using (analog bits) Diffusion to Generatively PreTrain sentences at a time instead of token by token. You can probably gain between 1 and 3 order of magnitude here if you write and optimize everything manually (or have your advanced network/compiler optimize your code for you)
-The third leap is reduced domain : You don't have a single network that you train on everything. Training one network by domain allows you to have a smaller network that compute faster. But also it allows you to focus your training on what matters to you : for example if you want to have a mathematics network, its parameters are not influenced a lot by showing it football pictures. There is at least 2 orders of magnitude there.
-The fourth one is external tool usage. It's kind of related to the first one but whereas in the first one is readily differentiable, this one necessitate some Reinforcement Learning (that's what decision transformer are used for).
-Compression : compress everywhere. The bottlenecks are memory bandwidth related. Work in compressed form when relevant. One order of magnitude
-Distributed training : Because the memory bandwidth of inside a GPU is in the order of TB/s where as the transfer to the GPU is in the order of 10GB/s. There is an advantage to have the parameters reside on the GPU but there is a limited quantity of memory in the GPU, so distributed training (something like petals.ml) allows you to increase your memory bandwidth by collaborating. So each actor can probably gain an order of magnitude. Provided that they can keep bad actors away.
-Use free resources : The other day Steam had 10M users with GPU waiting around doing nothing, just release a dwarf fortress mod with prettier pictures and use the compute for more important tasks.
-Remove any humans in the loop : it's faster to iterate when you don't have to rely any human, either for dataset construction or model building
:)
If you want to go with <1B model, you use a BERT which is bidirectional or a T5 that is easier to fine-tune on other tasks.
* https://www.kdnuggets.com/2021/02/gpt2-gpt3-openai-showdown....
* https://bakztfuture.substack.com/p/the-chasm-between-gpt-2-a...
https://github.com/karpathy/nanoGPT/blob/master/scaling_laws...
> We find that current large language models are significantly under-trained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant ... the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.
Assuming the GPT-3 authors know this, one could surmise they 10x'ed the number of training tokens also.
Edit: Should have kept reading. Sounds like GPT-3 was found to be undertrained.
What's the best source for these weights?
I still have a few ideas there (including another secret approach at better text generation) but it's hard to determine ROI.
Not everyone is trying to replicate CHATGPT results for certain tasks.
He's done it because he evidently loves it, and wants to share his hard-earned knowledge with the rest of the world.
He may be a product of the ivory tower, but he's been in the trenches. He knows firsthand how f-ing hard it is to ship a product.
And here he is, sharing useful personal code with everyone.
This github repo now has collected ~4K stars and its antecessor (minGPT) has collected ~11K stars over the past couple of years. In my experience, the number of people who clone, copy, view or otherwise use code from a repo is one to two orders of magnitude larger than the number of people who star it, so we can safely say that Andrej has helped at least a few hundred thousand -- but likely more than a million -- individuals around the world learn how to build and tinker with GPT models.
Remarkably, as I write this, no one else here has said thank you yet, so let me say it on everyone's behalf:
THANK YOU ANDREJ.
--
EDITS: I changed the wording in response to latexr's comments below.
> Normally, the number of people who clone or copy code from a repo is one to two orders of magnitude larger than the number of people who take the time to star it
Intuitively, I’m having trouble believing that. Starring takes considerably less effort than cloning or copying code. The “time to star” is a literal second, maybe two if you have to scroll up.
From anecdotal observation, repos with more forks and/or external contributors than stars are far from the norm. I’ve seen many mentioning they star repos as a way of bookmarking they seldom go back to, or as an easy way to send kudos to the developer even when they don’t use the project.
In no way is this a comment on the value of Andrej’s work (I’m not familiar with it). I am only interested in the source of your “orders of magnitude” claim, which if proven will update my mental model of the coding community.
Of course not everyone is the same, but I'd be surprised if overall clones were less than an order of magnitude more than forks or stars, and find two or even three orders of magnitude believable depending on the target group of the repo.
My estimate of "one to two orders of magnitude" is based on anecdotal evidence. I edited my comment to reflect as much.
I think the biggest difference with a clone and a star is that a star requires an account and some vested interest in the social network of Github. Anyone who is not interested in the social aspect can just bookmark it.
I guess this differs quite a lot by target demographic. A tool for GPT will probably get a lot more stars than a plugin for some consumer software simply because it is more targeted for the audience of people who have Github accounts.
(This is a tool that most people install and run without any interaction with GitHub, since it is in package managers)
Whilst starring is simpler, the incentive is much lower than that of cloning. Especially for projects you just want to use and not contribute to or follow.
In my many years of work, i have only starred less than 50 repos. I am sure i have cloned more than a thousand.
I seldom star, but neither you nor I can be extrapolated to the general community. I have thousands of stars in some repos, and I know a significant number of those users don’t even code, let alone clone repos or copy code, they’re interested in the final product. They have GitHub accounts because it’s the way to report bugs or make feature requests.
The OP made a claim. All I’m looking to understand is if it has data backing it up or it’s just a gut feeling, because if it’s the former I’ll have learned something and made a correction of my mental model of the world. Sharing more anecdotes will leave us stuck in the same situation.
By checking the usage statistics on that server, you can get an idea how many users there are, and typically it's far higher than the number of stars.
Starring is just not that useful to me so I can see why users or contributors would be much higher. I typically star repos if it's an unpopular or old repository that doesn't have NPM or Nuget packages.
He appears to be building a business and maintaining his profile. And there is nothing wrong with that - I admire him for for pursuing his career in this positive and helpful way.
But random folks do this sort of thing everyday with no such career goals and little recognition, so I'm not sure it is this specific contribution that needs to be called out.
I go the other way. I would like to thank anyone who releases open source code, whether they cause big ripples or not.
A million people building GPT models means that one in 8000 humans on earth has built one. That seems wildly off.
Linkedin has about 100.000 profiles of data scientists. Assume generously that the actual number is 10x higher. Not correcting for the fact that a data scientist isn't always a machine learning expert, etc etc, there's just no way every single one of them even KNOWS what a GPT-like model is.
BTW, I appreciate that you preceded your comment with "Pedantry time!" -- nice gesture :-)
Should I use this or minGPT?
It says it needs 8XA100 40GB node. What is that and where do I acquire it?
Could someone else train this and then send me the model? What would be required to run it as opposed to training it?
If you just want to play with a similar but much better model goto https://chat.openai.com
Unlike OpenWebText this will run in seconds. Finetuning takes very little time, e.g. on a single GPU just a few minutes. Run an example finetuning like:
For comparison GPT-3 has more than 1000x more params (175B) and training time was around 2 months on ~1500 V100 GPUs which costs millions of dollars in cloud compute costs. Gopher with 280B params was trained on 4096 TPU-v3 chips, Microsoft Megatron-Turing NLG 530B trained on 2240 NVIDIA A100 cards (each card costs ~15k USD). And the most mind blowing is PaLM from Google with 540B params and trained on 6144 TPU v4, which costs around 10-30M USD in cloud compute to train.
Divide your 30,000 word document into a hundred 300 word chuncks. For each chunk, give as input:
Please summarize the following text into 50 words:
[chunk]
Join all the outputs together, and you now have a shorter document. Repeat the process recursively.You can improve the results by doing the process again, but this time giving some context:
Please summarize the following text, an extract of a document about [1st attempt at a summary], into 50 words:
[chunk]Then that title can be used in the 2nd round, for example using a query of the form "The following is an extract from the Introduction section of a document about The benefits and disadvantages of nuclear power in sweden:"
30,000 words would be enough to finetune an existing model. If you did that, then the model would output text similar to the finetuning data. For example, if you finetuned it on shakespeare, then you might be able to use the model to make a new play, in shakespeare's style.
But you're right - the model finetuned on shakespeare would be good at writing a new play in the style of shakespeare, but would be bad at giving a critique of shakespeare's works.
I don't mind letting my machine churn for 2-3 weeks. But I'm not looking to buy another 1000$ GPU just because CUDA is the only compute library researchers understand
According to https://www.runpod.io/gpu-instance/pricing renting out 4x A100 40GB costs $3.56 per hour.
So that's $3.56 * 2 * 38hours = $270.56 then.
Also the opportunity to play around with all the parameters fairly cheaply to find improvements. The todo section of the readme gives a small taste of that. Making bigger models works for OpenAI, but maybe the rest of us manage to make small models just perform better instead.
If you are going to run a long training job, ensure you are creating checkpoints. Be sure to use persistent storage, EBS and ensure that you check the option that it doesn't get deleted if the instance is stopped, so your checkpoint remain in the disk and you can easily restart.
I haven't tried it but prices here are much cheaper. https://vast.ai/#pricing
Is it able to re-write articles? And where could I find a guide on how to train it?
This repo trains a model--how would I prompt it and print the generated output?
Looking at this feels like seeing the source code of a 64k demo, learning about Mode 13h and trying to replicate it in Turbo Pascal.
And, much like the old days of graphics programming, there's a good chance all of this knowledge will be mostly irrelevant soon, as the abstraction layers tend to come quicker and quicker and take care of the hard foundational work underneath. Then it'll be many of us here discussing whether or not it was good to have been with it from the start, to really get it, or whether playing with the highly-abstracted components is all that's needed to succeed with it.
Either way, super cool to see the pace here and I loved the "I only have a macbook" section.
Curious why HN didn't merge the submission as it usually does. Is there a "no, submit this again" option?
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
I know it sucks when your submission was earlier and gets overlooked! We should eventually have some sort of karma-sharing to take care of this. In the meantime, it at least evens out in the long run, since the reason is randomness.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...