Karpathy's MinGPT
github.com
github.com
You could say that it relies on PyTorch which is a lot of code, but most of the complexity comes from the need to do GPU/TPU/SIMD acceleration. A truly from-scratch CPU-only implementation in C would still not be a large amount of code.
I don't disagree that hardware acceleration is key in enabling these models, but I still find it interesting how simple the core techniques are.
You mean CPU implementation. C language isn't the culprit here, and low-level parts of ML frameworks are normally written in C or C++ as well and compiled with nvcc or equivalent.
I guess it's just something I have to get used to, but I wish there was an interface to do the same nd-array based logic that was designed more for developers like me rather than data scientists who perform surgery on 50d arrays all day long.
If you hand derive all equations in your program you are going to have a hard time ... but automatic differentiation (ad) and existing linear algebra libraries (for gradient descent/ADAM) make the the job pretty easy.
I recently dabbled in automatic (song-)mixing ... reimplemented the code of some open audio plugins in a way that they were auto-derivable and "trained" the whole thing on some multitracks with reference mixes. Sadly that didn't resulted in a good mix ... but in theory it should work.
What's your "simple deep learning" story?
https://github.com/Hello1024/shared-tensor
It does updates to weights based on 1 bit precision updates each iteration.
It would be fairly trivial to go to less than 1 bit precision too - simply set some threshold (eg 3), and wherever the difference between the weight on the server and the client is greater than 3, transmit a binary "1", else send a binary "0". Then entropy code all the resulting binary.
By adjusting the threshold up and down, you trade off the size of the data to send Vs precision.
https://arxiv.org/abs/2007.03970 https://arxiv.org/abs/1802.09941
Efficiency and resource cost is the big question though. You don't pay for the electricity or part wear that you don't use, and home computers or workstations may not be as efficient at performing a training run vs a task-specific setup. AI@home might end up costing even more, and increase the footprint of the model more, than doing it all together.
Part of the magic really needed is finding simpler ways to achieve the same levels of model robustness.
Sure part wear is relevant but I feel that most parts worldwide gets chucked way before they are worn out. Electric efficiency is probably quite worse though. Although you could possibly find opportunity in maximizing the load in regions that have surplus energy and/or renewable sources.
I don't understand this criticism of Transformers. Doesn't tracing (in both TorchScript and ONNX forms, which Transformers supports for exporting) just take the relevant model graph and freeze it? I don't think either contains the somewhat-weighty performance code.
Especially compared to StyleGAN, BERT, or pretty much anything else.
I used to hate the OpenAI GPT codebase: zero comments? no classes? What does "mlp" even mean? But over time, I find myself reverting to their style.
https://github.com/blue-season/pywarm/blob/master/examples/t...
https://github.com/pbloem/former/blob/master/former/transfor...
https://github.com/openai/blocksparse/blob/master/examples/t...
https://github.com/google/trax/blob/master/trax/models/trans...
The internet at large has been really good at taking opaque machine learning research and elucidating the details and recreating the results. I've seen a few posts/repositories/etc doing that for GPT-3 as well but man, 175B parameters is just so far out of reach for hobbyists. It's really a shame.
In time further research will likely make language models more efficient to where something GPT-3-like can be trained at the hobbyist level. Probably a blended model like SHA-RNN or something with dedicated memory in its architecture, so the model isn't burning precious weights on remembering e.g. Lincoln's birthday. In the meantime though it makes me sad that something as impressive as GPT-3 is solely the toy of corporations.
Millions of dollars in CAPEX and OPEX just for one model
https://www.hardwarezone.com.sg/tech-news-nvidia-dgx-a100-su...
That number is off. The DGX-2 consumes 10 kW at peak [0] and the DGX-2H consumes 12 kW at peak [1].
[0] https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
[1] https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Dat...
Input texts:
Go, rate thy minions, proud insulting boy!
Hither to London, to be crown'd our king. / Welcome, sweet prince, to London, to your chamber. / Post you to London, and you will find it so / Now to London, To see these honours in possession.
How will the country for these woful chances Misthink the king and not be satisfied!
Output:
Go, rating to London, with all these woful chances Misthink the king and not be satisfied!
It's true that the probability distribution is a sort of "edit distance". And GPT has already been used for text compression: https://bellard.org/nncp/gpt2tc.html so it seems not too far of a stretch to use it for similarity matching.
(Sure, perhaps there are more efficient or more effective techniques than using GPT for this, but I like the idea and am curious how it works.)
With Teansformer models it's quite easy to print out top Query-Key pairs to debug what happened, but that was not my intention.
EDIT: I'm running it on the IMDB dataset now ... just to see.
Prompt: This is my review of Lord of the Rings.
> I can't tell why the movie is a story with a lot of potential the main reason I want to see a movie that is Compared to the Baseball movie 10 Both and I can say it was not just a bad movie.
What kind of complexities is he talking about? Is it simply the complexity of having a batch dimension? (compared to more simple single input code)
For what it's worth, as a hiring manager who hires technical people, but not software engineers, these kinds of side projects can really help folks with a less formal education.
The upside is that if they DO understand it or at the very least learned something interesting or useful while working on the project it will help them a lot in an interview.
Folks who do what you're worried about exist, and they just don't get hired by people who aren't impressed with the shallow use of some new technology. There are also plenty of companies where that person will be fine, not really need to use GPT to do their job, and everyone will be happy anyways :)
For what it's worth I've gone back and forth over this in my career, and I do think it has a bit to do with the level of your assertion, but there's an amount of naïveté that creeps into resumes especially since advice is so widely varied, do you brag, sell yourself, show real projects, or your past titles.
Anyways, there's no right way, and I've decided that the job seeker is in a position of weakness to employers and the industry in general, and if people are seeking to better themselves - great.
When I interview and hire people, I'm the one who screens out people who lie or are incapable of doing their job, what they put on their resume is just part of the process.
Your medical example is a bad one, sorry, I'm not going to engage with it, there are gatekeepers in almost all industries, and I'm not proposing changing them by saying that someone putting Java on their resume is not the same as a doctor lying about their qualifications. One, because systems exist to vet those qualifications, and two because they're simply not on the same spectrum.
How many people have you screened, hired, and for how many roles? Is this a problem you've run into or experienced professionally, or just something that annoys you?
I've experienced both. When I ran a satellite, I employed interns and most were CS majors who claimed to know C, C++, and Java - I asked them to write strcpy in C as my interview, none of them could do it, and these people were juniors in college, so no, I'm not too worried about what people put on their resume because it's just not representative of their actual abilities ever.
If someone claims to know Java and gets a job with no actual test of their skill, that's the manager's fault.
Those projects are good things to press in a phone interview.
- Explain why an NN with n nodes across multiple hidden layers can model a more complex structure than an NN with n nodes but only one hidden layer.
- When using an NN, when is it appropriate and not appropriate to utilize a cross-entropy cost function?
- Why can a single perceptron not approximate an XOR operation?
- Why is neural network (NN) training data divided into three sets: training, generalisation, and validation? What is the purpose of each? Must the three sets be mutually exclusive?
Etc.
I'd be curious to hear your answer on that one.
https://twitter.com/elonmusk/status/1294374864657162240?s=20
Just because someone's working on a high-profile high-risk project does not in any way change their personal time: it's theirs to do what they want with.
But this is a personal project. Something that is unrelated to his job. So it doesn't really matter what "bar" it is because he's doing it for his own benefit, not Tesla's or "humanity's".
Unless you're implying he can't have personal hobbies anymore. Which was the exact point of my previous comment.
I'm saying that there is a bar and it's higher when lives depend on it.
There is a long list of activities that people could do in their personal time that would be considered inappropriate. What is done outside of work is relevant. It can form part of a character assessment which can have political and legal ramifications.
If Tesla ends up with a dangerous product then it becomes very relevant.
At least you've made it perfectly clear that you don't think people should have lives outside of work. Hope you also subscribe to the same ideology yourself and don't have any personal projects or hobbies (why are you on HN anyway, shouldn't you be working right now?) or else you'd easily be considered a hypocrite.
Either you're misunderstanding logic, law or society. If Karparthy fails he runs a real risk of having his life very closely examined. The rules are different for different people.
I do machine learning medical work for the coronavirus. I work at near optimal as humanly possible. Occasional hackernews is one of my few indulgences.
By your logic, you should have no indulgences, so you should probably stop posting.
Would you please stop posting in the flamewar style, and especially please stop snarking on HN? I realize it's an internet tradition, but we've all had lots of opportunity to see what its systemic effects are when it dominates the culture, and we don't want those here.