A ChatGPT clone, in 3000 bytes of C, backed by GPT-2 (2023)
nicholas.carlini.com
nicholas.carlini.com
If someone has a hint where the magic lies in, please explain it to me. Is it the GELU function or the model that‘s downloaded through the bash script?
But yeah, if gameplay equals the use a trained model, then the model is an asset bundle (the pak0.pak if you will) and training data is the original models, textures etc, and the training software is all the programs that are used in the asset production pipeline.
AI: How can I help you? Human: Who are you? AI: I am Alice. Human: Tell me something about a computer. AI: I am a computer model trained by OpenAI. How can I help you? Human: What is a computer? AI: I am a computer model trained by OpenAI. How can I help you? Human: Explain mathematical addition. AI: Explain mathematical multiplication. Human: 2+2 AI: 2+2 Human: Sum 2+2 AI: Sum 2+2 Human: What are your capabilities? AI: I am a computer model trained by OpenAI. How can I help you.
I'm wondering what caused the quantum leap between GPT-2 and 3? Bigger model, more data, or both? I know RLHF makes a huge difference, but even the base GPT-3 model (text completion) was very useful, given enough examples.
https://deepdreams.stavros.io/episodes/the-princess-the-fair...
So did you make it with that GPT-2 from this page?
Psst, don't tell anyone. Artificial Intelligence is the black magic we do to make money.
> You can then use this to create something like Chat GPT---just so long as you don't care about the quality of the output. (It's actually pretty terrible output, objectively speaking... But it does run.)
It's unusable and bears no relationship other than the name-dropping. But it's a program that compiles and runs.
By the looks of those in this discussion who are giving high praises about the capabilities of a project whose author admittedly states it doesn't really work, I guess the point is using buzzwords to bait the bandwagon types.
> By the looks of those in this discussion who are giving high praises about the capabilities of a project whose author admittedly states it doesn't really work, I guess the point is using buzzwords to bait the bandwagon types.
Instead of assuming everyone else is naive, perhaps consider another perspective?
Ok, allowed this time as punching up.
--
If you missed the code link (it's buried in the text): https://github.com/carlini/c-chat-gpt-2
https://www.cs.cmu.edu/afs/cs/project/ai-repository/ai/areas...
Splotch will compile it fine on modern unixen with few changes.
https://github.com/carlini/c-chat-gpt-2/blob/main/c_chat_gpt...
Thanks for the post!
UNARY(GELU, b / 2 * (1 + tanh(.7978845 * (b + .044715 * b * b * b))))
In contrast, the fast inverse square root really exploits the bit representation of a floating point input to cheaply compute an initial guess.
Could quantized weights on huggingface be used with this?
What type of problems or queries would this be really good at?
bash ./run.sh
AI: How can I help you? Human: who are you AI: I am alice. Human: what can you do for me? AI: I can help you. Human: how to say bird in Chinese? AI: bird in Chinese. Human: 2+2=? AI: 2+2= bird.
Bitter clashes between 4 and 2+2 birders around the world resulted in dozens injured, promoting the UN to call for peace between the two churches
Well, it knows how to be snarky!
Look at this article a different way. The author put a lot of information in a small context window so that it easier for readers to understand transformers. They included a useful code highlighter to ground it.
To soooo many people, even those with strong math/science, GPT is magic. This article opens the spell book, lays it out as a computation graph with all the fixing’s. The code isn’t abominable especially when paired with the prose.
It is a good piece of education and I will also call it art.
I see so many articles saying, or assuming, that 'those software people know what is happening', because it is programmed. I hear 'well they programmed it, so they know what is going on right?'.
This is pretty clearly showing, look it isn't a lot of code, and look what happens. Even if GPT2 isn't all that great, look at the amount of code, it is small, not some pile of millions of lines of complex programming to make it work.
- John Carmack, Lex Fridman Podcast (August 4th, 2022)
This was around 3 months before ChatGPT's initial release.
Timestamped: https://www.youtube.com/watch?v=I845O57ZSy4&t=14677s
Linux (and friends) do a lot of things on the command line. LLMs are good at writing text and using a command line interface. Creating LLMs and AIs takes relatively little work and is interesting, open source is especially good at this type of work. Therefore, I predict that Linux will keep pace with other operating systems when it comes to voice control.
BTW, Karpathy has a nice video tutorial about building an LLM: https://www.youtube.com/watch?v=kCc8FmEb1nY
As a claim, it's antique -- it would easily have been a view of Turing and others many decades ago.
And as a claim its as false then as now. Not least because all the code which is the actual algorithm for generating an LLM is all the code that goes into its data collection which is just being inlined (/cached) with weights.
However, more than that, it's an extremely impoverished view of general intelligence which eliminates any connection between intelligence and a body. All the "lines of code" beyond a single are concerned with I/O and device manipulation.
Thus this is just another way of repeating the antique superstition that intelligence has nothing to do with embodiment.
Frankly, when the total argument against a position consists of such "boo" words, I immediately suspect some projection of personal preferences.
But anyway I googled "intelligence vs embodiment" and found this quite nice summary albeit from 2012: https://pmc.ncbi.nlm.nih.gov/articles/PMC3512413/
The basic idea seems to be that human conscious experience is a mishmash of sensory-reflexive impulses and internal deliberation, combined into a sophisticated narrative. Simulating this kind of combination may help robots to move about in the world and to relate more closely to human experience. I have a lot of sympathy with this although I'm guarded about how much it really tells us about the potentials for AGI.
> Not least because all the code which is the actual algorithm for generating an LLM is all the code that goes into its data collection which is just being inlined (/cached) with weights.
Could this "data collection" code not potentially be put into a few thousand lines also?
This would mean hardcoding a model of the world. Which maybe, with the help of some breakthrough from physics, would be possibe (but I think the kind of breakthrough needed to get down to such a small size would be a theory of everything). But this means eliminating the self-learning part of current neural networks, which is what makes them attractive: you don't have to hardcode the rules, as the model will be able to learn them.
Absolutely.
We discussed this in the 90s and were of the opinion, then, that event state of the art NNs (in the 90s) wouldn't get much more complex because of their actual mathematical descriptions.
It's all the bells and whistles around managing training, and real-world input/output translation code.
The core or 'general case' always will be tiny.
You seem to be assuming, without any evidence, a narrow view of what embodiment entails.
The "antique superstition" claim is particularly ironic in this context, since it much more clearly applies to your own apparently irrational hasty conclusion.
Why It Would Be Preferable To Colonize Titan Instead Of Mars https://youtu.be/_InuOf8u7e4?si=hRO1ZYCZtbQXUuK9
If you fully watch it, or already know these issues, the notion of going to Mars any time soon seems outright foolish.
If you have a large enough platform, it might still be useful as fundraising campaign, a PR stunt, or as a way to rally the uninformed masses around a bold vision. But, that's about it.
At the other extreme, if you're happy to exclude 'libraries' then I could wrap these tens of thousands of lines to a bashscript, and claim I have written an artificial general intelligence in only 1 line.
BTW tinygrad shows that tens of thousands of lines are enough, it's skipping even most of the AMD kernel drivers and talks to the hardware directly.
Also most things are simple. It is the data and libraries and applications from those simple things that are interesting. Take for example C++. The actual language is a few dozen keywords. But the libraries and things built using those small things is quite large.
Also most drivers are written so that we can re-use the code. But with the right docs you can twiddle the hardware yourself. That was 'the 80s/90s way'. It was not until macos/os2/windows/xlib that we abstracted the hardware. Mostly so we could consistently reuse things.
https://old.reddit.com/r/MachineLearning/comments/2xcyrl/i_a...
But yeah, for C the LOC metric can be gamed to a silly degree
Python does have a "semicolon to combine multiple statements", and has further (lambda f: f(f)) expressions for complex expressions with local names and scopes.
(not that using either for those would result in pythonic code, but it is certainly not missing from the language)
1. Art becomes content 2. Automation to make life better becomes automation for profit 3. Life becomes meaningless achievement
At some point, there is always the counterpoint of balance and we don't have that.
We are definitely legitimate to lament about the good old time when art and mere texnical stuffs were clearly and sharply separated, while this new generation is ruining societies and our glorious civilizations with their careless use of words that mixes everything. Well, at least as legitimate as any old man in traceable history.
https://history.stackexchange.com/questions/28169/what-is-th...
If you cannot understand that there is something in man which responds to the challenge of this mountain and goes out to meet it, that the struggle is the struggle of life itself upward and forever upward, then you won't see why we go. What we get from this adventure is just sheer joy. And joy is, after all, the end of life. We do not live to eat and make money. We eat and make money to be able to live. That is what life means and what life is for.
--
George Mallory
The worldwide coverage of mountain climbing has even turned mountain climbing into a somewhat arbitrary thing in terms of "climbing everests": first it was to the top of the mountain (cool), then it was to the top in a certain time, then it was without supplementary oxygen, then it was climbing all 8,000m+ mountains, then it was doing THAT without oxygen...
At some point it's diminishing returns and becomes stupid -- and results in people just killing themselves, and AI is already on that level.
With every achievement there also comes a responsibility and our society is just achievement without responsibility.
Otherwise nobody would ever run a marathon again.
So while I agree that some people's motivations may be sound, it is how they are applied that is a perversion, and that Mallory's quotation has the germ of a phenomenon that is incredibly dangerous and isn't as deserving of awe as it is.
That process is the joy.
I just don't see the harm in what this person has done? Envisaging something that could be created, and then creating it, is what life is all about.
I just don't see how words like "perversion" and "distortion" can sensibly be applied to this work?