Tinygrad
github.com
github.com
I especially like that he outlines an actual plan for an AI chip startup that he thinks will work, and has an update explaining why he was subsequently convinced that it wouldn't work.
For example, Nvidia's compute and consumer GPU line diverged a long time ago. Modern A100s have literally only one SM capable of doing normal GPU tasks, probably to support running a display on whatever Quadro version they end up increasing. They diverged in really specific ways, for example the P100 has hardware scheduling, where as the 1080 does not (in the same way at least).
Another issue is the author spends a long time talking about how important software and ecosystem is, then completely misses that point when talking about their own CHIP - just because it is RISCV and compilers exist for that arch does not equal CUDA. Also, big re-order buffers cost area and heat that could be spent on more SMs. That's why in order to beat Nvidia you must get more specialized, they've picked their niche on the CPU-GPU-ASIC continuum, beating them at the same process node requires ditching some stuff of the stuff an Nvidia GPU. Which is why they've been specializing their arch with tensor cores.
It just also turned out those are useful for gaming with deep learning to upres the graphics, as that's easy to accelerate than driving quadraticlly higher resolutions.
what does hardware scheduling mean in this context?
Geohot is right about the AI accelerator market's problems, and a competitive 4-digit-dollars device is a great idea even if his initial strategy was way off. Although you could say almost the same thing about the high-performance CPU and GPU duopolies (Apple's chips don't count due to their proprietary OS lock-in, although I wish Asahi Linux luck at fixing that).
[edit] And beyond that, you have TSMC dominating the next-gen fab market, too.
Kinda weird to try to sneak millions of a worthless coin into your wallet.
Also posted a picture of them together in a Ferrari around this time lol (taken down)
In the spirit of learning, anyone else on his level do live streams or has a youtube channel?
Here's his last 7 hour stream coding Tinygrad.
I applaud this. Committing to keeping a project small and simple.
So many projects start small and simple, and before long they've been extended in many different directions and now have thousands of options and things to understand before you can get started.
>Each chapter consists of a walkthrough of a program that solves a canonical problem in software engineering in at most 500 source lines of code. We hope that the material in this book will help readers understand the varied approaches that engineers take when solving problems in different domains, and will serve as a basis for projects that extend or modify the contributions here.
"In bioinformatics, BLAST (basic local alignment search tool) is an algorithm and program for comparing primary biological sequence information, such as the amino-acid sequences of proteins or the nucleotides of DNA and/or RNA sequences." - could it be used to identify and group similar structures in NNs? (https://en.wikipedia.org/wiki/BLAST_(biotechnology))
I've been thinking of some kind of visual representation of weights, graphs etc. Images and evolving images as the substrate of the neural network. Pixels on a plane, planes affecting each other, activation recorded as brightness/color. Then we can use visual algorithms like SIFT (https://en.wikipedia.org/wiki/Scale-invariant_feature_transf...) to do cool stuff and grow better and better graphic-based NNs.
edit: It would be cool to see the evolution of neural networks over their training, and im transfer learning. Comparative neural network genomics.
> the full student code for minitorch. It is designed as a single repo that can be completed part by part following the guide book
> Basic Neural Networks and Modules ; Autodifferentiation for Scalars ; Tensors, Views, and Strides ; Parallel Tensor Operations ; GPU / CUDA Programming in NUMBA ; Convolutions and Pooling ; Advanced NN Functions
https://en.wikipedia.org/wiki/Minix gives more details.
It is freaking me out, remove this link.
What is the best way to ship training code in a game? Do I embed Python and PyTorch or something? Do I code my own NN training algorithm? Do I use a library such as Tinygrad?
I don’t know of any such libraries offhand and my guess would be that’s because the size of the library generally matters less if there’s a requirement to have a powerful GPU.
I usually prefer to to rewrite my training step as a pure function, so the model weights are just inputs and the gradient updates are outputs.
You need to serialize your computation graph in some way, so it can be run in C++ or some other low-level language. TensorFlow is known for doing this well since it was original design goal of the project. Some of the other frameworks that originally targeted researchers make this harder. Most mature frameworks have some way of doing this now, though, and projects like https://onnx.ai/ may solve this in general.
It gets more complicated if your model has dynamic control flow, but you get the idea.
You could also compile a neural net into a less python-tied format e.g. ONNX or torchscript. In general a siloed pytorch env would be massive, I'm assuming at least a gig or two.
For training you could just make two AI teams play against each other (off the cuff, 10-100k games should do the trick). Once you release the game, you could sample the best AI from each player's machine. Free distributed neuroevolution cluster ;)
I've wanted to do this for a while but I haven't made any proper games yet. Someday!
Edit: spelling
For example https://scholar.google.com/scholar?hl=en&as_sdt=0%2C21&q=gen... and that references a paper where they do this with MUGEN (Ai v Ai fighting game) http://irep.ntu.ac.uk/id/eprint/30021/1/PubSub7423_8186_Mart...
Genetic programming sounds like a complicated term, but it's basically an easy way to take successful characteristics and breed them into something else.
PyTorch (or Tensorflow or Keras) are the real options.
Why not just use Julia :)
> Apart from a criminal streak, Hotz shares with Raskolnikov, Dostoyevsky’s antihero, a predilection for instrumental reason and an urge to test his own mettle, to know himself by knowing his limits. As a young adult Hotz allowed himself to become addicted to prescription opiates almost as an experience in self-mastery. “I did it, I was addicted, and I quit,” he told me. “I think I had to have that experience. I don’t think I ever could have been the type who never tried it. Because in some ways I feel that if I’m not strong enough to defeat that and overcome it…” He paused for several beats before assuring me he’d never want anyone to follow his example. “In order to quit,” he continued, “it required me to rethink what I wanted out of life. After that, one of the biggest things that changed is I stopped caring about money.”
https://return.life/2022/03/07/george-hotz-comma-ride-or-die...
I assure you some people try themselves - and I do not see what is not "normal" about it. To experience, voluntary, then grow, is the norm.
-- either you go along well, hence that she agrees with your opinion is a weak test;
-- or you married her to test yourself, which would prove my point.
TND; QED. /J
Normal is what reflects the norm. That natural, observational norm (type, mode) and deontic, optimal norm ("as it should be") so typically ("normally") diverge, so the latter is found in the standard deviation, comes from the very point that was raised initially: growth («To experience, voluntarily, then grow, is the norm»), or the point where you are in it, proceeding towards the right extreme.
There is no escape from the norm (and its negative), you see: if good, then optimal norm, peripheral in the curve; if lacky, then observational norm in the centre. And between the two there is a sort of a continuity, thresholds aside...
That the "world" looks so abnormal, the bad way (hence you can call 'abnormal' the normal), is justified in such framework - especially when you look at it as a playground. And we just say, ok, if it were possible just to reduce the collateral damage...
I'm doing better now! The buprenorphine injection has made my life so much better.
Of course my trauma was one of the real driving forces behind that "experiment" and thought process. Really, my "lets find out what its like" was a rationalisation it seems.
Disruptiveness, for lack of a better term. These are people who are well-acquainted with finding the boundaries of systems and barreling through them.
I've got massive respect for everybody who addresses their own problems and fixes them. Way too many people only look for the problem in other people, but it's never that simple.
If I have a rare disease, I don't care if the doctor is nice. I want them to get the diagnosis correct.
Van Gogh painted brilliantly. And he cut off his ear.
Eminem is a great rapper. His themes can be violent.
It's open source so you could run it on your own hardware for free.
You have a complaint or question? Just check out their discord where you can talk to engineers working there instead of some outsourced chat bot.
They just figured out the minimum you would have to do to control a car and are making incremental improvements with the insane business strategy of charging more than it takes for them to build it.
Oh and if you're a business executive and want to partner, you can schedule a 30 minute phone call with comma's VP of Business Development for $1,000!
A whole amateur dissection of this guy's personality. It's so bloody awkward, just comment on the work and move on.
But it's great to know HN readers are all such paragons of virtue.
Love this ethos so much. Wish more projects would follow suit.
If you don't plan on ever doing anything but the core it does seem pretty reasonable though.
I also enjoyed watching the life streaming videos https://www.youtube.com/watch?v=Xtws3-Pk69o and the videos videos documenting the neural network ANE M1 chip reverse engineering efforts for this project. See video links here: https://news.ycombinator.com/item?id=30852818
What if they want to add a new feature that takes another ~1000 LoC? Can they just write it as a separate library and include it as a dependency? IMO trying to minimize LoC creates a perverse incentive to split up your package when it might not need it (in the same way maximizing LoC creates a perverse incentive to write the most verbose code possible, not necessarily the most readable/maintainable)
They don't.
The LoC limit exists as much to avoid scope creep and maintain focus as it is about anything else.
They like this round number. Fork. Why argue.
Subject: here's two extra lines of precious code (#307)
diff --git a/tinygrad/__init__.py b/tinygrad/__init__.py
index 0deab3e9..31ac75d5 100644
--- a/tinygrad/__init__.py
+++ b/tinygrad/__init__.py
@@ -1,3 +1 @@
-import tinygrad.optim
-import tinygrad.tensor
-import tinygrad.nn
+from tinygrad import optim, tensor, nn
There's a difference between "I will refuse features like X/Y/Z" and "I want the length of the code file to be N lines at most". The former tells you both explicitly and implicitly which features not to bother contributing. The latter is just nonsense.Import line is actually improved with this commit. LGTM, ship it.
We now return to your regularly scheduled bikeshedding.
Additionally, I feel like installing an additional tool is kind of against the spirit of 1kLOC simplicity.
Any notes on the actual code in the OP beyond the import lines formatting?
This affects the actual code, too! See, for example, https://github.com/geohot/tinygrad/commit/cfb7a4c41a2b6bcc09..., which includes this gem:
diff --git a/tinygrad/ops/ops_cpu.py b/tinygrad/ops/ops_cpu.py
index a454f56f..0686f810 100644
--- a/tinygrad/ops/ops_cpu.py
+++ b/tinygrad/ops/ops_cpu.py
@@ -2,22 +2,15 @@
from ..tensor import Function
class CPUBuffer(np.ndarray):
- def log(x):
- return np.log(x)
- def exp(x):
- return np.exp(x)
- def relu(x):
- return np.maximum(x, 0)
- def expand(x, shp):
- return np.broadcast_to(x, shp)
+ log = lambda x: np.log(x)
+ exp = lambda x: np.exp(x)
+ relu = lambda x: np.maximum(x, 0)
+ expand = lambda x,shp: np.broadcast_to(x, shp)
+ permute = lambda x,order: x.transpose(order)
+ type = lambda x,tt: x.astype(tt)
+ custompad = lambda x,padding: np.pad(x, padding)
def amax(x, *args, **kwargs):
return np.amax(x, *args, **kwargs)
- def permute(x, order):
- return x.transpose(order)
- def type(x, tt):
- return x.astype(tt)
- def custompad(x, padding):
- return np.pad(x, padding)
def toCPU(x):
return x
@staticmethod
The "actual code" is being golfed for no other reason than to keep the line count down, for the sake of a promise that never needed to be made. This commit recovered a total of seven lines wherein new code can be added, in exchange for readability. It is not just a quirk, not just some sort of interesting aside like one could argue for suckless' `dwm`: this promise actively makes the code worse over time, for no reason.In my opinion we are down the rabbit hole of arguing style over substance. I just think you are grasping at straws. Tools like black (if there was a just god it would just be part of the language, like gofmt) can easily settle stylistic choices.
Really not worth endlessly debating, what next, tabs vs spaces?
This commit seems fine to me as well. Consider the golfing serves the purpose of "the lols" if you must. Heck this particular commit looks slightly better with the change!
Excessive golfing to the detriment of code quality over time is a well known and wholly irrelevant point to make with something like Tinygrad.
If you are so convinced of what you are saying, fork and rewrite it as verbosely as you prefer and demonstrate some significant improvement to the project quality. Would be insightful.
If I were hell-bent on convincing you, I could fork their repo and add literally any feature requiring more than ten LOC and win. They would have to bend over backwards to get enough lines back to be able to merge whatever it is I wrote, with exponentially worse effects per individual line contributed.
But I'm not hell-bent on convincing you. You know that I'm not actually going to write code for you. You might think it's a clever rhetorical strategy, but it's not. Because I have other things to do, like a day job and personal projects, whose preemption of my jumping through your hoop are completely disconnected from the veracity of my line of reasoning.
And it isn't a rhetorical strategy. It is socratic questioning.
Should you sit down and attempt to write something out of spite (an excellent motivator, as good as any) it would go slightly differently than how you theorize.
In order to add a feature requiring more than ten LOC you'd actually have to come up with one in the first place :)
That would require reading Tinygrad, which shouldn't take very long for obvious reasons. Probably we've been arguing about this for longer.
At which point you'd realize there is not much to add to it and you are splitting hairs over whether it is 1200 lines or 999 lines.
In which case you could just run black on it and leave it at whatever line count it spits out. That wouldn't even require any additional coding on your part!
Then you can compare the black'd version to theirs and realize they made a very very small sacrifice as a lark. Thereby finally understanding the difference between the principal of "code golf bad" and the reality of "tinygrad l33t demo" you ol' fuddy duddy.
I'm not against something like a 64k or 4k intro competition, because the results are generally that you get to see someone take advantage of every inch of their machine, like in https://linusakesson.net/scene/a-mind-is-born/, filing down an executable byte by byte until it does what they wanted it to do. That's rad as hell.
The results here are that... someone has filed down their Python program in a way that makes it smaller. Not even in number of bytes, just in number of lines. In a really uninteresting way. It's not even like the Python examples on StackExchange's codegolf in that sense, because we could go the complete opposite route and merge most of these statements together with semicolons. It doesn't become more interesting as it gets smaller, or less interesting as it gets bigger. It's the same program regardless of how many lines it occupies, which makes it that much less interesting.
So, yes, to some extent we agree, you can just run this through automated tools and it will be as many or as few lines as you like. I see that as pointless because other restrictions would be much more stimulating, and not restricting yourself so heavily would probably lead to better (or at least more legible) results.
But sure, that'll just be this fuddy duddy's opinion.
Egads. No, it isn't the focus of the repo. You keep on missing the forest for the trees.
Tinygrad expresses its ideas succiently and clearly. It is not an exercise in code golfing. Close to the end of feature completeness they noticed it is hovering a round number and spent a minor effort pushing it down slightly for street cred.
Those changes did not impair readability or harm the project in any way. They did not go over board. You say they did, I reckon, just because you are allergic to any such changes - that is your personal subjective opinion. A waste of time you say! Bah humbug!
Well, yeah, so? Heh.
I say run it through black and make it objective, it is a great tool. It will reveal just how minor the difference is in this specific case.
I keep trying to focus on the substance. Did you actually read this project? Not through the eyes of a human linter? Was it difficult to follow? What are we arguing about if not the absolutely least interesting aspect of it really?
def log(x): return np.log(x)
def exp(x): return np.exp(x)
...
And keep the same line count, and the functions would keep their __name__.i suspect a part of the motivation for this library is learning and teaching.
and maybe a little flexing on overcomplicated autograd libraries.