The tiny corp raised $5.1M
geohot.github.io
geohot.github.io
This project fits the pattern of his previous projects: he gets excited about the currently hot thing in tech, makes his own knockoff version, generates a ton of buzz in the tech press for it, and then it fizzles out because he doesn't have the resources or attention span to actually make something at that scale.
In 2016, Tesla and self-driving cars led to his comma one project ("I could build a better vision system than Tesla autopilot in 3 months"). In 2020, Ethereum got hot and so he created "cheapETH". In 2022 it was Elon's Twitter, which led him to "fixing Twitter search". And in 2023 it's NVIDIA.
I'd love to see an alternative to CUDA / NVIDIA so I hope this one breaks the pattern, but I'd be very, very careful before giving him a deposit.
It seems like they've been assuming Twitter is the way it is because it was staffed by technically incompetent leftists, and if only they could apply their own get-things-done attitude and "neutral" politics, then the problem would be trivially fixable.
Where does this fallacy come from? Is it because of the illusory simplicity of the tweet format? Something like: "We just need to come up with the right algorithm and do an embarrassingly parallel run over these tiny 280-character chunks of text. How hard can that be. In my own Very Serious Day Job, I deal with oompabytes of very complex data. This tweet processing stuff should be child's play in comparison."
I've seen enough engineers presume they can easily become experts in law; I haven't seen many lawyers presume they can easily become experts in engineering.
Why?
Dunning-Kruger is, approximately "I'm good at the thing I do" (by someone who is actually incompetent).
What I'm talking about is "That thing that other people are doing is really easy; I'd be good at it" (the thing is not easy, and they would not be good at it).
If the person in the latter case actually ends up doing the allegedly easy thing, they may realise that actually they are not good at it, in which case it's not Dunning-Kruger. This is pretty common, I think; person barges in, saying "this will be easy, because I've decided the thing I'm good at is more difficult than it", admits it's not easy, and either leaves or learns. Alternatively of course they may retreat into full Dunning-Kruger; see the Musk Twitter debacle, which is _both_, say.
Very complicated algorithms and mathematical proofs can still be understood by a single person, and be explored by a small number of people who all know each other. Brain surgery is done by a small team of people. These are typical "smart people" occupations.
Something as simple as Twitter still needs machinery that spans across technical skills, needs 24 hour monitoring, and needs lawyer and accountant support, so nobody can actually to it.
People think they can do it, because it's easy to spin up a demo that sends messages to a few thousand people and then shut it down again. They don't think about how to scan for CSAM, or how to respond to foreign government censorship requests.
WhatsApp was 55 people big when they got acquired, and to me that sounds about right.
Twitter employed 7,500 people. 7,500!!!! So please tell me where the complexity lies? Surely not in the front-end code I can tell you that.
Let's compare it to something WAY-WAY-WAY more complex, like a game with multiplayer, awesome mod tools, etc.: ROBLOX: 2,200 employees. Do I need to mention they wrote their own physics simulation engine and keeping realtime multiplayer going?
So please, explain this to me: how is Twitter more than 3 times more complex than Roblox???
Maybe I'm wrong, that's very possible, I've been wrong in the past. But just explain this 1 thing then: Twitter needs more than 3 times the manpower than Roblox?
Then nothing happened. At least, nothing that I personally observed as a casual Twitter reader. The goalposts were moved to "it will go down with the New Year's Eve spike", and once again nothing happened. Then the narrative became "the cracks will only be noticeable in a few months", and here we are and yet again, nothing.
So Musk and Geohot came out as the saner voices of that whole debacle. Of course Geohot said exaggerated things like "you only need 40 engineers to run Twitter", but if it turns out it takes 300 engineers, then I would consider this as Geohot being proven mostly right.
I don’t think that qualifies as “nothing happened” when features used in high-profile events fail, with the CEO and a potential future president left on the line. Any other platform wouldn’t have struggled with a stream of this size.
I guess you might say that’s just one thing, and other than the CEO’s live streams not working, everything is fine. But there are numerous other examples of accumulating paper cuts and failures at Twitter. I think this is close to what most of those doomsayers expected would happen.
https://mashable.com/article/google-ai-maps-search-event-bin...
> the AI falsely said the James Webb Space Telescope took the first ever picture of an exoplanet
> During the announcement about a new Lens feature, the demo phone was misplaced and the presenter wasn't able to show the demo
> Google seemed to say, "let's pretend this never happened," and immediately made the livestream recording private after the event
Are you sure ? Others say 6.5 M listened to the livestream that was delayed 20 mins
There was a lot of "ooh, it will catastrophically fail within weeks", which was fundamentally an assumption that the previous team was entirely incompetent. (Any halfway decent team tries their hardest to build resilient systems, not things that need hand-holding all the time.)
The current trajectory is exactly on the expected failure path predicted by anybody who does actually work on large systems - a steady increase of smaller failures, punctuated by the occasional large failure. (Cf. DeSantis announcement)
In essence, a reduction in staff will result in worse SLO results. It will result in less coverage of edge cases (technical and UX). Smaller teams are more constrained to travel on "the happy path". And the fact that marginal utility of additional engineers decreases means you can usually reduce teams a lot before impacting that path.
In complex systems, reductions also mean you're more vulnerable to a black swan event being irrecoverable, but that still requires a black swan first.
I don't think anyone argued Twitter was run by technically incompetent people. Where was this, if so? By leftists, yes, and by far too many people, yes. Both were argued repeatedly. But those things are now proven objectively true. The Twitter files showed just how systematic their enforcement of left wing orthodoxy was, and Musk fired most of the staff yet the site kept trucking and even launching new changes which is more or less the definition of having been over-staffed.
That's definitely more outrageous than saying that frontend is trivial. Whatever, I never took him seriously anyway.
https://en.wikipedia.org/wiki/Hilbert%27s_paradox_of_the_Gra...
Maybe something that pretends to be the real numbers, like a matrioshka doll of infinite containers inside containers.
Related: the Schröder–Bernstein theorem [4], "if there exist injective functions f : A → B and g : B → A between the sets A and B, then there exists a bijective function h : A → B.".
Not related, but great: Max Cooper (sound) and Martin Krzywinski (visuals) did a splendid job visualising "ℵ_2" [5].
[1] https://en.wikipedia.org/wiki/Cardinality_of_the_continuum
[2] https://en.wiktionary.org/wiki/%E2%84%B6
[3] "Cardinalities and Bijections - Showing the Natural Numbers and the Integers are the same size", https://www.youtube.com/watch?v=kuJwmvW96Zs
[4] https://en.wikipedia.org/wiki/Schr%C3%B6der%E2%80%93Bernstei...
[5] "Max Cooper - Aleph 2 (Official Video by Martin Krzywinski)", https://www.youtube.com/watch?v=tNYfqklRehM
As for the set of real numbers, we have the subset of irrational numbers which are uncountably infinite (see cantors diagonalization argument) thus making the whole set of real numbers, a set whose cardinality is ℵ_1.
The annotated turing book goes into this pretty well in the first couple pages.
It turns out the 'continuum hypothesis' can be true or it can be false. Neither contradicts standard ZFC set theory: the hypothesis is 'independent'.
[1] Using the concept of polycomputing from There’s Plenty of Room Right Here: Biological Systems as Evolved, Overloaded, Multi-Scale Machines: "Form and function are tightly entwined in nature, and in some cases, in robotics as well. Thus, efforts to re-shape living systems for biomedical or bioengineering purposes require prediction and control of their function at multiple scales. This is challenging for many reasons, one of which is that living systems perform multiple functions in the same place at the same time. We refer to this as 'polycomputing'—the ability of the same substrate to simultaneously compute different things, and make those computational results available to different observers.", https://www.mdpi.com/2313-7673/8/1/110
On the relative fringes, there are serious studies on alternative interpretations. See for example
https://en.wikipedia.org/wiki/Constructivism_(philosophy_of_...
(You can skip to the part that discusses Cantor's arguments, but I suspect that if you haven't heard about related concepts you probably want to understand what it is first.)
If I had a dollar for every job that I didnt get where I estimated the correct degree of difficulty, and they laughed, went with the person who said it would be easy and they could bang it out in a day of sleeping - I would be rich.
The loud optimist wins the 5.1 mil every time.
As soon as there is a dead serious one, I've noticed everyone get serious and starts ignoring the rainbows and unicorns people. If it's a slap on the write the rainbows still win.
It's likely there was some humor and bravado, as is the culture of his east coast origins.
The truth of his engagement with Twitter was, Just based on my watching him and his live streams during that time, that he was looking for a thing to do while he ceded control over comma AI, to a new executive leadership group.
...that doesn't change a fact there are some failures where developers really should know better and design it less shit
Of course, he should know better than to throw claims like that.
I'm glad that the company exists. I'm glad that smaller car manufacturers were exploring integrating the self-driving software, too.
And, from a celebrity media perspective, I think your engagement at Twitter was underreported especially in gossip mongering pseudo serious technology fan sites like this one.
It was an interesting shift for you, as you mature out of your first big startup into kind of a journeyman phase, in my view. (Tasting different experiences.) Something that is relatable to large numbers of us.
I hadn't laughed so hard in a long time.
It's not like he decided to hop from self-driving cars to 'AI' because fads changed.
1. https://geohot.github.io//blog/jekyll/update/2022/10/29/the-...
AI is not going anywhere. This is not a fad like some of the others mentioned but more likely than not where the next decade of innovation is built on.
The reason the company might fail is because their main thesis, being that car manufacturers would just license the self driving tech to somebody else (like Comma), never came about. Car manufacturers are just too conservative. It was a perfectly reasonable bet to make though. Unfortunately they ended up in the business of selling hardware and giving away software for free when they wanted to be in the business of selling software.
I'm Comma user in one of my cars as well, and I do like it. But, when was last time you tried built-in driving assist in a Tesla-priced car?
Tesla's driver assist is nothing special nowadays.
Things that Comma handles seamlessly that the built-in cruise in both cars will not:
- Full stop and go
- Sharp turns on the highway that require slowing down (both built-in adaptive cruise modes will gladly just drive you off a cliff at 65 mph)
- Situations where the lane lines are hard to see or are implied
- Non-highway driving
- Not requiring me to touch the steering wheel every 20 seconds
Maybe those things work in higher end cars (though I'd say the Ioniq is a fairly high-end car), but then again with Comma you get it for ~$2k in a ton of cars instead of having to buy a luxury car.
It is true that if you are on a highway, with clear lane lines, the steering assist in both cars is certainly a lot better than nothing, but it's just not nearly matching the reliability and versatility of Comma in any sort of imperfect situation.
In many countries doing this will void your insurance.
> - Sharp turns on the highway that require slowing down (both built-in adaptive cruise modes will gladly just drive you off a cliff at 65 mph)
While it's probably given that this will happen, it's also an infrastructural failure. Just place a limited speed limit sign way before the sharp turn, or fix the road so it doesn't make a sharp turn.
https://newsroom.porsche.com/en/2023/company/porsche-mobiley...
It takes guts to put out such bold bets in writing. We've seen many (senior!) tech people sneer at George's at times naive optimism. I actually find the "how hard could be" attitude refreshing against "no no it is complicated you can't do that" gatekeeping. Because otherwise we will end up using big-tech for lack of alternative. It is not the critic who counts and all that..
Cruse and Waymo have invested billions, they need 10’s of billions in annual sales or their project is a failure.
Is that it? SF only? After billions invested? There is a laundry list of those that tried and failed especially with burning an insurmountable amount of VC money even with billions of their own money.
Lyft: Scrapped and sold their self-driving project. [0]
Uber: Scrapped their robot-taxi project and sold it off. [1]
Zoox: Once valued at $3BN, acquired by Amazon for $1BN after nearly going bankrupt and is still using specialised cars for self driving only in SF. [2]
Cruise: Acquired by GM and still using specialised cars for self driving in SF [3]
Drive.ai: Ran out of money and almost bankrupt and acquired by Apple. [4] No where to be found on the roads.
Waymo: Same situation as Cruise, but Google keeping them alive.
Comma has lasted longer than these over-valued companies and is already in lots of consumer grade vehicles beyond SF today and not in specialised cars and taxis unlike Cruise and Waymo who are still stuck in SF [5].
[0] https://www.nytimes.com/live/2021/04/26/business/stock-marke...
[1] https://www.npr.org/2020/12/08/944337751/uber-sells-its-auto...
[2] https://www.cnbc.com/2020/06/26/amazon-buys-self-driving-tec...
[3] https://fortune.com/2016/03/11/gm-buying-self-driving-tech-s...
[4] https://techcrunch.com/2019/06/25/self-driving-startup-drive...
[5] https://techcrunch.com/2023/05/18/cruise-waymo-near-approval...
which is the critical element everyone is in denial about, even to the point of saying Tesla has a long way to catch up.
1. They are living in a sf USA big city centric bubble. 2. They are very easily influenced by marketing. 3. They are just trolling.
Comma is a product you can buy all code is opensource. All others is just a service, where people could theoretically just be diving remote and they sell it as self driving.
Whereas comma is taking the right approach of a nimble team, iterate fast and ship a working product (even if not L4-5), get cash flow, next milestone.
George is courageous, inspiring, and highly intelligent. He stands for what he believes in. He stands up for himself and his beliefs, and talks back to powerful people.
How many tries did it take to invent scalable electricity, or the light bulb?
George Hotz is frikkin awesome
1. Soliciting others to do his work - and offering internships to others that can help, in a kind of internship MLM (He was an intern at the time).
2. Complaining about how Twitter doesn't run on his laptop
I know it seems like ages ago at the current pace, but ya gotta remember that GPT-3 was released in mid-2020.
But not to everyone.
I would argue that it was not until around December 2022 that the world at large got the opportunity to begin really using AI directly. With ChatGPT.
What kind of comment is this on a site that used to be called Startup News? Even if that doesn’t resonate with you, isn’t what he’s talking about pure hacker ethos anyway?
Agree with this insight. One thing Nvidia got right was a focus on software. They introduced CUDA [1] back in 2007 when the full set of use cases for it didn't seem very obvious. Then their GPUs had Tensor cores, and more complementary software like TensorRT to take full advantage of them post deep learning boom.
Right as Nvidia reported insane earnings beat too [2]. Would love more players in this space for sure.
[1] - https://en.wikipedia.org/wiki/CUDA [2] - https://www.cnbc.com/2023/05/24/nvidia-nvda-earnings-report-...
The SW story is a train wreck, though. The problem basically was that they couldn’t hire any good SW people. As I said I know the founders. They are both genuinely decent guys, they put their own money in so they have some (well, minimal) skin in the game, and they know a ton of expert-level embedded and systems coders with between 20 and 40 years of hard core experience. As far as I can tell, they weren't really able to get anyone that we know in common to join. I certainly did not, and no one I know did either. Last I heard they'd had to hire third choice guys in Europe to do the work and it wasn't going well.
There's a pretty good reason for it, and it comes down to a sociological problem. HW people don’t value SW people. It's just basically true and has been true everywhere I've looked. Maybe if you're doing a system (like a router or maybe a drone) then the HW people will begrudgingly admit that the SW is a major part of the delivery, but that isn't true for chip companies (including chips-on-reference-boards).
You can rest assured that at a chip company, all of the high comp people in the company are going to be on the ASIC team and the SW team will never be on the same tier. The argument is always the same, no matter how many times it bites the companies on the ass and sends them careening into the dumpster: “yes, but the chip without SW is the chip! we can buy SW, if we have to. SW without the chip has zero value.”
Almost every chip company ends up like that, and the kind of low level, experienced SW people that work in the space know to avoid them and work at systems companies instead.
As far as I've been able to determine, with _maybe_ the exception of Cerebras - maybe - this is the situation that has played out at all of the 201x AI chip companies. They get founded by ASIC guys, most of whom have more than a small chip on their shoulder about the relative value of ASICs-vs-SW. These guys are all ex-SGI, ex-Sun, ex-Google, ex-Nvidia, ex-Intel HW guys who saw SW people making a lot more, not just in broader industry terms over the last few years, but at hardware-focused companies. In general, ASIC guys make less than SW guys unless they are the very narrow set of top level architects. IMHO from a value creation standpoint, that is _super unfair_ and I am not here to justify it, but it is how it is. The result poisons ASIC companies. SW people who know what needs to be don't won't go to them most of the time, for good reason, and so they fail.
So I will say, given that, starting with SW first is brilliant.
> In 2022, FuriosaAI remained the only startup to submit results in MLPerf Inference... This time, through purely enhancements in the compiler, our team was able to double the performance on the exact same silicon.
Maybe it helped they are based in South Korea. Other places to work in South Korea doing system programming is not very attractive.
Fair enough. I’m old so these conversations are never personal. I tried to help them by steering younger, less experienced but very high potential engineers toward them, but in the end they failed utterly to put together a viable SW team.
I don’t know any of their investors in terms that would let me ask, but I do know a bunch who passed, and most of them had concluded that it would end up just being more silicon on a crowded market. Being 10-20x better than nvidia isn’t the point if the market is about to be flooded with a dozen other chips against whom you are maybe 1.5-3x better. Without nailing the go to market needs, which means “make the stuff people have work, don’t make the customer learn new” etc. you have nothing. That’s all a Sw problem.
It’s actually worse, because the engineers in the space are actually pretty bad. A lot of what they have actually barely works to begin with, being a bunch of cobbled together python frameworks of dubious engineering quality and all of the hassle of the ecosystem. So the amount of mental space for “different” is almost zero even if you ignore that they’ve been burned (AMD) before, which you can’t.
The past decade or so, they haven't been able to create any good software for their hardware. They made small improvements but the competition, Nvidia, has also made improvements to their already good software.
It too the point where their software is the reason why most people/companies don't use their products. Their drivers for their customer products are just as bad.
They are very competitive in hardware, but Nvidia dominates them at software which make companies buy Nvidia. No one wants to deal with the pain of AMD software.
AMD is a better company to work with than Nvidia, but it not worth it when it comes to dealing with their software lol.
Look, the reality of software is that it’s actually dozens of markets, from random bloated web crap to building safety-focused critical systems. Each of those areas values different skills and has different knowledge.
But I can attest that AMDs problems - and for that matter, Intel’s since I am familiar with both - is that the companies see software outside of a very few niches as an afterthought and relatively a cost of doing business rather than a value.
It should in fact amaze you that Nvidia has pulled off being dominant. Jensen is pretty well known to have relatively poor understanding of software and Nvidia for a long time had a reputation exactly as I describe in one of my other comments. But Nvidia stock appreciation made it possible for them to become an attractive destination despite their core corporate mentality and accidentally created a company that ended up with a strong software culture.
AMD could solve this problem tomorrow with the right level of investment and a willingness to bulldoze from above the corporate politics that prevent them from having a markwt-leading software team.
Yes, it is hard. I have worked in companies like that before. I am not just saying “hey, go hire some rockstar coders” because that answer is ALWAYS bullshit. Software people are especially prone to thinking that’s the answer and that isn’t what I am saying.
But there are ways for companies in their situation to structure things to make it work out. The specifics are not appropriate for a public forum as it would identify a number of former employers who were successful in fixing their issues.
They tried at some point. Like another commenter pointed out elsewhere HW people just don't care about SW. They think HW is superior and SW is the joke. I don't think much of the culture has changed.
A lot of their understanding = workaround the HW, make it work, etc.
I haven't confirmed, but I strongly assume that either their graphics drivers or something Ubuntu does with Wayland are not fine.
In this specific case, I was unable to get to a console and had actual things to get done so I wasn't going to debug it.
So in the gaming market, Nvidia still commands a huge premium. No AMD card can play Cyberpunk on overdrive 4k.
IMHO, unless you have a $1000+ GPU, RayTracing is still not worth the performance hit. I prefer playing at 100+ FPS with RayTracing off, than turning it on and have my frame rate cut in half.
This makes very little sense. Even if he was able to achieve his goals, consumer GPU hardware is bounded by network and memory, so it's a bad target to optimize. Fast device-to-device communication is only available on datacenter GPUs, and is essential for models training like LLaMA, Stable Diffusion, etc. Amdahl's law strikes again.
The quality, arrogance and ignorance of enterprise SDLC was akin to that of a 1st year grad.
Many of us would absolutely be just as contrarian towards general society as he is, and idolize contrarian figured like Musk.
Doesn't mean that the work he is doing is not valid. It's a shame though because his ideology is very likely to hinder the progress in his companies (like for example policy against remote work).
I suspect they are going to be needed for quite some time to come.
https://spectrumnews1.com/ap-top-news/2020/10/08/waymo-remov...
https://www.azfamily.com/2022/08/29/waymo-launches-its-self-...
There was also a waymo press released this year that stated all rides in Arizona are now backup driver free, although I can't find it.
It is like it is wrong to stop working on one idea and start working on the other.
You need to be a team player and work across different groups e.g. Product, Testing, SRE etc in order to successfully get features into Production.
So being a talented engineer is useful but having high emotional intelligence and being able to negotiate and collaborate is far more important.
For such a smart guy, locking yourself out of a ton of talent by requiring software developers to be on-site in 2023 seems...out of character, to put it politely.
(Rephrased, my original post was a bit too ad hominem and accumulating downvotes rapidly. I wanted to delete this entire comment but apparently HN no longer allows comments to be deleted.)
There is a high chance that you are probably neurodivergent to some extent.
So instead of WFH, where you can remove distractions and work on your own time, you are now forced to abide by someone else's schedule, take time commuting, e.t.c
In office work is for people/positions that require hands on work with hardware, or you are hiring for replaceable positions where people don't have dedication to the cause and are going to do as little work as possible for the same pay. Tinygrad is neither.
I doubt they need mass volumes of employees at this stage and they maybe want to work closely with the people they choose?
Remote work is available to everyone on GitHub. If you submit a bunch of good PRs and show me you are easy to work with, I'm down to pay per project.
Source https://twitter.com/realGeorgeHotz/status/166153013618397184...
I'm not anti remote, I'm anti full time remote. It's hard to build a culture
In the same way Comma went from "Our goal is to solve AI, Comma body is the next big thing" to George peacing out because now they are just doing the busy work to make more money.
If he's anti full-time remote, then his pool of candidates is still limited to those who live in San Diego or very close to it.
I mean a lot of smart people seem to do their hacking by themselves. I'm thinking like Fabrice Bellard. This is at least a step beyond that.
> One Humanity is 20,000 Tampas.
I'll never think of humanity the same way!
[0]: https://geohot.github.io/blog/jekyll/update/2023/04/26/a-per...
> I promise it’s better than the chip you taped out! It has 58B transistors on TSMC N5, and it’s like the 20th generation chip made by the company, 3rd in this series. Why are you so arrogant that you think you can make a better chip? And then, if no one uses this one, why would they use yours?
> So why does no one use it? The software is terrible!
> Forget all that software. The RDNA3 Instruction Set is well documented. The hardware is great. We are going to write our own software.
So why not just fix AMD accelerators in pytorch? Both ROCm and pytorch are open sourced. Isn't the point of the OSS community to use the community to solve problems? Shouldn't this be the killer advantage over CUDA? Making a new library doesn't democratize access to the 123 (fp16-)TFLOP accelerator. You fix pytorch and suddenly all the existing code has access to these accelerators. Millions of people now have This then puts significant pressure on Nvidia, as they can't corner the DL market. But it is a catch-22 because the DL market already is mostly Nvidia so it takes priority. Isn't this EXACTLY where OSS is supposed to help? I get Hotz wants to make money, and there's nothing wrong with that (it also complements his other company), but the arguments here seem more for fixing ROCm and specifically the pytorch implementation.
The mission is great, but AMD is in a much better position to compete with AMD. They caught up in the gamer's market (mostly) but have a long way to go for scientific work (which is what Nvidia is shifting focus to). This is realistically the only way to drive GPU prices down. Intel tried their hand (including in supercomputers) but failed too. I have to think there's a reason that's not obvious to most of us as to why this is happening.
Note 1:
I will add that supercomputers like Frontier (current #1) do use AMDs and a lot of the hope has been that this will fund the optimization from two places: 1) DOE optimizing their own code because that's the machine that they have access to and 2) AMD using the contract money to hire more devs. But this doesn't seem to be happening fast enough (I know some grad students working on ROCm).
Note 2:
There's a clear difference in how AMD and Nvidia measure TFLOPS. techpowerup shows AMD at 2-3x Nvidia, but performance is similar. Either AMD is crazy underutilized or something is wrong. Does anyone know the answer?
> It's pretty hard to really beat NVIDIA for developer support though, they've invested a lot of work into the CUDA ecosystem over the years and it shows.
Yup, strong CUDA community and dev support. That said, more ergonomic domain specific languages like Mojo might finally give CUDA some competition though - it's still a very high bar for sure.
When we're presented with problems where the two potential answers are "it's a lot harder than it looks" and "the people working on it are idiots" I tend to lean towards the former. But hey, when it is the latter there's usually a good market opportunity. Just I've found that domain expertise is seeing the nuance that you miss when looking at 10k ft.
The change in the landscape I see now is that the models are big enough and useful enough that the commercial appetite for inference is expanding rapidly, hardware supply will continue to be constrained, and so tools that can reduce production inference cost by a percentage are starting to become a straight forward sale (and thus justify the infrastructure investment). This is not based on any inside info but when I look at companies like Modular and Octo that's a big part of why I think they probably will have some success.
Because there's no real evidence that AMD cares about this problem, and without them caring your efforts may well be replaced by whatever AMD does next in the space. Their Brooks language[1] is abandoned, OpenCL doesn't compare well, ROCm is like the Sharepoint of GPU APIs (it ticks boxes but doesn't actually work very well).
> So why not just fix AMD accelerators in pytorch
Why not just buy NVidia? They care deeply about the space, will actually help you if you have trouble, etc etc.
Even using Google TPUs is better: Google will help you too.
While everyone using NVidia isn't great for the market as a whole as an individual company or person it makes a lot of sense.
Read "The Red Team (AMD)" section in the linked article:
> The software is called ROCm, it’s open source, and supposedly it works with PyTorch. Though I’ve tried 3 times in the last couple years to build it, and every time it didn’t build out of the box, I struggled to fix it, got it built, and it either segfaulted or returned the wrong answer. In comparison, I have probably built CUDA PyTorch 10 times and never had a single issue.
This is geohot. He knows how to build software, and how to fix problems.
Note that "Our short term goal is to get AMD on MLPerf using the tinygrad framework."
> There's a clear difference in how AMD and Nvidia measure TFLOPS. techpowerup shows AMD at 2-3x Nvidia, but performance is similar. Either AMD is crazy underutilized or something is wrong. Does anyone know the answer?
From the linked article:
> That’s the kernel space, the user space isn’t better. The compiler is so bad that clpeak only gets half the max possible FLOPS. And clpeak is a completely contrived workload attempting to maximize FLOPS, never mind how many FLOPS you get on a real program
This is a nonsequitor. This feels like when my uncle learns that I know how to program he asks me to build a website. These are two different things. I do ML and scientific computing, I'm not your guy. Hotz is a wiz kid but why should we expect his talents to be universal? Generalists don't exist.
And we're talking the guy who tweeted about believing that the integers and reals have the same cardinality right? Between that and his tweets on quantum we definitely have strong evidence that his jailbreaking skills don't translate to math or physics.
He's clearly good at what he does. There's no doubt about that. But why should I believe that his skills translate to other domains?
STOP MAKING GODS OUT OF MEN. Seriously, can we stop this? What does Stanning accomplish? It's creepy. It's creepy if it is BTS, Bieber, Elon, Robert Downey Jr, or Hotz.
> Read "The Red Team (AMD)" section in the linked article:
Clearly I did, I quoted from it. You quoted from the next section (So why does no one use it?).
You definitely shouldn't trust what geohot says about infinitary mathematics or (god forbids) quantum mechanics. On the other hand, you generally should trust what he says about machine learning software stack.
This project isn't even about skill in ML, which demonstrates misunderstandings. The project requires writing accelerator code. Go learn CUDA and tell me how different it is. It isn't something you're going to pick up in a weekend, or a month, and realistically not even a year. A lot of people can write kernels, not a lot of people can do it well.
If you haven't worked with someone who's smarter and more motivated than you are, then I can see how you'd draw that conclusion, but if you have, then you'd know that there are full stack developers out there who do both better than you. It's humbling to code in their repos. I've never worked with geohot so I don't know if he is such a person, but they're out there.
No of course not. But this is literally his field of expertise, and there's plenty of reasons to think he knows what he is doing. Specifically, the combination of reverse engineering and writing ML libraries means I'd certainly expect he's had reasonable experience compiling things.
RDNA 3 has dual-issue that basically isn't used by the compiler so half the FPUs are idle.
It doesn't fit the business model. I mean sure, they'll sell AMD computers now like a bootleg Puget Systems. But why buy from the bootleg when I can just buy from the real thing (or AWS) and run tinygrad on there if I want?
So the play is, get people using your framework (tinygrad), then pivot to making AI chips for it:
> In the limit, it’s a chip company, but there’s a lot of intermediates along the way.
Seems far fetched but good luck to them.
Memory bandwidth (sparsity helps) and networking connectivity (Nvidia bought Mellanox and other networking companies) are important too. They are also using a lot of die space on raytracing stuff that they don't waste on the datacenter versions presumably.
Intel's software is much better (MKL, vtune, etc for GPU) and getting better.
Where is this number coming from? The number of spikes per second?
Edit: doing a quick search, it doesn’t seem like there’s a consensus on the order of magnitude of this. Here’s a summary of various estimates: https://aiimpacts.org/brain-performance-in-flops/
Used 20 PFLOPS of compute to simulate it.
My prediction is AMD is already working on this internally, except more oriented around PyTorch not Hotz's Tinygrad, which I doubt will get much traction.
https://pytorch.org/blog/pytorch-for-amd-rocm-platform-now-a...
>The software is called ROCm, it’s open source, and supposedly it works with PyTorch. Though I’ve tried 3 times in the last couple years to build it, and every time it didn’t build out of the box, I struggled to fix it, got it built, and it either segfaulted or returned the wrong answer. In comparison, I have probably built CUDA PyTorch 10 times and never had a single issue.
I had the same experience ~3 months ago. Gave up and switched to Nvidia 3090s for my workloads.
The parent post is surprised that they still aren't making the appropriate investments to make it work. They kind of started to do that a few years ago, but then it fell on the wayside without reaching even table stakes, which in my opinion would require providing a ROCm distribution that works out of the box for most of their recent consumer cards (i.e. those cards which the enthusiasts/students/advocates/researchers might use while choosing which software stack to learn, and afterward base corporate compute cluster purchasing decisions on whether they support the software they wrote for e.g. CUDA+Pytorch), and they seem to be failing at that.
AMD have a lot of money to lose.
George Hotz $5M OSS company - well, not so much.
To be fair, the previous page has a bit more details on the hardware.
I can run LLaMA 65B GPTQ4b on my $2300 PC (built from used parts, 128GB RAM, Dual RTX 3090 @ PCIe 4.0x8 + NVLink), and according to the GPTQ paper(§) the quality of the model will not suffer much at all by the quantization.
Just saying, open source is squeezing an amazing amount of LLM goodness out of commodity hardware.
I looked into this previously[1] but wasn't super confident it's possible or what hardware is required (2x x8 pcie and official SLI support?). AFAICT still would look like two GPUs to the system.
[1] https://discuss.pytorch.org/t/is-there-will-have-total-48g-m...
The GPUs have gotten improved thermal pads. The 8 core Ryzen 3700X is a 65W model. It appears to be fast enough not to be a bottleneck for this purpose. 1200W PSU.
I may swap the fan in front of the GPUs for a high-rpm model.
Also for longer runs i throttle the GPU power draw. It doesn't cost much performance.
He made a very good point about how this isn’t general purpose computing. The tensors and the layers are static. There’s an opportunity for a new type of optimization at the hardware level.
I don’t know much about Google’s TPUs, except that they use a fraction of the power used by a GPU.
For this experiment though, my sincere hope is that all the bugs are software only. Supporting argument - if they were hardware bugs, the buggy instructions would not have worked during gameplay.
And then we get a computer that... how do I interact with it? Will it have its own OS? Some flavor of linux? Is the intent to work on it directly, or use it as an inference server, and talk over a network?
Sure they could just fork it and try to continue development themselves, but the community and momentum might very well not go with them.
$15k seems like a kind of low - if you've tried to spec out a server with similar quantities of accelerators recently, you'd have trouble hitting that figure, even if using consumer grade GPUs.
Most people just would never take that risk and will stick to their well-paying job.
I took my one shot! Any more is too risky for quite a while.
Can someone educate me why that is the case? Does `x[y]` require a Turing-complete kernel to compute?
https://tenstorrent.com/research/tenstorrent-raises-over-200...
The tinybox
738 FP16 TFLOPS
144 GB GPU RAM
5.76 TB/s RAM bandwidth
30 GB/s model load bandwidth (big llama loads in around 4 seconds)
AMD EPYC CPU
1600W (one 120V outlet)
Runs 65B FP16 LLaMA out of the box (using tinygrad, subject to software development risks)
$15,000
I have no idea what I am doing but here goes!
By [1] we have 156 FP16 TFLOPS, taking their non "*" (* = with sparsity) value. So you need 5 So $40,000? pls the other stuff, and someone to make a profit putting it together say $50,000?
So this setup is 3 times cheaper for the same.
If I am allowed to use the sparsity value it is 1.5 times cheaper.
[1] https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
It's just too easy for anyone to throw together a Supermicro machine with 6x GPUs in it, which is what it sounds like they'll be doing.
My guess is they'll end up creating some premium extensions to the software and selling that to make money. Or maybe they can sell an enterprise cluster manager type thing that comes with support. He's good at software so it makes sense for him to sell software.
And maybe the box will sell well initially just as a "dev kit" type thing.
Have you seen what a DGXA100 costs? It starts at $199k for 8 40GB A100's, which have a list price of $10k each. So the GPU costs are $80k. What do you get for the extra $120k? 1TB ram, 2 2TB NVMe OS drives, 4 4TB NVME general storage, and 8x200Gbit infiniband. I would guess no more than 20k all of the remaining hardware. So that's a ~$100k computer selling for $200k. And that's with NVDA likely making massive margins already on the A100 and the Infiniband hardware.
The reality is that companies want to buy complete solutions, not to build and manage their own hardware. A $15k a computer that's $10k in parts is not a large markup at all for something like this.
NVIDIA's advantage is that they're a proprietary company and they're the ones actually making the chips they're putting in a box.
That's very far away from a random little open source startup slapping third-party GPUs in a generic box.
Price: $15,000.
If they had a "lite" model that sold for $1500, and were actually shipping....
HPC compute is well advanced past just slapping GPUs into generic supermicro servers anyway. Without semi-custom hardware and equivalents to nvlink/nvswitch AMD won't ever be competitive in the HPC space.
Don't waste your money.
Buy 6 RTX 4090's and a decent ECC-memory server, and call it a day.
> Unfortunately, this advantage is thrown away the minute you have something like CUDA in your stack. Once you are calling in to Turing complete kernels, you can no longer reason about their behavior. You fall back to caching, warp scheduling, and branch prediction.
> tinygrad is a simple framework with a PyTorch like frontend that will take you all the way to the hardware, without allowing terrible Turing completeness to creep in.
I like his thinking here, constraining the software to something less than Turing complete so as to minimize complexity and maximize performance. I hope this approach succeeds as he anticipates.
Let's say you want to run LLaMA. LLaMA is a tiny amount of code, say, 300 lines. LLaMA is static. It doesn't matter people will implement LLaMA with PyTorch and not tinygrad, geohot can port LLaMA to tinygrad himself. In fact, he already did, it's in tinygrad repository.
What I am saying is while running all models ever invented is harder than running LLaMA and Stable Diffusion (Stable Diffusion port is also in tinygrad repository), that's not necessarily trivializing the problem. It is noticing that you don't need to solve the full problem, there is enough demand for solving the trivial subset.
While developers will choose usability, users will choose cheap price. If they can run what they want on cheaper hardware, they will. I already have seen this happening: people don't buy NVIDIA to run Leela Chess Zero, they just run it on their hardware. It doesn't matter everyone working on LC0 model is using NVIDIA, that's irrelevant to users. LC0 model is fixed and tiny, people already ported the model to OpenCL, OpenCL port is performant, it runs well on AMD. The same will happen to text and image generation models.
I recall reading about avoiding Turing completeness for similar reasons to avoid the halting problem.
> Other times these programmers apply the rule of least power—they deliberately use a computer language that is not quite fully Turing-complete. Frequently, these are languages that guarantee all subroutines finish, such as Coq.
Shows you what is possible in 2.5 years. Keeps me motivated to learn.
Even if adopting "hardware/software co-design", leadership might be hardware people, and they might not understand that there's tons more to systems software engineering than the class they had in school or the Web&app development that their 5 year-old can do. That misunderstanding can exhibit in product concepts, resource allocation, scheduling, decisions on technical arguments, etc.
(Granted, the stereotypical software techbro image in popular culture probably doesn't help the respect situation.)
I'm always grateful for his user feedback, it led to next-level improvements. Thanks geohot
The reason they sometimes win is huge structural or incentive issues in the large companies. And they don't win often.
AMD has failed to fix the problem for years. Is it because the business structure doesn't incentivize it? Is it because the company's entrenched culture is opposed to compensation that might attract the right talent (often out of "fairness")? Is it because there's some internal owner for the function that keeps fucking up but for political reasons the CEO won't replace them and no one else can work on the thing?
Any of these are possible. I have seen - personally witnessed - all of these and more at large companies. We don't know the reason, but we can sort of guess as to the shape of it.
It’s really as simple as that and it still hasn’t changed so nvidia is dominating them in AI as a result.
They should've been adding tensor cores and neural acceleration to their CPUs. The need for headed graphics cards is moot and wasteful. NVIDIA solved this with the A100.
NVIDIA may spin into a mainstream enterprise CPU and systems vendor as a sales channel for converged CPU-GPU solutions beyond what they're already doing.
Of course nothing is perfect and you can never have 100% trust to someone else hardware, but it's defenetely step in right direction.
That being said I am still sceptical.
if you're a research scientist or grad student, to a certain extent a lot of projects are "greenfield" so it's easy to jump on a new framework if it is nice to use and offers some advantage.