That's worth 100M. And they won't even send us 2 ~100k boxes. In what world does that make sense, except in a world where decisions are made based on pride instead of ROI. Culture issue."
That's worth 100M. And they won't even send us 2 ~100k boxes. In what world does that make sense, except in a world where decisions are made based on pride instead of ROI. Culture issue."
Take the free offer, prove everyone wrong and then start to tell us how great you are. https://x.com/HotAisle/status/1880507210217750550
He picked his problem better. The whole reason that tinygrad is, well, tiny, is that it limits the amount of overhead to onboard people and perform maintenance and rewrites. My strong impression is that the ROCm codebase is simply much too large for AMD's dev resources. You're trying to race NVidia on their turf with less resources. It's brave, but foolish.
I can see how Tinygrad could succeed. The story makes sense. AMD's doesn't, neither logically nor empirically. NVidia would have to seriously fumble.
Worked for AMD in the CPU market.
That said I'm deeply worried about anyone whose based their company on amd gpus. The only reason why they do well in hpc is because there's an army of dreadfully underpaid and over performing grand students to pick up the slack from AMD. Trying to do that in a corporate environment is company suicide.
> Chiplets are also enabled by TSMC technology, CoWoS.
Interesting, my mistake. Thank you for pointing that out!
Sony Interactive and Microsoft XBox seem to be doing great without an army of underpaid students. AMD does great at the top and bottom: the corporates in the middle that are unwilling or unable to pay people to author/tweak their software for AMD GPUs will do better going with Nvidia, which has great OOTB software, and a premium to go with it.
I suppose if AMD had infinite resources, it'd fix this post-haste.
That's the entire point of AMD partnering with larger companies, rather than going all-in with consumers and small startups at this point in time.
This would end up costing maybe tens of millions at most, but the potential return is indeed measured in billions.
And yep, lots of people like geohot are (to put it mildly) eccentric. So deal with it. They are not merely your customers, they are your freaking sales people.
As it is, I work in a startup that does a bit of AI vision-related stuff. I'm not going to even touch AMD because I don't want to deal with divas on the AMD board in future. NVidia is more expensive right now, but they're far more predictable.
Do you really want all AI hardware and software dominated by a monopoly? We're not looking to "beat" Nvidia, we are looking to offer a compelling alternative. MI300x is compelling. MI355x is even more compelling.
If there is another company out there making a compelling product, send them my way!
> It most definitely is about “beating” NVIDIA.
Hard disagree, but we are just going to have to agree to disagree on that.
I'm willing to try AMD, and I even built an AMD-based machine to experiment with AI workflows. So far it has been failing miserably. I don't care that MI300X is compelling when I can't make samples work both on my desktop and on a cloud-based MI300X. I don't care about their academic collaborations, I'm not in the business of producing papers.
I'll just pay for H100 in the cloud to be sure that I will be able to run the resulting models on my 3090 locally and/or deploy to 4090 clusters.
If AMD shows some sense, commits to long-term support for their hardware with reasonable feature-parity across multiple generations, I'll reconsider them.
And AMD has a history of doing that! Their CPU division is _excellent_, they are renowned for having long-term support for motherboard socket types. I remember being able to buy a motherboard and then not worrying about upgrading the CPU for the next 3-4 years.
Anush was actively looking for feedback on this on github today...
https://www.reddit.com/r/ROCm/comments/1i5aatx/rocm_feedback...
All AMD had to do was support open standards. They could have added OpenCL/SYCL/Vulkan Compute backends to Tensorflow and Pytorch and covered 80% of ML use cases. Instead of differentiating themselves with actual working software, they decided to become an inferior copy of NVIDIA.
I recently switched from Tensorflow to Tinygrad for personal projects and haven't looked back. The performance is similar to Tensorflow with JIT [0]. The difference is that instead of spending 5 hours fixing things when NVIDIA's proprietary kernel modules update or I need a new box, it actually Just Works when I do "pip install tinygrad".
0: https://cprimozic.net/notes/posts/machine-learning-benchmark...
So it is all shit, but tinygrad saves the day?
I don't know of any other autograd libraries with a non-CUDA backend, but I'd be interested to learn about them.
That doesn't help if the drivers are buggy. AMD needs to send hardware to their own driver developers.
However it would also raise future revenue, which should be what's reflected by the market.
So it would still be something that's good for the company, but not nearly 100B good.
And how's that been going? The AMD stock price compared to NVidia seems to speak volumes about the efficacy of these projects.
IREE has been around for 5 years, without producing anything overtly practical. They seem to be focused more on academic jobs and citations. It's also focused on the general case of a compiler for "all" AI-type tasks, supporting everything from WASM to CUDA.
OpenXLA seems to be a bit more practical, but I spent the last 2 hours trying to make it work on my AMD card (Radeon Pro W7900) and failing.
I personally don't like Tinygrad's approach of doing their own thing rather than integrating into PyTorch/JAX/..., but it at least is _practical_ with a reasonable end-goal. Is it going to be successful? Who knows. But it's more practical than anything AMD has done within the recent 5 years.
Those academic publications are a sign that the people involved actually know what they’re doing, and are making sure their work holds up to scrutiny.
I've been hearing about MLIR and OpenXLA for years through Tensorflow, but I've never seen an actual application using them. What out there makes use of them? I'd originally hoped it'd allow Tensorflow to support alternate backends, but that doesn't seem to be the case.
0: https://cprimozic.net/notes/posts/machine-learning-benchmark...