Also, for me, the answer is no: they're not any good. I use PyTorch. I use custom CUDA code. The reality is brutal but simple: if it doesn't run CUDA, no serious ML research will use it for training anything else than toy models.
Also, for me, the answer is no: they're not any good. I use PyTorch. I use custom CUDA code. The reality is brutal but simple: if it doesn't run CUDA, no serious ML research will use it for training anything else than toy models.
They will use it to edit code, deploy and remotely run jobs. Serious deep learning on laptops is a non-starter, on any laptop, on account of heat dissipation.
Training Tesla's FSD neural net on a laptop? A student training some models for courses or self study?
Serious schools have computing clusters. ML researchers might be interested to know which undergrad is actually training models complex enough to benefit from better hardware.
Being able to do light stuff on one's own consumer hardware without having to buy something new is still incredibly helpful to students and other people trying to learn, as well as hobbyists though.
> these still aren’t machines made for deep learning. Don’t get me wrong, you can use the MBP for any basic deep learning tasks, but there are better machines in the same price range if you’ll do deep learning daily.
The article has 4 short sentences in the 'conclusion' section which can be found, as expected, at the end of the article.
it really isn't "buried"
“But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.”
I agree, but from what I have seen they also perform poorly against collab. A 1660 Ti 4gb RAM laptop gpu, runs about as fast as Collab CPU does. [1]
I have not been able to find how the RTX 2080 Super (Laptop) and upcoming RTX 3080 (Laptop) compares to Collab. The 3080 has over 3x as many cores as the 1660 Ti.
[1] https://towardsdatascience.com/google-colab-how-does-it-comp...
According to the chart on that page you linked to the Lenovo Legion (with the 1660ti) is about 30% faster than the Colab GPU. The Lenovo 480s (CPU only, no GPU) is about the same speed as the Colab CPU. So if you can train your model on the on 4GB of VRAM, a laptop with a discrete GPU (if you exclude the really basic business laptop GPUs) may be useful.
It can come with a RTX 2060 (Max Q / low voltage) which supports CUDA; as well as Ryzen 9 4900HS, which was the fastest mobile CPU until the 5000 series mobile Ryzen came out.
The 2060 + 4900HS has gone on sale at Best Buy for $1199 multiple times late last year (just check slickdeals).
I think it’s absolute steal of a machine for that price. Not to mention that it’s beautifully built, is lightweight, has an amazing matte color accurate screen, etc.
There’s also a fan sub for this particular laptop with over 13k subscribers (as of today) on Reddit: https://www.reddit.com/r/ZephyrusG14/
However, if your goal is just to have a machine on which you can actually test your CUDA code, then the Zephyris G14 laptop is perfect. It's portable, has a nice screen, etc.
It's "optional" in the sense that things still calculate correctly on CPU without it, but at a 1000x performance penalty. Or you could skip it if you had 64GB of GPU RAM, which you cannot buy (yet).
So if you actually want to work with this on GPUs that are commercially available, you need it.
Are there any examples where custom cuda code implements some op that can't be written in Pytorch/TF/Jax/etc? That would have provided a better support to your claim that M1 needs to be able to run cuda.
Note the cuda kernels in the original repo were added in August 2017. It might have been the case at the time they needed them, but again, if you need to do something like that today, you're probably an outlier. Modern DL libraries have a pretty vast assortment of ops. There have been a few cases in the last couple of years when I thought I'd need write a custom op in cuda (e.g. np.unpackbits) but every time I found a way to implement it with native Pytorch ops which was fast enough for my purposes.
If you're doing DL/CV research, can you give an example from your own work where you really need to run custom cuda code today?
(downvoted for asking a honest question. Ahhh the things you see on Apple related posts... Emotion driven bunch)
Its like Apple sell you the iPhone, but also the iOS APIs and application model so you can run iOS apps. Once you run iOS apps, you are in Apple's ecosystem, both the ISVs and the user are hard to leave. Its like saying when will Apple officially make iOS APIs and libraries run on Android phones. Both will never happen.
And getting email to the internet at large was no mean feat: I remember having to do a slew of UUCP addressing to get an email from AT&T to the internet (something along the lines of "astevens@redhill3!ihnp4@mit.edu). It was the wild west.
Researchers want to get their work done. They don't want to fight against their tools.
There's also the issue that by relying on proprietary frameworks that work now, you might be painting yourself into a corner if Nvidia changes something in the future and then you have to adapt to them because you have no choice.
When nvidia sees a need, they can change CUDA over night to address it, and they pay people to do that.
When you need to do the same in Vulkan, that’s a multi year process till your extension is “open”. A teaser here whose job is to get something done with ML has better things to do than going through that.
oneAPI DPC++ Features Included in SYCL 2020 Final Spec [https://newsroom.intel.com/articles/oneapi-dpc-features-2020...]
Arstechnia's write up on OneAPI provides a good overview [https://arstechnica.com/gadgets/2020/09/intel-heidelberg-uni...]
[edit: added arstechnia reference & link]
GPGPU has been a thing for 20 years now, and I still can't easily write code that works on Nvidia and AMD and ship it to consumers on Windows. From what I've seen OpenCL seems to be dying, AMD doesn't care about compute on Windows or on their Radeon cards, and Cuda continues to be the only real option year after year in this growing segment. Why would anyone buy anything else than Nvidia if they're using Photoshop, Blender, DaVinci Resolve, or other compute heavy consumer software? Maybe it's unrealistic to hope that any library can fix this and just rename GPGPU to Nvidia compute and be done with it.
If you're using Blender it can absolutely make a ton of sense to use AMD hardware. For most of the time where Blender supported GPGPU AMD was the best choice, and I set up rendering servers with AMD hardware for that express purpose.
I feel like a big part of this attitude is from not actually having tried it. Because SYCL works fine on AMD. In fact, you have more backend options for AMD than NVidia.
If you control your own hardware and software stack maybe an AMD CDNA card is fine, but if you want to ship software to end users it seems to be difficult to even know what will work. So you use cross platform code for a worse experience on Nvidia and spotty support on AMD, or only Cuda and accept that it's Nvidia only but will give you a better experience.
I haven't done a lot of GPGPU programming, but I've tried to look at it from time to time, and I've been disheartened by it every time. Nvidia's handling of OpenCL, AMD's disregard for SPIR. This is what an AMD representative had to say in 2019 [2]:
"For intermediate language, we are currently focusing on direct-to-ISA compilation w/o an intervening IR - it's just LLVMIR to GCN ISA. [...] Future work could include SPIRV support if we address other markets but not currently in the plans."
[0] https://github.com/RadeonOpenCompute/ROCm/issues/1180#issuec...
[1] https://community.amd.com/t5/opencl/spir-support-in-new-driv...
[2] https://github.com/RadeonOpenCompute/ROCm-OpenCL-Runtime/iss...
While developing a small competitor to Tensorflow back then (Leaf), we were one of the few frameworks that also tried to support OpenCL, but the additional dev work made it unfeasible.