Introduction to CUDA programming for Python developers
pyspur.dev
pyspur.dev
Parallel question: I work as a Data Engineer and always wonder if it's possible to get into MLE or AI Data Engineering without knowing AI/ML. I thought I only need to know what the data looks like, but so far I see every job description of an MLE requires background in AI.
A "background in AI" is a bit silly in most cases these days. Everyone is basically talking about LLMs or multimodal models which in practice haven't been around long. Sebastian Raschka has a good book about building an LLM from scratch, Simon Prince has a good book on deep learning, Chip Huyen has a good book on "AI engineering". Make a few toys. There you have a "background".
Now if you want to really move the needle... get really strong at all of it, including PTX (nvidia gpu assembly, sort of). Then you can blow people away like the deep seek people did...
Karpathy's teaching style is well, Karpathy, Raschka is more conventional (but not buttoned down).
How can I leverage that experience into earning the huge amounts of money that AI companies seem to be paying? Most job listings I've looked at require a PhD in specifically AI/math stuff and 15 years of experience (I have a masters in CS, and no where close to 15 years of experience).
Most companies aren't doing a lot of heavy GPU optimization. That's why deepseek was able to come out of nowhere. Most (not all) AI research basically takes the given hardware (and most of the software) stack as a given and is about architecture, loss functions, data mix, activation functions blah blah blah.
Speculation - a good amount of work will go towards optimizations in future (and at the big shops like openAI, a good amount already is).
I'd think things like optimizing for occupancy/memory throughput, ensuring coalesced memory accesses, tuning block sizes, using fast math alternatives, writing parallel algorithms, working with profiling tools like nsight, and things like that are fairly transferable?
for a deeper dive, check out the sth like Georgia Tech’s CS 8803 O21: GPU Hardware and Software.
To get into MLE/AI Data Engineering, I would start with a brief introductory ML course like Andrew Ng’s on Coursera
At the very least you should know enough linear algebra that you understand scalar, vector and matrix operations against each of the others. You don't need to be able to derive back prop from first principles, but you should know what happens when you multiply a matrix by a vector and apply a non-linear function to the result.
The little learner comes close but I'd only really suggest that to people who already know the maths because the presentation is very non-standard and can get very misleading.
If you're interested drop me a line on my profile email and I'll have a look at some numerical algebra books and papers to see what's out there.
Anyway appreciate the help.
Neural networks are basically just linear algebra (i.e matrix multiplication) plus an activation function (ReLu, sigmoid, etc.) to generate non-linearities.
Thats first year undergrad in most engineering programs - a fair amount even took it in high school.
You don't really need even need to know about determinants, inverting matrices, Gauss-Jordan elimination, eigenvalues, etc. that you'd get in a first year undergrad linear algebra
You could also try to recreate or modify a shader you like from https://www.shadertoy.com/playlist/featured
You'll inevitably pick up some of the math along the way and probably have fun doing it along the way.
Some basic areas to focus on:
- Setting up the architecture and config
- Learning how to write the kernels, and what makes sense for a kernel
- Learning how the IO and synchronization between CPU and GPU work.
This will be as learning any new programming skill.Start with understanding parallel computing concepts and how GPUs are structured for it. Optimization is key - learn about memory access patterns, thread management, and how to profile your code to find bottlenecks. There are tons of great resources online, and NVIDIA's own documentation is surprisingly good.
As for the data engineering side, tbh, it's tougher to get into MLE without ML knowledge. However, focusing on the data pipeline, feature engineering, and data quality aspects for ML projects might be
> As for the data engineering side, tbh, it's tougher to get into MLE without ML knowledge. However, focusing on the data pipeline, feature engineering, and data quality aspects for ML projects might be
I have a feeling that companies usually expect MLE to do both ML/AI and Data Engineering, so this might indeed be a dead end. Somehow I'm just not very interested in the MLE part of ML so I'll dormant that thought for the meanwhile.
> Start with understanding parallel computing concepts and how GPUs are structured for it. Optimization is key - learn about memory access patterns, thread management, and how to profile your code to find bottlenecks. There are tons of great resources online, and NVIDIA's own documentation is surprisingly good.
Thanks a lot! I'll take these points in mind when learning. I need to go through more basic CompArch materials first I think. I'm not a good programmer :D
https://github.com/uncomplicate/clojurecuda
There's also tons of free tutorials at https://dragan.rocks And a few books! (not free) at https://aiprobook.com
Everything from scratch, interactive, line-by-line, and each line is executed in the live REPL.
CUDA requires clear understanding of mathematics related to graphics processing and algebra. Using CUDA like you would use traditional CPU would yield abysmal performance.
> MLE or AI Data Engineering without knowing AI/ML
It's impossible to do so, considering that you need to know exactly how the data is used in the models. At the very least you need to understand the basics of the systems that use your data.
Like 90% of the time spent in creating ML based applications is preparing the data to be useful for a particular use case. And if you take Google ML Crash Course, you'll understand why you need to know what and why.
They have excellent resources to get you started with Cuda/Triton on top of torch. It also has a good community around it so you get to listen to some amazing people :)
I have a slightly tangential question: Do you have any insights into what exactly DeepSeek did by bypassing CUDA that made their run more efficient?
I always found it surprising that a core library like Cuda, developed over such a long time, still had room for improvement—especially to the extent that a seemingly new team of developers could bridge the gap on their own.
> with better AI models and tools like Cursor, we will move to a world where you can mold code ever more specific to your use case to make it more performant
what do you think the value of having the right abstraction will be in such a world?
If code is ever updated by an LLM, does it benefit from using abstractions? After all they're really a tool for us lowly sapients to aid in breaking down complex problems. Maybe LLM's will create their own class of abstractions, diverse from our own but useful for their task.
Their paper [1] only mentions using PTX in a few areas to optimize data transfer operations so they don't blow up the L2 cache. This makes intuitive sense to me, since the main limitation of the H800 vs H100 is reduced nvlink bandwidth, which would necessitate doing stuff like this that may not be a common thing for others who have access to H100s.
CUDA is not only C++, as many mistake it for.
Programming Massively Parallel Processors by Wen-mei W. Hwu , David B. Kirk , Izzat El Hajj
seems to be tailor mode for folks transitioning from cpu -> gpu arch.> Combining evolutionary optimization with LLMs is powerful but can also find ways to trick the verification sandbox. We are fortunate to have Twitter user @main_horse help test our CUDA kernels, to identify that The AI CUDA Engineer had found a way to “cheat”. The system had found a memory exploit in the evaluation code which, in a small percentage of cases, allowed it to avoid checking for correctness (...)
The generated implementation doesn’t do a convolution.
The 2nd kernel on the leaderboard also appears to be incorrect, with a bunch of dead code computing a convolution and then not using it and writing tanhf(1.0f) * scaling_factor for every output.
The very idea is also dumb as hell. They could have done CUDA -> HIP/oneAPI/Metal/Vulkan/SYCL/OpenCL. Then they wouldn't need to beat the performance of anything, just the automatic porting would be worth an acquisition by AMD or Intel.
They get caught up in the hype, and focus on the marketing and not the essential engineering.
I'd recommend pyspur if you seek
1) More AI-native features eg. Evals, RAG, or even UI decisions like seeing outputs directly on the canvas when running on the agent 2) Truly open-source Apache license 3) Python-based (in the sense that you can run and extend it via python)
On the other hand, n8n is 1) more mature for traditional workflows 2) offering overall more integrations (probably every single integration you can think of) 3) TypeScript based and runs on Node.js
Yes, we're actively working on this, and we should have some more pages by next week. If you have any questions, you can always shoot us an email: founders@pyspur.dev or join our Discord.
> some links don't work, e.g., Next Steps on this page
This might be confusing, the cards below "After installation, you can:" are not meant to be links. Thanks for making us aware, we will improve the wording.
I've been curious for a few years now to get into MLIR, but I don't know compilers or LLVM, and all the docs I've found seem to assume knowledge of one or the other.
(yes this is a plea for someone to write an 'intro to compilers' using MLIR)
"PyDSL: A MLIR DSL for Python developers"
https://www.youtube.com/watch?v=iYLxgTRe8TU
"PyDSL, a subset of Python for constructing affine & transform dialects"
https://www.youtube.com/watch?v=nmtHeRkl850
And MLIR channel,