8 karma · joined March 6, 2025
I was wrong!!
We're now seeing multi-agent systems that take your PyTorch code and spit out CUDA or Triton kernels with 2x to 14x speedups over torch.compile(mode='max-autotune-no-cudagraphs'). Not on toy benchmarks. On real models like Llama-3.1-8B, Whisper, and Stable Diffusion
So I built code editor for CUDA that does it all:
Profile and benchmark your kernels in real-time while you code
Emulate multi-GPU without the hardware
Get AI optimization suggestions that actually understand your GPU "you can use local llm to cost you 0$"
It's free to use if you use your local LLM :D Still needs a lot of refinement, so feel free to share anything you'd like to see in it
so built an emulator instead
you give it a kernel, it predicts execution time on any gpu without running it. h100, a100, v100, whatever.
how: scraped specs for 50+ nvidia gpus, built tile-based simulator that models memory bandwidth, occupancy, and sm scheduling. validated against 12 real gpus and the mean error 1.2%
doesn't work for: dynamic parallelism, multi-gpu, tiny kernels under 1us but I will figure it out soon
if anyone's solved this differently?
we couldn’t afford that, so we built an emulator that predicts how your kernel runs on any GPU like H100, A100, 4090, or V100 without running a single line
it’s not a guess, it gives real numbers
2.4ms on RTX 4090, 5.1ms on V100 within 1% of hardware
how it works
- NeuSight (99%) splits the kernel into tiles, simulates each one using real GPU specs like 132 SMs on H100 or 10 on 1060, checks occupancy, bandwidth, wave scheduling
- NCU Baseline (95–98%) if you profiled once, we scale it across GPUs, Hopper is 1.05x Ada, Ampere 0.92x, all measured manually
- Analytical (85–92%) roofline model fallback, works even without source code
we validated on 47 kernels across 12 GPUs
accuracy stayed above 98%, occupancy predictions were almost perfect
one team saved $18k in GPU cloud time
another found bugs on an A100 they didn’t own
still missing dynamic parallelism, multi-GPU, and tensor core perfection but we’re getting there
happy to go into the math or architecture details if anyone’s curious
So we built an emulator. give it your kernel code and it tells you exactly how it runs on any GPU. H100, A100, RTX 4090, V100, whatever "without running a single line"
Not a rough estimate. Real numbers. 2.4ms on RTX 4090, 5.1ms on V100. Within 1% of actual hardware.
How it works:
We have three emulators. Each trades accuracy for speed:
1. NeuSight (99% accurate): Breaks your kernel into tiles, simulates each one. Uses real GPU specs from our database (132 SMs for H100, 10 SMs for GTX 1060, etc). Checks occupancy, memory bandwidth, wave scheduling
2. NCU Baseline (95-98% accurate): If you already profiled on one GPU, we scale to others. Hopper is 1.05x faster than Ada at compute. Ampere is 0.92x. We measured it all
3. Analytical (85-92% accurate): Fast backup using roofline model. Works even without source code
We validated with 47 test kernels on 12 real GPUs. Results: 98-99% accuracy on execution time. Occupancy prediction is basically perfect
One team saved $18,000 in cloud costs. Another caught bugs on an A100 they don't even own
What doesn't work yet: Dynamic parallelism, multi-GPU, perfect tensor core modeling but nothing is impossible, we will figure it out soon
Built into RightNow AI (our CUDA editor). Free to try
https://www.rightnowai.co/blog/building-99-accurate-gpu-emul...
We're a small team and we spent 3 months on this
Happy to answer questions!!
Everything is either tiny or garbage quality.
So I built my own. 744K articles, 244 million words. Spent months cleaning it properly.
It's 8.7GB of good Arabic text covering everything. Made it completely free.
GitHub: https://github.com/RightNow-AI/rightnow-arabic-llm-corpus
Anyone building Arabic AI models?
- 743K articles
- 244M words
- 1.5M unique words
- Cleaned, deduplicated, JSONL format, LLM-ready
It’s ready for GPT, BERT, LLaMA, or any Arabic NLP research.
For too long, Arabic AI lagged behind. Most LLMs ignored our language, and high-quality datasets didn’t exist. That changes today.
More GPUs → smaller gains
That’s the law.
The only move left: shift the intercept.
Make the same FLOPs buy lower loss.
I broke it down:
- Why scaling stalls
- What actually shifts the curve
- A one-week playbook to run now
This got me thinking: we spend so much time iterating manually when profiling could guide optimizations automatically. Curious if others have tried this approach, or if there are better ways to tackle kernel tuning at scale.
It automatically optimizes CUDA kernels. Depending on the kernel, you can get 10–50x speedups :D
How it works:
- Finds bottlenecks in your kernel
- Generates a few optimized variants
- Benchmarks them
- Keeps the fastest one
Everything is open, so you can look under the hood and tweak anything.
How it works: - Analyze kernel patterns - Generate AI-optimized variants - Benchmark automatically - Ship the fastest version
You can see and modify everything. Full control and transparency.
Open source here: https://github.com/RightNow-AI/rightnow-cli
For full AI-assisted optimization across your entire codebase, check RightNow AI Code Editor: https://rightnowai.co
It’s a GPU-first editor that:
Profiles your kernels live without leaving the editor
Lets AI rewrite and tune them for your exact GPU
Autocompletes based on your hardware, not generic code
Supports multi-GPU setups out of the box
Works with local or cloud LLMs
In tests, we’ve seen up to 180× speedups on the same machine with no manual tuning.
Beta is open now, and we’ll open source it soon once it’s stable.
So, I built RightNow AI solo while still at university. It’s a CUDA code editor that profiles kernels live inside the editor, writes kernels optimized for your exact GPU, autocompletes code based on your hardware, shows inline profiling for debugging and optimization, and lets you choose between local AI or cloud models for privacy.
It’s free during the preview period. You can use your own API keys or run local LLMs with Ollama, etc.
In testing, kernels generated in RightNow AI have run up to 179x faster on the same hardware, with no manual tuning!!
Join the waitlist and get early access before it fills up
I just put the product out there for anyone who wants to try it for FREE and share feedback. I already have a good number of real users, and they’re happy with it