HNHacker News
TopNewBestAskShowJobs

jaberjaber23

8 karma · joined March 6, 2025

i build things
submissionscomments
jaberjaber23··on The Concatative Language XY
interesting mix of k and joy. the queue manipulation primitives like -> and => have no equivalent in joy, lets you do things like call/cc in a few lines
jaberjaber23··on Maze Algorithms (2017)
seconding the jamis buck book, its one of the few programming books i actually finished. the way he explains each algorithm with visualizations makes it stick
jaberjaber23··on Doing gigabit Ethernet over my British phone wires
had the same thing, electrician just used whatever was on the truck. swapped the faceplates and got instant gigabit
jaberjaber23··on [dead]
Six months ago, if you asked me whether an LLM could write a CUDA kernel that actually beats PyTorch's compiler, I would have said no. The optimization space is too complex. Too many hardware details. Too easy to write something that compiles but runs slower than the baseline

I was wrong!!

We're now seeing multi-agent systems that take your PyTorch code and spit out CUDA or Triton kernels with 2x to 14x speedups over torch.compile(mode='max-autotune-no-cudagraphs'). Not on toy benchmarks. On real models like Llama-3.1-8B, Whisper, and Stable Diffusion

jaberjaber23··on [dead]
amazing paper
jaberjaber23··on CUDA Programming Is Cooked
Wow man, will give it a try
jaberjaber23··on New Era for CUDA Development
Impressive work, will give it a try!!
jaberjaber23··on [dead]
I was using 4 sometimes 6 different tools just to write CUDA. vs code for coding, nsight for profiling, many custom tools for benchmarking and debugging, plus pen to calc the performance "I was cooked"

So I built code editor for CUDA that does it all:

Profile and benchmark your kernels in real-time while you code

Emulate multi-GPU without the hardware

Get AI optimization suggestions that actually understand your GPU "you can use local llm to cost you 0$"

It's free to use if you use your local LLM :D Still needs a lot of refinement, so feel free to share anything you'd like to see in it

jaberjaber23··on Who invented deep residual learning?
science repeats itself
jaberjaber23··on How to sequence your DNA for <$2k
Nanopore’s getting closer
jaberjaber23··on I Built Claude Code for CUDA Using Python
amazing!!
jaberjaber23··on Structured Procrastination (1995)
most people don’t procrastinate because they’re lazy, they procrastinate because their brain rejects meaningless work
jaberjaber23··on Why do LLMs freak out over the seahorse emoji?
llms don’t actually freak out over seahorses. it’s just a mismatch. the model thinks “seahorse emoji” is real, but the output system doesn’t have a token for it. it tries to show what it means, realizes it can’t, and spirals trying to fix itself
jaberjaber23··on [dead]
testing cuda kernels on different gpus costs $7k/month in cloud rentals

so built an emulator instead

you give it a kernel, it predicts execution time on any gpu without running it. h100, a100, v100, whatever.

how: scraped specs for 50+ nvidia gpus, built tile-based simulator that models memory bandwidth, occupancy, and sm scheduling. validated against 12 real gpus and the mean error 1.2%

doesn't work for: dynamic parallelism, multi-gpu, tiny kernels under 1us but I will figure it out soon

if anyone's solved this differently?

jaberjaber23··on [dead]
Testing CUDA kernels on 15 GPUs costs thousands every month

we couldn’t afford that, so we built an emulator that predicts how your kernel runs on any GPU like H100, A100, 4090, or V100 without running a single line

it’s not a guess, it gives real numbers

2.4ms on RTX 4090, 5.1ms on V100 within 1% of hardware

how it works

- NeuSight (99%) splits the kernel into tiles, simulates each one using real GPU specs like 132 SMs on H100 or 10 on 1060, checks occupancy, bandwidth, wave scheduling

- NCU Baseline (95–98%) if you profiled once, we scale it across GPUs, Hopper is 1.05x Ada, Ampere 0.92x, all measured manually

- Analytical (85–92%) roofline model fallback, works even without source code

we validated on 47 kernels across 12 GPUs

accuracy stayed above 98%, occupancy predictions were almost perfect

one team saved $18k in GPU cloud time

another found bugs on an A100 they didn’t own

still missing dynamic parallelism, multi-GPU, and tensor core perfection but we’re getting there

happy to go into the math or architecture details if anyone’s curious

jaberjaber23··on [dead]
Testing CUDA kernels on 15 different GPUs costs $3,000/month. We couldn't afford that!!

So we built an emulator. give it your kernel code and it tells you exactly how it runs on any GPU. H100, A100, RTX 4090, V100, whatever "without running a single line"

Not a rough estimate. Real numbers. 2.4ms on RTX 4090, 5.1ms on V100. Within 1% of actual hardware.

How it works:

We have three emulators. Each trades accuracy for speed:

1. NeuSight (99% accurate): Breaks your kernel into tiles, simulates each one. Uses real GPU specs from our database (132 SMs for H100, 10 SMs for GTX 1060, etc). Checks occupancy, memory bandwidth, wave scheduling

2. NCU Baseline (95-98% accurate): If you already profiled on one GPU, we scale to others. Hopper is 1.05x faster than Ada at compute. Ampere is 0.92x. We measured it all

3. Analytical (85-92% accurate): Fast backup using roofline model. Works even without source code

We validated with 47 test kernels on 12 real GPUs. Results: 98-99% accuracy on execution time. Occupancy prediction is basically perfect

One team saved $18,000 in cloud costs. Another caught bugs on an A100 they don't even own

What doesn't work yet: Dynamic parallelism, multi-GPU, perfect tensor core modeling but nothing is impossible, we will figure it out soon

Built into RightNow AI (our CUDA editor). Free to try

https://www.rightnowai.co/blog/building-99-accurate-gpu-emul...

We're a small team and we spent 3 months on this

Happy to answer questions!!

jaberjaber23··on GPUs need their own editor
exactly
jaberjaber23··on [dead]
Arabic AI training data is terrible.

Everything is either tiny or garbage quality.

So I built my own. 744K articles, 244 million words. Spent months cleaning it properly.

It's 8.7GB of good Arabic text covering everything. Made it completely free.

GitHub: https://github.com/RightNow-AI/rightnow-arabic-llm-corpus

Anyone building Arabic AI models?

jaberjaber23··on [dead]
I built one of the largest Arabic LLM datasets:

- 743K articles

- 244M words

- 1.5M unique words

- Cleaned, deduplicated, JSONL format, LLM-ready

It’s ready for GPT, BERT, LLaMA, or any Arabic NLP research.

For too long, Arabic AI lagged behind. Most LLMs ignored our language, and high-quality datasets didn’t exist. That changes today.

jaberjaber23··on [dead]
AI hits a scaling wall

More GPUs → smaller gains

That’s the law.

The only move left: shift the intercept.

Make the same FLOPs buy lower loss.

I broke it down:

- Why scaling stalls

- What actually shifts the curve

- A one-week playbook to run now

jaberjaber23··on [dead]
I’ve spent years manually tuning CUDA kernels, and it’s exhausting. Recently, I experimented with real-time profiling to see bottlenecks instantly. The results were surprising—small tweaks yielded huge performance gains.

This got me thinking: we spend so much time iterating manually when profiling could guide optimizations automatically. Curious if others have tried this approach, or if there are better ways to tackle kernel tuning at scale.

jaberjaber23··on [dead]
Just open-sourced CUDA CLI :)

It automatically optimizes CUDA kernels. Depending on the kernel, you can get 10–50x speedups :D

How it works:

- Finds bottlenecks in your kernel

- Generates a few optimized variants

- Benchmarks them

- Keeps the fastest one

Everything is open, so you can look under the hood and tweak anything.

jaberjaber23··on Open Source AI CUDA Kernel Optimizer
RightNow CLI is now open source. It automatically optimizes CUDA kernels using AI. Depending on the kernel, it can achieve 10 to 50x speedups by removing bottlenecks and tuning for your GPU.

How it works: - Analyze kernel patterns - Generate AI-optimized variants - Benchmark automatically - Ship the fastest version

You can see and modify everything. Full control and transparency.

Open source here: https://github.com/RightNow-AI/rightnow-cli

For full AI-assisted optimization across your entire codebase, check RightNow AI Code Editor: https://rightnowai.co

jaberjaber23··on [dead]
Most developers guess why their CUDA code is slow. I wanted answers, so I built RightNow AI.

It’s a GPU-first editor that:

Profiles your kernels live without leaving the editor

Lets AI rewrite and tune them for your exact GPU

Autocompletes based on your hardware, not generic code

Supports multi-GPU setups out of the box

Works with local or cloud LLMs

In tests, we’ve seen up to 180× speedups on the same machine with no manual tuning.

Beta is open now, and we’ll open source it soon once it’s stable.

jaberjaber23··on [dead]
Most CUDA developers waste hours trying to speed up kernels because they can’t see the full architecture and behavior of their GPU.

So, I built RightNow AI solo while still at university. It’s a CUDA code editor that profiles kernels live inside the editor, writes kernels optimized for your exact GPU, autocompletes code based on your hardware, shows inline profiling for debugging and optimization, and lets you choose between local AI or cloud models for privacy.

It’s free during the preview period. You can use your own API keys or run local LLMs with Ollama, etc.

In testing, kernels generated in RightNow AI have run up to 179x faster on the same hardware, with no manual tuning!!

Join the waitlist and get early access before it fills up

jaberjaber23··on We Made CUDA Optimization Suck Less
what's up guys, take it easy. Just to clarify: I didn’t add any reviews myself. I’m building this SOLO and barely have time to finish the product, let alone fake comments. I wasn’t even aware of the Product Hunt stuff until people here pointed it out!!!!

I just put the product out there for anyone who wants to try it for FREE and share feedback. I already have a good number of real users, and they’re happy with it

jaberjaber23··on We Made CUDA Optimization Suck Less
I really appreciate that!! thanks:D
jaberjaber23··on We Made CUDA Optimization Suck Less
Do you have a place where we can chat? Linkedin,....
jaberjaber23··on We Made CUDA Optimization Suck Less
absolutely. it really depends on the kernel type, target architecture, and what you're optimizing for. the 2x-4x isn’t the limit, it's just what users often see out of the box. we do real-time profiling on actual GPUs, so you get results based on real performance on a specific arch, not guesses. when the baseline is rough, we’ve seen well over 10x
jaberjaber23··on We Made CUDA Optimization Suck Less
We profile and optimize kernels live on real GPUs!! So we’re different than unsloth
Page 1 of 2Next →