HNHacker News
TopNewBestAskShowJobs

neilmovva

312 karma · joined August 7, 2014

Co-founder of Blyss.dev : end-to-end encrypted AI. YC W23. neil@blyss.dev
submissionscomments
neilmovva··on So you want to use OpenRouter?
hi, I'm one of the founders of Sail. I'm very sorry that you had a bad experience with us! We are serious about serving models correctly, and always publish a link to the exact HF checkpoint we're using for each model in our docs. If you ever have an issue like this again, please send a note to support@sailresearch.com and we'll make it right with a detailed postmortem.
neilmovva··on TPUs vs. GPUs and why Google is positioned to win AI race in the long term
I think Hopper's native matmul tile is 64x64, and Blackwell is 128x128.

see this blog for a reference on Blackwell:

https://hazyresearch.stanford.edu/blog/2025-03-15-tk-blackwe...

neilmovva··on Qwen3-Omni: Native Omni AI model for text, image and video
The multilingual example in the launch graphic has Qwen3 producing the text:

> "Bonjour, pourriez-vous me dire comment se rendreà la place Tian'anmen?"

translation: "Hello, could you tell me how to get to Tiananmen Square?"

a bold choice!

neilmovva··on Writing Speed-of-Light Flash Attention for 5090 in CUDA C++
Today, training in "low precision" probably means computing FP8 x FP8 -> FP32. The FP32 accumulation is still important, but otherwise yes this works, especially if we're talking about MXFP8 as supported on Blackwell [0].

What's less proven is a recipe using MXFP4 x MXFP4 -> FP32 compute, e.g. [1], which needs more involved techniques to work. But if you get it to work stably, that pathway is running at full throughput on 5090.

[0]: https://arxiv.org/abs/2506.08027 [1]: https://arxiv.org/abs/2502.20586

neilmovva··on Writing Speed-of-Light Flash Attention for 5090 in CUDA C++
Not really:

5090: 210 TF / $2k == 105 TF/$k

B200: 2250 TF / $40k == 56 TF/$k

Getting only 2x the FLOPs per dollar probably isn't worth the hassle of having to rack 10x as many GPUs, while having no NVLink.

neilmovva··on Writing Speed-of-Light Flash Attention for 5090 in CUDA C++
I was surprised to see 5090's theoretical BF16 TFLOPs at just 209.5. That's not even 10% of the server Blackwell (B200 is 2250, and GB200 is 2500). B200 costs around $30-40k per GPU, so they are pretty close in performance per dollar.

Starting with 4090, NVIDIA limits the performance of tensor cores on gaming cards, specifically for ops that might be used in ML training. FP8 and FP16 matmuls run at full speed if accumulating in FP16 (I've never seen anyone use this), but only half speed when accumulating in FP32. This restriction is not present for lower precision matmuls like FP4, and is removed entirely on the workstation-class cards like RTX Pro 6000.

It doesn't seem worth it to use NVIDIA gaming cards as a "cheaper FLOPs" alternative anymore (e.g. diffusion models could have been cheaper to run on 3090 than A100). They are generous with memory bandwidth though, nearly 2TB/s on 5090 is amazing!

neilmovva··on Intel Gaudi 3 AI Accelerator
A bit surprised that they're using HBM2e, which is what Nvidia A100 (80GB) used back in 2020. But Intel is using 8 stacks here, so Gaudi 3 achieves comparable total bandwidth (3.7TB/s) to H100 (3.4TB/s) which uses 5 stacks of HBM3. Hopefully the older HBM has better supply - HBM3 is hard to get right now!

The Gaudi 3 multi-chip package also looks interesting. I see 2 central compute dies, 8 HBM die stacks, and then 6 small dies interleaved between the HBM stacks - curious to know whether those are also functional, or just structural elements for mechanical support.

neilmovva··on Show HN: 80% faster, 50% less memory, 0% loss of accuracy Llama finetuning
I agree that synchronization causes overhead, so 2x GPUs won't achieve the ideal 0.5x total runtime. But here, taking your Alpaca benchmark as an example, we are seeing 2x GPUs get 3.6x runtime with Huggingface, or 1.15x with Unsloth Max.

In other words, every benchmark, in either HF or Unsloth, is slower in absolute terms when going from 1 to 2 GPUs. That makes me think something is wrong with the test.

Could you share your benchmark code?

neilmovva··on Show HN: 80% faster, 50% less memory, 0% loss of accuracy Llama finetuning
promising results, excited to try it out!

question on the perf benchmarks: why do all the results with 2 GPUs & DDP take longer than the single GPU case? Both benchmarks do the same amount of work, one training epoch, so this negative scaling is surprising.

neilmovva··on SRAM in AI: The Future of Memory
I can't universally agree with the headline statement. The article focuses on the pros of SRAM, which are real -- peak bandwidth (e.g. 5 TB/s out of the H100’s L2) and lower energy per bit transferred (the rule of thumb I remember is ~10x lower than HBM).

But the companies who already bet big on SRAM in AI, Cerebras in particular and Graphcore to lesser extent, aren’t obviously running away with the AI performance crown. Seems like LLMs need more memory capacity than anyone expected, to the point where even HBM stacks somewhat limit the scale of current models. Maybe the next version of Cerebras WSE can get closer to 100GB of on-chip memory, and serve some useful LLMs very efficiently - excited to see what they can do with more modern processes!

I think innovation in SRAM packaging, like AMD’s stacking in “3D V-Cache”, is also promising and might play a larger role in AI accelerators going forward. But it’s important to note that while both are SRAM, performance of AMD’s stacked L3 is not yet comparable to a GPU’s centralized L2 - it’s more like HBM today in latency and bandwidth.

neilmovva··on Sandy Bridge: Setting Intel’s modern foundation
I like this review:

https://www.lighterra.com/papers/modernmicroprocessors/

A bit dated, but the major ideas used in current CPUs are all covered!

neilmovva··on The New Super Commute
> moves to Austin because it is less “vulnerable to climate change”

> commutes by plane

hmmm

neilmovva··on Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any use of H100.

The new "NVL" variant adds ~20% more memory per GPU by enabling the sixth HBM stack (previously only five out of six were used). Additionally, GPUs now come in pairs with 600GB/s bandwidth between the paired devices. However, the pair then uses PCIe as the sole interface to the rest of the system. This topology is an interesting hybrid of the previous DGX (put all GPUs onto a unified NVLink graph), and the more traditional PCIe accelerator cards (star topology of PCIe links, host CPU is the root node). Probably not an issue, I think PCIe 5.0 x16 is already fast enough to not bottleneck multi-GPU training too much.

neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Yes, we all place a lot of trust in cloud vendors today. FHE is a way to move the trust boundary back to the client - let the server be as malicious or insecure as it wants. Raw compute could even become much cheaper, since any machine anywhere can be a supplier in the market for untrusted CPU time.
neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Thanks for the feedback, I understand your hesitation. We don't just want to advertise guarantees - we want you to never trust third-party servers again. Fully homomorphic encryption makes this possible by never letting sensitive data even leave your device. Our job is to make this new cryptography a web standard as ubiquitous as TLS.
neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Thanks for checking it out! Responses inline:

> That sounds like loading the entire database every time

Yup, we do perform computation over the entire database for every read - there is zero correlation between the server's work and the client's query. We currently serve queries to a 1 GB database in under 1 second. For much larger databases (100+ GB), this becomes more a question of cost: we can stay fast (1 sec) with more expense, or go slower (e.g. 5 sec) and stay cheap.

> Cryptographic proof that the server is not spying

If you trust your client software [0], then you can be sure that your request isn't decrypted anywhere outside your device. Even a malicious Blyss server cannot determine your query, because it never got a chance to see it.

[0] This level of security depends entirely on having a trusted client. Our client software is open source, and we plan to have it formally audited. We'll also publish signed desktop apps so you can be sure that you're running the same client every time.

neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Thanks! Yup, it's not always practical to make a huge number of queries when you expect many of them to come back empty. Instead, we first perform private lookups against a Bloom filter, to find out which keys actually hold data (e.g. messages). Then, we privately retrieve only the useful keys.

The Bloom filter is also served over Blyss, so the server still learns nothing about which keys you're interested in. We implemented this system for our private password checker, which tests passwords against almost a billion breached credentials: https://playground.blyss.dev/passwords

neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Thanks! Yup, private retrieval is interesting as a product because it's a fundamentally new capability; there aren't really competitors we can show incremental improvements against. If you're still interested in the space, we'd be happy to compare notes! Feel free to email us: founders AT blyss.dev
neilmovva··on Launch HN: Blyss (YC W23) – Homomorphic encryption as a service
Our FHE scheme uses lots of Number Theoretic Transforms (NTTs), which are pretty computationally expensive. NTT is a good candidate for acceleration, and there is quite a bit of interest from the zk community in doing so (https://www.zprize.io/prizes/accelerating-ntt-operations-on-...).

From a hardware perspective, NTT can be done in parallel, but has a fairly large working set of data (~512 MB) with lots of unstructured accesses. This is too big to fit in even the largest CPU L3 caches, so DRAM bandwidth is still relevant. It may be eventually be feasible to build an ASIC with this much on-chip memory, but in the meantime, GPUs do a pretty decent job with their massive HBM bandwidth.

neilmovva··on Running large language models like ChatGPT on a single GPU
While OPT-175B is great to have publicly available, it needs a lot more training to achieve good results. Meta trained OPT on 180B tokens, compared to 300B that GPT-3 saw. And the Chinchilla scaling laws suggest that almost 4T tokens would be required to get the most bang for compute buck.

And on top of that, there are some questions on the quality of open source data (The Pile) vs OpenAI’s proprietary dataset, which they seem to have spent a lot of effort cleaning. So: open source models are probably data-constrained, in both quantity and quality.

neilmovva··on Intel Core i9-13900T CPU benchmarks show faster than 12900K 125W performance
I was curious about exactly how much burst heat could be absorbed, so I asked WolframAlpha [0]. In a 15" workstation laptop, I think the CPU could quite reasonably pull +100 watts over steady-state TDP for 30 seconds. (100-gram aluminum heatsink absorbing 100W * 30sec -> ∆T = +33ºC)

[0] https://www.wolframalpha.com/input?i=%28100w+*+30+sec%29+%2F...

neilmovva··on Chipmaker Analog Devices to acquire Maxim Integrated for $21B
Latest in a trend of silicon industry consolidation. A few other major moves in the embedded market over the last five years:

NXP + Freescale in 2015

Microchip + Atmel in 2016

ON Semi + Fairchild in 2016

Infineon + Cypress in 2020

neilmovva··on Zen 2 Missives – AMD now delivering efficiencies that are double that of Intel
The author's comments on cache sizes are a bit reductive. Not all "L3" is created equal, and designers always make tradeoffs between capacity and latency.

In particular, the EPYC processors achieve such high cache capacities by splitting L3 into slices across multiple silicon dies, and accessing non-local L3 incurs huge interconnect latency - 132ns on latest EPYC vs 37ns on current Xeon [1]. Even DDR4 on Intel (90ns) is faster than much of an EPYC chip's L3 cache.

Intel's monolithic die strategy keeps worst case latency low, but increases costs significantly and totally precludes caches in the hundreds of MB. Depending on workload, that may or may not be the right choice.

[1] https://www.anandtech.com/show/14694/amd-rome-epyc-2nd-gen/7

neilmovva··on A Look at the AMD Zen 2 Core
Actually, the L3 cache is also sharded across chiplets, so there's a small (~8MB) local portion of L3 that is fast, while remote slices will have to go over AMD's interdie connection fabric and incur a serious latency penalty. On first gen Epyc/Threadripper, nonlocal L3 hits were almost as slow as DRAM at ~100ns (!).
neilmovva··on The Dark Silicon Problem and What It Means for CPU Designers (2013)
That still means multiple silicon dies, which we have known how to do for a while (see: Intel Core 2 Quad from 2006, and more recently AMD Epyc).

Having more dies lets you dissipate more heat, but then it's kinda hard to build low-latency / high-bandwidth interconnects between the dies. Inter-die buses go over a PCB or interposer, which impose higher parasitic capacitance and make it difficult/expensive to run wide interfaces. That's why techniques like "dark silicon" allocation are important - it allows us to get more perf in a single die.

neilmovva··on Cloud TPUs in Beta
It matters when defining parallel work distribution. Unless memory bandwidth is homogeneous across the whole board (i.e. each TPU on a board gets 600 GB/s to its peers), we can't do model parallelism across ASICs efficiently, and must fall back to data parallelism. Which is fine, until you run into limits on maximum batchsize (e.g. up to 8192, as FAIR was able to manage [1] with some tweaks to SGD).

[1] https://arxiv.org/abs/1706.02677

neilmovva··on Intel delays Cannonlake 10nm processor for the third time
Intel's "10nm" has roughly twice the transistor density vs. Samsung/TSMC "10nm", so I wouldn't compare based on the advertised process names.
neilmovva··on HP Updates Z8 Workstations: Up to 56 Cores, 3 TB RAM, 9 PCIe Slots, 1700W
The fact that "2TB is addressable" is irrelevant. Putting NAND on the board doesn't improve latency/bandwidth nearly enough to function like vram. Nvidia has also supported unified virtual memory since Pascal, meaning you can address your cpu's ram in GPU code. The "SSG" card still has the 2tb of flash on a PCIe interface, so not much difference from existing systems beyond marketing. I'd expect very few real world perf wins.
neilmovva··on AMD unveils its Vega 10 GPU architecture
GPUs are generally not latency optimized anyway - if an app is sensitive to tens of microseconds then it belongs on the CPU, broadly speaking.

AMD's advantage here is getting drop-in throughput, by switching the SSD directly to the GPU. Thus, the onboard SSD doesn't waste host PCIe lanes (4 lanes, typically) that are getting hard to come by in regular desktop computing systems.

neilmovva··on AMD unveils its Vega 10 GPU architecture
Vega is already drawing near 300W, and is so high up on the voltage/frequency curve that even a measly 5-10% gain in core clock can easily cost over 100W more.

Here, AMD is once again (see last year's Polaris) a victim of the inferior GlobalFoundries 14nm LPP process. TSMC 16nm would have been much better in perf/W, but sadly a very restrictive wafer supply agreement locks AMD to GloFo for the time being.

Page 1 of 2Next →