I needed this info, thanks for putting it up. Can this really be an issue for every data center?
1,647 karma · joined March 15, 2011
Founder at Flybrix, Dir. DSP at General Radar, Founding Engineer at hCaptcha, Zigfu YC S11
amirhirsch.com
I needed this info, thanks for putting it up. Can this really be an issue for every data center?
Also the machine is well north of 100K when you include the RF ADCs and DACs in there that run a radar.
Worst case, I have multiple.
HuggingFace has incredible reach but poor UX, and PyTorch installs remain fragile. There’s real space here for a platform that makes this all seamless maybe even something that auto-updates a local SSD with fresh models to try every day.
Consider examples using building tools like screwing in a drywall screw, or hammering a nail, using a paint roller, caulking a sink, minor plumbing repair with a torch and solder. These differ enough in terms of forces, state changes, and combined dexterity/acuity (two-handed proprioception) from the windex, sandwich and key examples
Ikea product assembly for gold medal.
The person who showed a speed-up indicated over a week of prior experience with cursor while all others under a week.
The specific time sucks measured in the study are addressable with improved technology like faster LLMs and improved methodology like running parallel agents—the study was done in March running Claude 3.7 and before Claude Code.
We also should value the perception of having worked 20% less even if you actually spent more time. Time flies when you’re having fun!
In San Francisco we just call it “a Waymo”
Xilinx routinely had more I/O (SerDes, 100/200/400G MACs on-die) and at times now more HBM bandwidth than contemporary GPUs. Also deterministic latency and perfectly acceptable DSP primitives.
The gap has always been the software.
Of course NVidia wasn’t such an obvious hit either, the flubbed the tablet market due to yield issues and ultimately it really only went exponential in 2014. I invested heavily in NVidia 2007-2014 because of the CUDA edge they had, but sold my $40K of stock at my cost-basis.
I currently do DSP for radar, and implemented the same system on FPGA and in CUDA 2020-2023. I know as a fact that the FFT performance of an $9000 FPGA was equal to a $16000 A100 that also needed a $10000 computer in 2022 (the types on FPGA were fixed point instead of float so no apples-to-apples but definitely application equivalent)
Even once you get that LED blinking, changing a clock speed for that blinking LED should be near instantaneous but more likely requires a rebuilding the whole project. Fundamentally the vendors don’t view their chips as something designed to run programs, and this legacy hardware design mentality plagues their whole business.
Something important here: Xilinx could and should have been where NVidia is today. They were certainly aware of the competitive accelerated computing market as early as 2005, and fundamentally failed to make a software architecture competitive with CUDA.
Before CUDA even existed I interned at Xilinx working on the beginnings of their HLS C compiler. My (decade older) fraternity brother led the C compiler team at Altera. We almost went into making a spreadsheet compiler for FPGA (my masters thesis) together but 2007 ended up being a terrible year to sell accelerated computing to Wall Street.
And that was before Claude Code.
It hits the request per minute limit instantly and then you wait a minute.
Here's a fun experiment for someone: 1) Give N people K fake credit cards to enter into a form, and have them solve a captcha 2) Take recorded keyboard and mouse data similar to the captcha 3) Train a neural network model to identify
I've been out of this for 6 years but I bet transformers rock this problem now.