HNHacker News
TopNewBestAskShowJobs

amirhirsch

1,647 karma · joined March 15, 2011

Digital Logician, Laser Whisperer

Founder at Flybrix, Dir. DSP at General Radar, Founding Engineer at hCaptcha, Zigfu YC S11

amirhirsch.com

submissionscomments
amirhirsch··on I got an Nvidia GH200 server for €7.5k on Reddit and converted it to a desktop
# Tell the driver to completely ignore the NVLINK and it should allow the GPUs to initialise independently over PCIe !!!! This took a week of work to find, thanks Reddit!

I needed this info, thanks for putting it up. Can this really be an issue for every data center?

amirhirsch··on Why xor eax, eax?
I implemented a PDP-11 in 2007-10 and I can still read PDP-11 Octal
amirhirsch··on Every mathematician has only a few tricks (2020)
I am so into optimizing fast polynomial multiplication I assure you there is nothing that will demoralize me from creating a slightly more optimized version.
amirhirsch··on Are you stuck in movie logic?
The phenomenon described in movies has a name called “Idiot Plot” (https://en.wikipedia.org/wiki/Idiot_plot) an older term which Roger Ebert popularized. Feels missing from blogpost.
amirhirsch··on AI adoption in US adds ~900k tons of CO₂ annually, study finds
What if the goal of writing about how “AI is bad for the environment” (because of the energy and water it uses) is to identify gullible people and on-ramp them into a lifetime of media manipulation?
amirhirsch··on Benchmarking leading AI agents against Google reCAPTCHA v2
This was done on the re-captcha demo page no invisible fingerprinting, behavioral test, or user classification.
amirhirsch··on The MP3.com Rescue Barge Barge
i found some of my own music here. not the mp3s i lost forever though.
amirhirsch··on Getting DeepSeek-OCR working on an Nvidia Spark via brute force with Claude Code
it’s all running inside a Proxmox VM with IOMMU and GPU passthrough. It’s as safe as doing the same on any cloud system.

Also the machine is well north of 100K when you include the RF ADCs and DACs in there that run a radar.

Worst case, I have multiple.

amirhirsch··on Getting DeepSeek-OCR working on an Nvidia Spark via brute force with Claude Code
I also use Claude Code to install CUDA and PyTorch and HuggingFace models on my quad A100 machine. Shouldn't feel like debugging a 2000s Linux driver.

HuggingFace has incredible reach but poor UX, and PyTorch installs remain fragile. There’s real space here for a platform that makes this all seamless maybe even something that auto-updates a local SSD with fresh models to try every day.

amirhirsch··on OpenAI researcher announced GPT-5 math breakthrough that never happened
It is not about whether some team somewhere is dabbling in math, it’s about institutional emphasis. When serious math work exists, it tends to be checked before being broadcast. The fact that this didn’t happen suggests that genuine mathematical rigor isn’t a central focus.
amirhirsch··on The Trinary Dream Endures
Not quite. Leakage current in CMOS circuits became the dominant source of power consumption around the 90 nm and 65 nm nodes, long before quantum tunneling was a major factor, and often exceeded dynamic switching power. This led to the introduction of multiple threshold-voltage devices and body-biasing techniques to dynamically adjust Vt and curb static leakage.
amirhirsch··on OpenAI researcher announced GPT-5 math breakthrough that never happened
The people involved are very smart and must know that AI doing novel math is a canary for AGI. A serious effort around solving open problems would not fuck up this kind of announcement.
amirhirsch··on OpenAI researcher announced GPT-5 math breakthrough that never happened
The sad truth about this incident is that it reveals that OpenAI does not have a serious effort to actually work on unsolved math problems.
amirhirsch··on Benjie's Humanoid Olympic Games
I like these benchmarks and the videos are funny!

Consider examples using building tools like screwing in a drywall screw, or hammering a nail, using a paint roller, caulking a sink, minor plumbing repair with a torch and solder. These differ enough in terms of forces, state changes, and combined dexterity/acuity (two-handed proprioception) from the windex, sandwich and key examples

Ikea product assembly for gold medal.

amirhirsch··on Discrete Fourier Transform
If you are doing polynomial multiplication and want exact integer convolution then you can’t leave the precision of your FFT up to chance so the alternative to an FFT over complex roots of unity is to use something called a Number Theoretic Transform (NTT) that relies on nth roots of unity in finite integer rings
amirhirsch··on Zero ASIC releases Wildebeest, the highest performance FPGA synthesis tool
For the curious: the process of discovering the logic and route timings for an FPGA device is to use ring oscillators (three series inverters) and compare counters against a known clock. Place the inverters all over to test every LUT, and use every routing path to test each path's timing.
amirhirsch··on Programming language inventor or serial killer? (2003)
Which are you?
amirhirsch··on AI coding
Just putting this up as a reference the next time this comes up on HN. The study data shows that the median task is 1.5hrs and the 15 minutes that developers think they saved was actually more than that many minutes less researching, testing, and writing code and the 15 minutes longer they actually spent was more idle (5mins) and waiting for AI (5 mins) and reviewing and prompting dominating their work over actually writing code.

The person who showed a speed-up indicated over a week of prior experience with cursor while all others under a week.

amirhirsch··on AI coding
That METR study gets a lot of traction for its headline; and I doubt many people read the whole thing—it was long—but the data showed a 50% speed up for the one dev with the most experience with Cursor/AI, suggesting a learning curve and also wild statistical variation on a small sample set. An errata later suggested another dev who did not have a speedup had not represented their experience correctly, but still strongly draws into question the significance of the findings.

The specific time sucks measured in the study are addressable with improved technology like faster LLMs and improved methodology like running parallel agents—the study was done in March running Claude 3.7 and before Claude Code.

We also should value the perception of having worked 20% less even if you actually spent more time. Time flies when you’re having fun!

amirhirsch··on GigaByte CXL memory expansion card with up to 512GB DRAM
The i in that logo seems like it’s hurting the A
amirhirsch··on Self-driving cars begin testing on NYC streets
I prefer the subway over the street traffic in NYC but often have a problem finding a bathroom. So startup idea: a Waymo but you poop in it.

In San Francisco we just call it “a Waymo”

amirhirsch··on Pebble Time 2 Design Reveal [video]
This is awesome Eric! I'd want to give something like this to my kids, any way to add a tracker?
amirhirsch··on Let's get real about the one-person billion dollar company
If you aren't in a private group chats with folks talking about AI tools and one-person-unicorn ideas are you even a real founder?
amirhirsch··on The Amaranth hardware description language
This isn’t the full story though, like I (professionally, as a consultant) analyzed GOPs/$ and /Watt for big multi chip GPU or FPGA systems from 2006-2011.

Xilinx routinely had more I/O (SerDes, 100/200/400G MACs on-die) and at times now more HBM bandwidth than contemporary GPUs. Also deterministic latency and perfectly acceptable DSP primitives.

The gap has always been the software.

Of course NVidia wasn’t such an obvious hit either, the flubbed the tablet market due to yield issues and ultimately it really only went exponential in 2014. I invested heavily in NVidia 2007-2014 because of the CUDA edge they had, but sold my $40K of stock at my cost-basis.

I currently do DSP for radar, and implemented the same system on FPGA and in CUDA 2020-2023. I know as a fact that the FFT performance of an $9000 FPGA was equal to a $16000 A100 that also needed a $10000 computer in 2022 (the types on FPGA were fixed point instead of float so no apples-to-apples but definitely application equivalent)

amirhirsch··on The Amaranth hardware description language
You can tell the veteran status of FPGA devs by the quality of their rants about the tools. The big FPGA companies have no quality metrics for developer experience. You should be able to make an LED blink within a minute of powering up a board and not after a day of downloading and installing stuff. It used to be possible to quickly start with Vivado on AWS cloud, and I was using that workflow for years, although recent licensing changes presented a speed-bump there, and I ended up going with a local install for my recent project.

Even once you get that LED blinking, changing a clock speed for that blinking LED should be near instantaneous but more likely requires a rebuilding the whole project. Fundamentally the vendors don’t view their chips as something designed to run programs, and this legacy hardware design mentality plagues their whole business.

Something important here: Xilinx could and should have been where NVidia is today. They were certainly aware of the competitive accelerated computing market as early as 2005, and fundamentally failed to make a software architecture competitive with CUDA.

Before CUDA even existed I interned at Xilinx working on the beginnings of their HLS C compiler. My (decade older) fraternity brother led the C compiler team at Altera. We almost went into making a spreadsheet compiler for FPGA (my masters thesis) together but 2007 ended up being a terrible year to sell accelerated computing to Wall Street.

amirhirsch··on AI promised efficiency. Instead, it's making us work harder
The blogosphere (am I dating myself?) keeps bringing up the METR study (https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...) without really understanding the result. The guy with experience had a huge boost. You are reading the results wrong if your conclusion is this blog.

And that was before Claude Code.

amirhirsch··on Cerebras Code
i ended up getting it working through copying the transformer in this issue: https://github.com/musistudio/claude-code-router/issues/407

It hits the request per minute limit instantly and then you wait a minute.

amirhirsch··on Cerebras Code
API Error: 422 {"error":{"message":"Error from provider: {\"message\":\"body.messages.0.system.content: Input should be a valid string\",\"type\":\"invalid_request_error\",\"param\":\"validation_error\",\"code\":\"wrong_api_format\"}
amirhirsch··on Cerebras Code
the distinction is from weekly limits of claude code.
amirhirsch··on OpenAI's ChatGPT Agent casually clicks through "I am not a robot" verification
I don't know exactly what they do now, bloom filters was a thing then, also lots of heuristic approaches based on the bots we detected. the OP agent example actually would fail the very first test I deployed which looked for basic characteristics of the mouse movement

Here's a fun experiment for someone: 1) Give N people K fake credit cards to enter into a form, and have them solve a captcha 2) Take recorded keyboard and mouse data similar to the captcha 3) Train a neural network model to identify

I've been out of this for 6 years but I bet transformers rock this problem now.

← PreviousPage 3 of 13Next →