HNHacker News
TopNewBestAskShowJobs

fooblaster

747 karma · joined April 25, 2016

submissionscomments
fooblaster··on Grand Theft Oil Futures: Insider traders keep making a killing at our expense
Airlines can't raise the prices of tickets sold months ago. There is still financial reason to hedge.
fooblaster··on Amazon chips no longer just a side dish, they're a $20B biz
I think the revenue claims around future trainium spend is complete bullshit.
fooblaster··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
Honestly, this is the AI software I actually look forward to seeing. No hype about it being too dangerous to release. No IPO pumping hype. No subscription fees. I am so pumped to try this!
fooblaster··on Tesla 'Full Self-Driving' crashed through railroad gate seconds before train
Tesla has not pulled the driver. It's just not comparable.
fooblaster··on Artemis II crew see first glimpse of far side of Moon [video]
It's fine to not be interested, but this time one of the astronauts is black
fooblaster··on Solar and batteries can power the world
where are you? that is a massive amount of solar in any place at a reasonably low latitude. Is your house enormous or are you heating your house with resistive heating?
fooblaster··on First 6 days of Iran war cost $11.3B
This thing is far from over. Iran will indefinitely be able to block the straight. The us will be stuck in this defensive position for months, until it pulls out and effectively loses the war.

It's clear we are going to lose, because we cannot topple the regime without putting troops on the ground, which we will never do. Setting that as a war aim doomed this whole effort from the start.

fooblaster··on MacBook Pro with M5 Pro and M5 Max
what inference runtime are you using? You mentioned mlx but I didn't think anyone was using that for local llms
fooblaster··on 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern
Again, it wasn't exactly a huge sink of resources. There was no genius gamble from jensen like you are suggesting. I suspect your view here is intrinsically tied to your need to feel like you and others who are in your position are responsible for your own success, when in fact it's mostly about luck.
fooblaster··on 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern
CUDA was profitable very early because of oil and gas code, like reverse time migration and the like. There was no act of incredible foresight from jensen. In fact, I recall him threatening to kill the program if large projects that made it not profitable failed, like the Titan super computer at oak ridge.
fooblaster··on 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern
It was definitely luck, greg. And Nvidia didn't invent deep learning, deep learning found nvidias investment in CUDA.
fooblaster··on Defining Safe Hardware Design [pdf]
I was really happy to see that blue spec was fully open sourced in recent years. Does anyone have experience with a non trivial project with it? Does it have any traction anymore in real silicon development.
fooblaster··on Apple picks Gemini to power Siri
calling neural engine the best is pretty silly. the best perhaps of what is uniformly a failed class of ip blocks - mobile inference NPU hardware. edge inference on apple is dominated by cpus and metal, which don't use their NPU.
fooblaster··on Developing a BLAS Library for the AMD AI Engine [pdf]
Looks like they have made some progress on a native model in recent months: https://github.com/amd/IRON/tree/devel
fooblaster··on CES 2026: Taking the Lids Off AMD's Venice and MI400 SoCs
sms is the Nvidia definition of processor, and cuda device properties returns it, not anything else. If you want a marketing number, use cuda cores, it doesn't consistently match to anything in the hardware design.
fooblaster··on CES 2026: Taking the Lids Off AMD's Venice and MI400 SoCs
b200 is 148 sms, so no
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
The versal stuff isn't really an FPGA anymore. The chips have PL on them, but many don't. The consumer NPUs from AMD are the same versal aie cores with no PL. They just aren't configurable blocks in fabric anymore and don't have the same programming model. So I'm not contradicting myself here.

That being said, versal aie for ml has been a terrible failure. The reasons for why are complicated. One reason is because the memory hierarchy for SRAM is not a unified pool. It's partitioned into tiles and can't be accessed by all cores. additionally, access of this SRAM is only via dma engines and not directly from the cores. Thirdly, the datapaths for feeding the VLIW cores are statically set, and require a software configuration to change at runtime which is slow. Programming this thing makes the cell processor look like a cakewalk. You gotta program dma engines, you program hundreds of VLIW cores, you need to explicitly setup on chip network fabric. I could go on.

Anyway, my point is FPGAs aren't getting ML slices. Some FPGAs do have a completely separate thing that can do ML, but what is shipped is terrible. Hopefully that makes sense.

fooblaster··on Why is the Gmail app 700 MB?
10 MB mail app. 690 MB local llm to write snarky emails for you.
fooblaster··on Corundum – open-source FPGA-based NIC and platform for in-network compute
wow, say more..
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
I'd like to know more. I expect these systems are 8xvh1782. Is that true? What's the theoretical math throughput - my expectation is that it isn't very high per chip. How is performance in the prefill stage when inference is actually math limited?
fooblaster··on Developing a BLAS Library for the AMD AI Engine [pdf]
This architecture is likely going to be a dead end for AMD. It has been in the wild for several years, yet still has no open programming model, multiple compiler stacks with poor software support. I find it likely that AMD drops this architecture and unifies their ML support around their GPGPU hardware.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
The amd npu and versal ML tiles (same underlying architecture) have been an complete failure. Dynamic programming models like cu tile do not work on them at all, be cause they require an entirely static graph to function. AMD is going to walk away from their NPU architecture and unify around their GPU IP on inference products in the future.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
I wouldn't trust any benchmarks on the vendors site. Microsoft went down this path for years with FPGAs and wrote off the entire effort.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
Yep, but they are still 50x faster than any fpga.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
Show me a single FPGA that can outperform a B200 at matrix multiplication (or even come close) at any usable precision.

B200 can do 10 peta ops at fp8, theoretically.

I do agree memory bandwidth is also a problem for most FPGA setups, but xilinx ships HBM with some skus and they are not competitive at inference as far as I know.

fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
Yeah, I wouldn't have guessed it would be helping me write systemverilog.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
FPGAs will never rival gpus or TPUs for inference. The main reason is that GPUs aren't really gpus anymore. 50% of the die area or more is for fixed function matrix multiplication units and associated dedicated storage. This just isn't general purpose anymore. FPGAs cannot rival this with their configurable DSP slices. They would need dedicated systolic blocks, which they aren't getting. The closest thing is the versal ML tiles, and those are entire peoxessors, not FPGA blocks. Those have failed by being impossible to program.
fooblaster··on TinyTinyTPU: 2×2 systolic-array TPU-style matrix-multiply unit deployed on FPGA
Great! How do you program it?
fooblaster··on OpenAI's cash burn will be one of the big bubble questions of 2026
Well, anthropic just purchased a million TPUs from Google because even with a healthy margin from Google, it's far more cost effective because of Nvidia's insane markup. That speaks for itself. Nvidia will not drop their margin because it will tank their stock price. it's half of the reason for all this circular financing - lowering their effective margin without lowering it on paper.
fooblaster··on OpenAI's cash burn will be one of the big bubble questions of 2026
If you think they are going to catch up with Google's software and hardware ecosystem on their first chip, you may be underestimating how hard this is. Google is on TPU v7. meta has already tried with MTIA v1 and v2. those haven't been deployed at scale for inference.
← PreviousPage 3 of 10Next →