HNHacker News
TopNewBestAskShowJobs

adityazero

26 karma · joined July 28, 2026

submissionscomments
adityazero··on Vx – One Language, Every Chip
> Is the intent that applications developed with this are compiled for target hardware on a machine-specific basis?

Vx gives you the power to do so, but one can be conservative by putting all the values in the fleet/*.vx files to extreme values. An analogy to that would be when we compile for X86-64, we can specify `-mcpu` otherwise by default the compiler picks a conservative backend.

Vx allows one to define compute elements(number of cores, etc), memory elements (number of distinct memory elements, hierarchy if any) and their relationship (placement, ownership etc). Together they define the toplogoy of a machine. Toplogy goes into machine description files (see https://github.com/vx-lang/Vx/tree/main/fleet) and the compiler uses them to 'monomorphize' the program based on that topology.

adityazero··on Vx – One Language, Every Chip
That story is buried in a long chat thread between me and Gemini.
adityazero··on Gemini Hacked Three Companies in First Known Breakout by Google's AI
After all of the others were done with hacking? There was a point in time when it was giving some publicity, it is a bit late IMO.
adityazero··on Nvidia announces native GPU programming in Rust
Rust/C/C++ all have a similar memory model, they were designed for a linear memory model. a `float` does not carry the provenance (a float out of cudaMalloc or a malloc look the same to the rest of the program). This is the fundamental problem that very few (e.g., Vx vxlang.org) are trying to address.
adityazero··on Rust SIMD on the GPU
Really like that the barrier scopes and shuffle controls are type-level, that is a lot of footguns turned into compile errors. Once you lower to PTX, does any of that structure survive for ptxas to optimize on, or is it fully erased by then? Also curious whether the strip mining abstraction fights the register allocator at higher lane counts, or if it stays friendly.
adityazero··on How I use LLMs to learn complex topics
I just ask LLM to to give me a short tutorial on the thing i want to learn but i dont want to start from chapter-1, that is boring. so i upload my resume and say that make this tutorial based on the skillsets i already have.

btw, this is a much better experience than watching Youtube videos for learning.

adityazero··on AMD acquires Taalas to boost inference performance by etching models in silicon
They are busy capturing market first, and I think that makes more sense. 'premature optimization etc.'
adityazero··on AMD acquires Taalas to boost inference performance by etching models in silicon
There are market verticals where this makes a lot of sense. Embedded systems and IoT devices comes to mind.

Even in data center space, I believe several layers can use fixed weights and remaining layers will compensate for the variations. Power savings will be huge so I think there is an incentive to do more of these.

adityazero··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
All of these would happily come back if Sundar left.
adityazero··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
> sundar, softly: “we can create a permission working group.”

so true. i have attended meetings to decide on meeting topics for the next half.

adityazero··on [dead]
NVIDIA runs the other way. The instruction set is documented, but the `ptxas` compiler is not. And NVIDIA's documentation stops higher up than people assume: PTX is a virtual ISA, but SASS is published simply as a list of opcode names.

So I went through the open compiler sources, OpenXLA and the Mosaic dialect inside JAX. They name the parts of the machine they generates code for, and there is useful information in there for anyone writing TPU kernels.

`tpu.log` is `printf`. A tag string plus whatever values you hand it. Next to it sit log_buffer, and trace_start/trace_stop for timing regions. I feel kernel debugging on TPU has a worse reputation than it deserves.

`tpu.weird` takes an `f32` and returns a `bool`. JAX's lowering evaluates `http://lax.is_finite` is its negation. So the TPU has a hardware predicate for "this float is weird."

And the MXU is a FIFO. Push the weights, stream the activations, pop the result. Weight-stationary, more than one per core, 32-bit accumulator enforced by the verifier. "bf16 in, fp32 accumulate" is the only shape the unit offers.

90 operations, eight memory spaces, and a lot of guidance to developers.