HNHacker News
TopNewBestAskShowJobs

zackangelo

375 karma · joined July 23, 2012

building mixlayer, zack at mixlayer.com
submissionscomments
zackangelo··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
GPT-OSS will run even faster on Blackwell chips because of its hardware support for fp4.

If anyone is working on training or inference in Rust, I'm currently working on adding fp8 and fp4 support to cudarc[0] and candle[1]. This is being done so I can support these models in our inference engine for Mixlayer[2].

[0] https://github.com/coreylowman/cudarc/pull/449 [1] https://github.com/huggingface/candle/pull/2989 [2] https://mixlayer.com

zackangelo··on SQLx – Rust SQL Toolkit
Is something like SeaQuery[0] what you're talking about?

[0] https://github.com/SeaQL/sea-query/

zackangelo··on Qwen3-Coder: Agentic coding in the world
Draft model doesn’t degrade quality!
zackangelo··on Show HN: We made our own inference engine for Apple Silicon
We also wrote our inference engine in rust for mixlayer, happy to answer any questions from those trying to do the same.

Looks like this uses ndarray and mpsgraph (which I did not know about!), we opted to use candle instead.

zackangelo··on Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
Typically a combination of expert level parallelism and tensor level parallelism is used.

For the big MLP tensors they would be split across GPUs in a cluster. Then for the MoE parts you would spread the experts across the GPUs and route to them based on which experts are active (there would likely be more than one if the batch size is > 1).

zackangelo··on Launch HN: Morph (YC S23) – Apply AI code edits at 4,500 tokens/sec
For anyone more curious about how this works, Fireworks wrote a blog post about it last year (I think):

https://fireworks.ai/blog/cursor

zackangelo··on Life of an inference request (vLLM V1): How LLMs are served efficiently at scale
In your forward pass section you give a lot of emphasis to FlashAttention, but it might be worth mentioning Paged Attention as well (which was the paper written by the vLLM authors and I believe was the genesis of the project). PA-style block tables are now supported in most fused attention kernels, but vLLM originally came up with it and it's the main reason why vLLM has such high throughput!
zackangelo··on Introducing Gemma 3n
Yes, this is true. A lot of times labs will hold back necessary infrastructure pieces that allow them to train huge models reliably and on a practical time scale. For example, many have custom alternatives to Nvidia’s NCCL library to do fast distributed matrix math.

Deepseek published a lot of their work in this area earlier this year and as a result the barrier isn’t as high as it used to be.

zackangelo··on Introducing Gemma 3n
In the README of the linked library they have a code snippet showing how to have a conversation with the model.

Also, even if it were for fine tuning, that would require an implementation of the model’s forward pass (which is all that’s necessary to run it).

zackangelo··on Introducing Gemma 3n
Your reply adds more confusion, imo.

The inference code and model architecture IS open source[0] and there are many other high quality open source implementations of the model (in many cases contributed by Google engineers[1]). To your point: they do not publish the data used to train the model so you can't re-create it from scratch.

[0] https://github.com/google-deepmind/gemma [1] https://github.com/vllm-project/vllm/pull/2964

zackangelo··on OpenAI dropped the price of o3 by 80%
Is that input tokens or output tokens/s?
zackangelo··on Cloud Run GPUs, now GA, makes running AI workloads easier for everyone
I think it’s a typo, looks pretty close to their 8xH100 prices.
zackangelo··on DeepSeek may have used Google's Gemini to train its latest model
There might be a plateau coming but I’m not sure that will be the reason.

It seems counterintuitive but there is some research suggesting that using synthetic data might actually be productive.

zackangelo··on Why DeepSeek is cheap at scale but expensive to run locally
H200s are pretty easy to get now. If you switched I'm guessing you'd get a nice bump because the nccl allreduce on the big mlps wouldn't have to cross infiniband.
zackangelo··on Why DeepSeek is cheap at scale but expensive to run locally
This definitely happens, and I'm surprised it's not talked about more often. Some attention kernels are more susceptible to this than others (I've found that paged attention is better than just naive attention, for example).
zackangelo··on Net-Negative Cursor
I've had excellent experience with several models writing Rust. Wonder if there's just a particular issue with Tauri? I'm primarily writing code on top of the Candle ML framework.
zackangelo··on Deepseek R1-0528
Yeah, to run the full precision model you need either two 8xH100 nodes connected via Infiniband or one 8xH200 node or one 8xB200 node.

Not for the GPU poor, to be sure.

zackangelo··on How Does Claude 4 Think? – Sholto Douglas and Trenton Bricken
Are you familiar with min_p sampling?

Kind of funny that it was introduced randomly on Reddit a couple of years ago instead of in a journal or something[0]. But I believe it's widely implemented and used now.

[0] https://www.reddit.com/r/LocalLLaMA/comments/17vonjo/your_se...

zackangelo··on An intro to DeepSeek's distributed file system
The 3FS chunk engine is written in Rust.
zackangelo··on Multi-Token Attention
What codec were you using for the audio data?
zackangelo··on DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
My go to is Programming Massively Parallel Processors by Wen-Mei Hwu, excellent really approachable introduction. [0]

[0] https://a.co/d/9fmbZqg

zackangelo··on Ask HN: Is anyone still using Dreamweaver?
Allaire Homesite anyone?

It was a sad day for me when it got bought and integrated into Dreamweaver.

zackangelo··on Dismissed Nuclear Bomb Specialists Recalled by Energy Department
Something something Chesterton s fence
zackangelo··on We Were Wrong About GPUs
would love for you to test a serverless llm product i'm working on, zack [at] mixlayer.com
zackangelo··on WASM-Native Orchestration
I've been developing on top of wasm (wasmtime, specifically) for several years now.

I personally have my doubts about how broadly the components specification (which this cloud platform seems to depend on) will be adopted. Maybe I'm just not very smart, but it feels like one of the things I loved the most about WebAssembly, its simplicity, is being lost.

To get even basic things done with WASI, you now have to figure out which flavor to pull in (preview 1, preview 2, ... need any async [0], that's coming preview 3?), learn a new IDL language [1] that defines the component's spec, figure out its codegen toolchain [2]. It's all very reminiscent of COM, CORBA and all the other times this has been tried in the past.

In any case, wasmtime is an amazing project and I'm grateful for it. They've already had to bifurcate the library to support WASI and components. I just hope non-WASI/component code retains its status as a first class citizen in the project.

[0] https://docs.google.com/presentation/d/1MNVOZ8hdofO3tI0szg_i... [1] https://github.com/WebAssembly/component-model/blob/main/des... [2] https://github.com/bytecodealliance/wit-bindgen

zackangelo··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
The authors specifically recommend against using a system prompt in the model card.
zackangelo··on All Garmin Connect services are down
Was it at Strangeloop by any chance? I think I remember that talk too!
zackangelo··on Why HNSW is not the answer and disk-based alternatives might be more practical
lol good to know! I've had the luxury of only needing the first one :)
zackangelo··on Why HNSW is not the answer and disk-based alternatives might be more practical
fwiw faiss, although a bit unwieldy, has an optimized full scan search built into it as well
zackangelo··on Why HNSW is not the answer and disk-based alternatives might be more practical
This is so true. A plain old exhaustive SIMD-optimized similarity search will do just fine in many cases and not have any of the approximation tradeoffs of HNSW.
← PreviousPage 2 of 6Next →