HNHacker News
TopNewBestAskShowJobs

zackangelo

376 karma · joined July 23, 2012

building mixlayer, zack at mixlayer.com
submissionscomments
zackangelo··on Speculative Decoding in vLLM on AMD GPUs
I just finished overhauling our speculative decoding implementation for Mixlayer, so maybe I can help.

I think the piece of information that might make this click for you is the model outputs the probability distribution for all intermediate tokens even during prefill.

So for example, let's say you prefill the prompt "The quick brown fox" (and for the sake of simplicity, let's say each word is a single token). The model outputs a tensor that is [4, $vocabulary_size]. The first dimension is a token index into the input and the 2nd dimension assigns a probability to each token in the vocabulary. So even during prefill, we can look at the prediction logits for all of the intermediate tokens. That is, we can look at what the model would have predicted after "quick" and "brown", not just the tail token "fox".

In the single token autoregressive case, we just look at the next token prediction for "fox". But in the speculative decoding case we can use this information to compare the distribution of the draft model against the target model. In the greedy decoding case (ie, no sampling) we just make sure the highest probability token matches in draft and target. If we have sampling params like temperature and top-P, we have to apply something called Leviathan rejection sampling to the distribution. This basically allows us make sure the distribution is the same even if the exact probabilities are not and accept or reject draft tokens on that.

zackangelo··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Something a lot of model providers don't talk about: any time an engine uses speculative decoding the throughput will depend on how much your output token distribution matches what the draft model was trained on.

The DFlash2 draft model we're using was trained on a lot of code, so if you use it in a coding agent you'll probably notice it run a lot faster (we've seen it break 300 tok/s).

zackangelo··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
just added 8 more H200s to the cluster, if you (or anyone else) runs into issues please feel free to drop me a message: zack at mixlayer.com
zackangelo··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
apologies we just got a sudden burst of new users and traffic, it's scaling up now.
zackangelo··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).

https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.

zackangelo··on Rivian OS 2
What was wrong with your R1S?

I had an R1T Launch Edition for a little over a year and it was hands down one of the best cars I’ve ever owned. The only issue I had was with the fob proximity sensor. The only reason I sold it was it was a bit too big for me.

zackangelo··on DFlash 2: Keep Drafting Parallel
DFlash is lossless so this would be a bug in the implementation if it is indeed a regression against the target model.
zackangelo··on Performance per dollar is getting faster and cheaper
Blackwell supports nvfp4 natively.
zackangelo··on Two Qwen3 models on one DGX Spark: the residency math
what was the concurrency limitation? that node should be able to support a lot more
zackangelo··on Kimi K2.7-Code: open-source coding model with better token efficiency
I don't believe safetensors has a native int4 dtype, so they packed 4 int4s into a bf16 in this checkpoint.
zackangelo··on The real cost of owning a home
If you're in SF and weighing this decision, it's easy to get tilted in the buy direction because the rental stock is so horrific. Landlords have very little incentive to update properties or provide basic amenities that people take for granted in other major cities (good luck getting a washer/dryer).
zackangelo··on Qwen3.7-Max: The Agent Frontier
With the 3.5 release, the Plus model was just a rebrand of the open weight 397B. But I suspect that will change going forward. They haven’t released the weights for 3.6 but they did make it available through a few US providers.
zackangelo··on I’ve joined Anthropic
absolutely not, take Kimi K2.6 for a spin
zackangelo··on Mistral Medium 3.5
Isn't Kimi K2.6 natively INT4?
zackangelo··on HashiCorp co-founder says GitHub 'no longer a place for serious work'
I don’t think this is true across Blizzard. Overwatch is the best it’s ever been.
zackangelo··on Parallel agents in Zed
I give them a try about twice a year. I write a lot of Rust which should be squarely in their wheelhouse.

This last time I was pleasantly surprised to find they mostly fixed their SSH remote editing support. But then it started truncating rustc inline error messages and I couldn’t figure out how to view the whole thing easily. When you’re just trying to get something done little bits like this can add up quickly. Punted back to Cursor for now.

zackangelo··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
They are but the IDE needs to be integrated with them.

Qwen specifically calls out FIM (“fill in the middle”) support on the model card and you can see it getting confused and posting the control tokens in the example here.

zackangelo··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
17b per token. So when you’re generating a single stream of text (“decoding”) 17b parameters are active.

If you’re decoding multiple streams, it will be 17b per stream (some tokens will use the same expert, so there is some overlap).

When the model is ingesting the prompt (“prefilling”) it’s looking at many tokens at once, so the number of active parameters will be larger.

zackangelo··on GPU memory snapshots: sub-second startup (2025)
This uses Nvidia’s CUDA snapshot API under the hood, but you have to pair it with a host side snapshot as well. Modal uses gVisor for this, which is notoriously high overhead.

Does anyone know of a more efficient alternative if you’re running a trusted container?

zackangelo··on macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
You’re right I misunderstood.

I’m not sure if it would be of much utility because this would presumably be for tensor parallel workloads. In that case you want the ranks in your cluster to be uniform or else everything will be forced to run at the speed of the slowest rank.

You could run pipeline parallel but not sure it’d be that much better than what we already have.

zackangelo··on macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
Sparks are built for this and actually have Connect-X 7 NICs built in! You just need to get the SFPs for them. This means you can natively cluster them at 200Gbps.
zackangelo··on macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
No you use tensor parallelism in both cases.

The way it typically works in an attention block is: smaller portions of the Q, K and V linear layers are assigned to each node and are processed independently. Attention, rope norm etc is run on the node-specific output of that. Then, when the output linear layer is applied an "all reduce" is computed which combines the output of all the nodes.

EDIT: just realized it wasn't clear -- this means that each node ends up holding a portion of the KV cache specific to its KV tensor shards. This can change based on the specific style of attention (e.g., in GQA where there are fewer KV heads than ranks you end up having to do some replication etc)

zackangelo··on Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
What 1T parameter base model have you seen from any of those labs?
zackangelo··on NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference
Wouldn't you be able to test nccl if you had 2 of these?
zackangelo··on Launch HN: LlamaFarm (YC W22) – Open-source framework for distributed AI
Just a bit of feedback:

> Instead of one brittle giant, we orchestrate a Mixture of Experts…

“mixture of experts” is a specific term of art that describes an architectural detail of a type of transformer model. It’s definitely not using smaller specialized models for individual tasks. Experts in an MoE model are actually routed to on a per token basis, not on a per task or per generation basis.

I know it’s tempting to co-opt this term because it would fit nicely for what you’re trying to do but it just adds confusion.

zackangelo··on Apps SDK
Because it depends on how much better “best” is. If it’s only incrementally better than open source models that have other advantages, why would you bother?

OpenAI’s moat will only come from the products they built on top. Theoretically their products will be better because they’ll be more vertically integrated with the underlying models. It’s not unlike Apple’s playbook with regard to hardwares and software integration.

zackangelo··on From multi-head to latent attention: The evolution of attention mechanisms
Not quite a frontier model but definitely built by a frontier lab: Grok 2 was recently open sourced and I believe it uses a fairly standard MHA architecture with MoE.
zackangelo··on Mosh Mobile Shell
I feel a bit silly for not noticing this before. Over the last year or so I've often wondered when ssh added protocol-level support for session resume. I'd open my laptop on a new network and everything would be ready to go. But of course, it's nothing to do with ssh, it's just that I started using tailscale.
zackangelo··on Writing Speed-of-Light Flash Attention for 5090 in CUDA C++
Curious what issues you were having. The kernel should compile natively if you pass nvcc the correct arch flags, although it probably won't take advantage of any new hardware features.
zackangelo··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
GPT-OSS will run even faster on Blackwell chips because of its hardware support for fp4.

If anyone is working on training or inference in Rust, I'm currently working on adding fp8 and fp4 support to cudarc[0] and candle[1]. This is being done so I can support these models in our inference engine for Mixlayer[2].

[0] https://github.com/coreylowman/cudarc/pull/449 [1] https://github.com/huggingface/candle/pull/2989 [2] https://mixlayer.com

Page 1 of 6Next →