HNHacker News
TopNewBestAskShowJobs

zackangelo

375 karma · joined July 23, 2012

building mixlayer, zack at mixlayer.com
submissionscomments
zackangelo··on Open source inference time compute example from HuggingFace
Where did you see that? I thought they used an 8b model for their reward model?

> To guide our search strategies, we used RLHFlow/Llama3.1-8B-PRM-Deepseek-Data, an 8B reward model that has been trained using process supervision

zackangelo··on Ask HN: How is the M4 MacBook Pro Nano-texture screen for long coding so far?
Perhaps I don't have such a keen eye for things like font rendering, but the nano texture display has been a pure step forward for me. No negatives at all.
zackangelo··on ChatGPT Pro
A few thoughts:

* Will this be the start of enshittification of the base ChatGPT offering?

* There may also be some complementary products announced this month that make the $200 worth it

* Is this the start of a bigger industry trend of prices more closely aligning to the underlying costs of running the model? I suspect a lot of the big players have been running their inference infrastructure at a loss.

zackangelo··on The Rock VX Gas Canister Build (2022)
Interestingly, the actual prop was auctioned off in 2022 for more than $18K! [0]

[0] https://propstoreauction.com/lot-details/index/catalog/319/l...

zackangelo··on Samurai: Adapting Segment Anything Model for Zero-Shot Visual Tracking
I’ve been writing all of our transformer implementations in Rust using the Candle crate and it’s been great.

While dealing with CUDA and GPUs on servers is never a joy, deploying fully contained Rust binaries instead of a morass of python scripts has improved the situation for me significantly.

Getting Samurai running on Candle shouldn’t be that large of an undertaking. I believe there’s already a SAM implementation.

zackangelo··on Launch HN: Human Layer (YC F24) – Human-in-the-Loop API for AI Systems
Congrats on the launch Dex!
zackangelo··on Marshall Brain has died
So grateful that HSW existed when I was younger. As a teenager, I couldn't afford to get the timing belt and water pump replaced on my car so I had to figure out how to do it myself. I bought the service manual from AutoZone but I needed something to closer to an introduction to even be begin to understand it. He seemed to love explaining how car engines work and that series of articles was exactly what I needed at that time to get started.

RIP Marshall, I hope you knew what an inspiration you were.

zackangelo··on Using gRPC for (local) inter-process communication (2021)
Had to reach for a new IPC mechanism recently to implement a multi-GPU LLM inference server.

My original implementation just pinned one GPU to its own thread then used message passing between them in the same process but Nvidia's NCCL library hates this for reasons I haven't fully figured out yet.

I considered gRPC for IPC since I was already using it for the server's API but dismissed it because it was an order of magnitude slower and I didn't want to drag async into the child PIDs.

Serializing the tensors between processes and using the Servo team's ipc-channel crate[0] has worked surprisingly well. If you're using Rust and need a drop-in (ish) replacement for the standard library's channels, give it a shot.

[0] https://github.com/servo/ipc-channel

zackangelo··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
Related recent discussion on twitter: https://x.com/Teknium1/status/1858987850739728635

Looks like other folks get 80 tok/s with max batch size, that's surprising to me but vLLM is definitely more optimized than my implementation.

zackangelo··on Show HN: Rust library for numerical integration of real-valued functions
> Does the function oscillate over the region of integration? If it does, then make sure that the step size is chosen to be smaller than the wave length of the function.

Nyquist limit, but for numerical integration?

zackangelo··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
It’s been a minute so my memory might be off but I think when I ran 70b at fp16 it just barely fit on a 2x A100 80GB cluster but quickly OOMed as the context/kv cache grew.

So if I had to guess a 96GB H100 could probably run it at fp8 as long as you didn’t need a big context window. If you’re doing speculative decoding it probably won’t fit because you also need weights and kv cache for the draft model.

zackangelo··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
Ah, makes a lot more sense now.
zackangelo··on Extending the context length to 1M tokens
Would you care to share your prompts?

They posted a haystack benchmark in the blog post that seems too good to be true.

zackangelo··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
This is astonishingly fast. I’m struggling to get over 100 tok/s on my own Llama 3.1 70b implementation on an 8x H100 cluster.

I’m curious how they’re doing it. Obviously the standard bag of tricks (eg, speculative decoding, flash attention) won’t get you close. It seems like at a minimum you’d have to do multi-node inference and maybe some kind of sparse attention mechanism?

zackangelo··on You could have designed state of the art positional encoding
I'm surprised this is the case! I've been working on a rope implementation for my own project (needed to account for padding in unique situations) and even an off by one error usually causes the model to produce non-sensical output.
zackangelo··on ThunderKittens: Simple, fast, and adorable AI kernels
I’m working on an inference platform that allows for tokens to be appended to the context after some tokens have been generated. If there’s other sequences in the batch, it means they’ll have to be padded. Currently this means I can’t use FlashAttention because it doesn’t support arbitrary masks/padding masks… can ThunderKittens help me?
zackangelo··on Millions may rely on groundwater contaminated with PFAS for drinking water
Do you explicitly list that a PFAS test was omitted for a municipality? I looked up Austin and didn’t see it listed.
zackangelo··on Quantized Llama models with increased speed and a reduced memory footprint
I'd love it if you checked out what we've been working on.

It's still in early stages, but might be usable for something you're trying to build. Here's an example (this buffers the entire JSON object, but you can also gen as you go): https://docs.mixlayer.com/examples/json-output

zackangelo··on Quantized Llama models with increased speed and a reduced memory footprint
With mixlayer, because the round trip time to the model is so short, you can alternate between appending known tokens of the JSON output and values you want the model to generate. I think this works better than constraining the sampling in a lot of cases.

We haven’t built a state machine over JSON schema that uses this approach yet but it’s on the way.

zackangelo··on Serving 70B-Scale LLMs Efficiently on Low-Resource Edge Devices [pdf]
I've only read the abstract but they don't mention quantizing the weights or otherwise trying to shrink the model in any way.

They're claiming to be able to efficiently run larger models without loading the entire thing into GPU memory. If they're using the same weights, the same architecture and just using tensor parallel operations to perform the forward pass that would imply no loss in quality.

I'm sure there are trade-offs but they're not clear by just looking at the abstract.

zackangelo··on Show HN: Kameo – Fault-tolerant async actors built on Tokio
This limitation is common to most implementations of the actor model. In fact, I think a lot of people would consider it a feature, not a limitation because it allows you to reason about your concurrent behavior in a more straightforward way.
zackangelo··on Moshi: A speech-text foundation model for real time dialogue
We used candle[0], which uses cudarc and the metal crate under the hood. That means we run on nvidia hardware in production and can test locally on macbooks with smaller models.

I would certainly like to use non nvidia hardware but at this point it's not a priority. The subset of tensor operations needed to run the forward pass of LLMs isn't as large as you'd think though.

[0] https://github.com/huggingface/candle

zackangelo··on Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
I'm building this type of functionality on top of Llama models if you're interested: https://docs.mixlayer.com/examples/json-output
zackangelo··on Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
This is incorrect:

> With text-only inputs, the Llama 3.2 Vision Models can do tool-calling exactly like their Llama 3.1 Text Model counterparts. You can use either the system or user prompts to provide the function definitions.

> Currently the vision models don’t support tool-calling with text+image inputs.

They support it, but not when an image is submitted in the prompt. I'd be curious to see what the model does. Meta typically sets conservative expectations around this type of behavior (e.g., they say that the 3.1 8b model won't do multiple tool calls, but in my experience it does so just fine).

zackangelo··on Moshi: A speech-text foundation model for real time dialogue
Just trying to stay focused on launching first (https://docs.mixlayer.com) and keeping early customers happy, but would love to open source some of this work.

It'd probably be a separate crate from candle. If you haven't checked it out yet, mistral.rs implements some of these things (https://github.com/EricLBuehler/mistral.rs). Eric hasn't done multi-GPU inference yet, but I know it's on his roadmap. Not sure if it helped, but I shared an early version of my llama 3.1 implementation with him.

zackangelo··on Moshi: A speech-text foundation model for real time dialogue
Yeah, I’ve had to rewrite continuous batching and other scheduling logic. That and multi-GPU inference have been the hardest things to build.

I’ll need to get paged attention working as well, but I think I can launch without it.

zackangelo··on Moshi: A speech-text foundation model for real time dialogue
Their inference server is written in Rust using huggingface’s Candle crate. One of the Moshi authors is also the primary author of Candle.

We’ve also been building our inference stack on top of Candle, I’m really happy with it.

zackangelo··on Show HN: Hestus – AI Copilot for CAD
I wonder if the prevalence of coding LLMs will bring renewed attention to code-based modeling systems like OpenSCAD. It seems much easier to get them to generate code vs. translating into GUI interactions or modifying some other internal state directly.
zackangelo··on Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights
Is this tied to a specific framework like pytorch or an inference server like vLLM?

Our inference stack is built using candle in Rust, how hard would it be to integrate?

zackangelo··on We're Cutting L40S Prices in Half
L40S has 48GB of RAM, curious how they're able to run Llama 3.1 70B on it. The weights alone would exceed this. Maybe they mean quantized/fp8?

I just had to implement GPU clustering in my inference stack to support Llama 3.1 70b, and even then I needed 2xA100 80GB SXMs.

I was initially running my inference servers on fly.io because they were so easy to get started with. But I eventually moved elsewhere because the prices were so high. I pointed out to someone there that e-mailed me that it was really expensive vs. others and they basically just waved me away.

For reference, you can get an A100 SXM 80GB spot instance on google cloud right now for $2.04/hr ($5.07 regular).

← PreviousPage 3 of 6Next →