HNHacker News
TopNewBestAskShowJobs

anerli

225 karma · joined June 4, 2023

submissionscomments
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Will be adding this one very soon!
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).

llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
These numbers are just estimates based on your hardware and may differ from actual performance. It's hard to get an accurate measurement until it's actually downloaded and running. They also don't account for gains from speculative decoding. Working on changes to make this more clear.

For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
While it's great to see tok/s go up as high as possible, I think it's important to consider the actual usability of these models when you quantize down to something like 2-bit. From what we've tested it seems like going below 4-bit quickly leads to serious issues with thinking, tool calls, and overall model coherence.

Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Yeah we heavily leverage coding agents for optimizing our kernels. Since it's highly verifiable and takes time to measure we often leave multiple running and improving performance on different model architectures.

Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
No specific benchmarks for Strix Halo yet but planning to release more results for different hardware and models soon!

Qwen3.8-Flash-Next support will also be added very soon.

Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.

Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
As mentioned here https://news.ycombinator.com/item?id=49912327 we'll eventually build an inference cloud for hybrid workloads. For now though, we're focused on making the inference engine great!
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thanks for reporting this issue - assessing is not supposed to take more than a minute or so. This is not strictly necessary but filters out models that don't fit in memory and gives speed estimates. This should ideally be a very short step so skipping hopefully wouldn't feel necessary if we patch this.

Could you share your hardware and OS details to help us identify what might be the issue here?

There's also a github issue open on this topic if you want to leave a comment there: https://github.com/magnitudedev/magnitude/issues/142

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
From the looks of it, this seems focused on datacenter/batch inference, and doesn't tune its kernels to the specific hardware and workload where inference is being run like Magnitude does.

Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Temperature control is a good callout! Feel free to open an issue for any features that you'd like to see in Magnitude

https://github.com/magnitudedev/magnitude/issues

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Issues should already be open for anyone to report!

https://github.com/magnitudedev/magnitude/issues

Let me know if you keep running into problems for some reason

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
We are actively benchmarking our Vulkan kernels to ROCm implementations in other engines to ensure that we can reach the performance ceiling with them. Vulkan is much more portable and also works on non-AMD hardware even though it can be more awkward to write kernels for. If we find that Vulkan is not sufficient for reaching the same performance as ROCm, we'll consider adding it as a backend
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
I would say the overall idea of trying to achieve performant inference for agent workloads is the strongest commonality with Wafer.

It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
We have a retrieval benchmark based on RULER which we've been using to ensure that the model maintains complete awareness of the full context window.

All our benchmarks are open source so you can check it out here if you'd like: https://github.com/magnitudedev/magnitude/blob/main/inferenc...

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.

However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
We support Vulkan as well, we just didn't mention it in the benchmark. When AMD or Strix Halo is detected the engine will use Vulkan.

Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thanks for reporting the issue.

Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? https://github.com/magnitudedev/magnitude/issues

As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Yeah, generally being able to focus on specific architectures lets you optimize better for those. However models of the same family (for example Qwen 3.5/3.6/ some 3.8 models) share the same architecture, so you only need to optimize once and new models can use the same kernels. There's also shared algorithms and kernels that can be optimized once and used across different families, so it's a bit nuanced.

We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Yeah these are all things that we directly tackle!

Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Compared to MLX - we've done some rough benchmarking and we are outperforming any of the MLX-based engines we've compared to so far. Going to do more in depth benchmarking and release it soon.
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...

anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
anerli··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.
anerli··on The M×N problem of tool calling and open-source models
In my experience it's actually very doable to do reliable tool calling with a generic response format across models. You just need to disable native tool calling completely and provide a clearly defined response/tool format that conforms well to pretraining across a variety of models (e.g. XML-like syntaxes).

For example: ``` <think>Let me take a look at that</think> <read path="foo.txt"/> ```

The hard part is building a streaming XML parser that handles these responses robustly, can adjust for edge cases, and normalizes predictable mishaps in history in order to ensure continued response format adherance.

anerli··on Code Mode: the better way to use MCP
yup, we've been using this approach with our product to make composing different integrations easier for the LLM and also give it the flexibility of code. Main difference is we use quick-js instead of v8 isolates. Seeing a TS interface instead of ugly JSON schema and simply writing code is far simpler for the LLM
anerli··on Pure-vision browser agent scores 94% on WebVoyager (SOTA)
Hey HN, Anders and Tom from Magnitude (YC S25) here. On our last Show HN post about our open-source browser agent, someone left a comment - "there are multiple similar projects like this posted here daily, and this one likely isn't the best". So we asked ourselves, are they right? We decided to run on WebVoyager (a well known benchmark for browser agents) to test ourselves. We scored 94%, beating all other browser agents and making Magnitude state-of-the-art.

You can view the entire run here: https://magnitude-webvoyager.vercel.app/

The original WebVoyager benchmark was meant to demonstrate a new technique for interacting with the browser by annotating the DOM. Since then, vision models have come a long way in terms of accuracy and visual understanding. Our pure-vision approach with our framework and today's models surpasses the hybrid DOM strategies used by the original WebVoyager paper and other agents like browser-use.

So why does pure-vision beat hybrid DOM approaches?

- Generalizes far better - handles canvas elements, iframes, drag-and-drop, precise text selection, and many other scenarios elegantly where hybrid DOM would struggle and need to implement hacks for those cases to work

- Easier for the LLM - we think LLM performance is roughly proportional to prompt clarity. If the prompt contains a crowded screenshot with loads of colored boxes + a long list of element labels and is asked to pick one, vs given a clean screenshot + where do you want to click - the latter seems far easier

We believe another reason for our success is that we can still hook into the browser as needed. We can use browser-native actions like tab switching, can look at network traffic to know when a page is ready, or use the DOM for other purposes like data extraction. Computer use agents like Operator or Claude Computer Use on the other hand are limited to generic mouse and keyboard controls.

It's worth mentioning that WebVoyager is a strange and flawed benchmark. It contains many tasks that depend on the current date (and need their dates updated), tasks that depend on the time of day, and some tasks that are impossible or too ambiguous to properly evaluate. In the repo we detailed exactly the patches we made to the original WebVoyager benchmark such that each task is at least theoretically possible.

Why does this all matter? People are trying to adopt agents for real use cases, but they often fail to make it to production. We want to enable developers to build with production-ready browser agents - which is why it's important to get the fundamental interaction paradigm right. We think this benchmark is a step in the right direction, showing that pure-vision has best-in-class performance in the browser domain. Curious to hear what others think about this, would love to get your feedback!

anerli··on Show HN: Magnitude – Open-source AI browser automation framework
So there’s a very big difference in the sort of vision approach that browser-use does vs. what we do

browser-use is still strongly coupled to the DOM for interaction because of the set-of-marks approach it uses (for context - those little rainbow boxes you see around the elements). This means it’s very difficult to get it to reliably do interactions outside of straightforward click/type like drag and drop, interacting with canvas, etc.

Since we interact based purely on what we see on the screen using pixel coordinates, those sort of interactions are a lot more natural to us and perform much more reliably. If you don't believe me, I encourage you to try to get both Magnitude and browser-use to drag and drop cards on a Kanban board :)

Regardless, best of luck!

anerli··on Show HN: Sink – Sync any directory with any device on your local network
^ syncthing is nice
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
We do believe in a hybrid approach where a fast/deterministic representation is saved - but think there is a more seamless way were the framework itself is high level and manages these details by caching the underlying actions that can run
Page 1 of 3Next →