HNHacker News
TopNewBestAskShowJobs

francisjp

10 karma · joined September 18, 2026

submissionscomments
francisjp··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.

llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.

francisjp··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
To OP: great work on the release! I am generally interested in this kind of optimization work.

Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.

Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.

francisjp··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thanks for all of your exploration in public Simon.

Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.

Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: https://github.com/ggml-org/llama.cpp/pull/27461/changes