HNHacker News
TopNewBestAskShowJobs

philipportner

56 karma · joined July 5, 2025

ml systems, compilers, databases phd at tu berlin
submissionscomments
philipportner··on OpenArch – PyTorch implementations of modern LLM architectures
Yes. https://inferencex.semianalysis.com provides some comparisons wrt. certain cost metrics. Some commonly used ones are vLLM, TensorRT-LLM, and SGLang. These three at least are open source, and all come with an Apache 2.0 license.

Some providers also have to implement their own engines, e.g., Cerebras has their own inference serving stack for their wafer-scale chips, as does Google for their TPUs (XLA compiler).

philipportner··on Ask HN: Why can I only downvote select comments?
Wasn't aware of all of those, thanks for sharing.
philipportner··on LLMs are making me lose my savviness
> Are Rubber Ducks offloading thinking?

Just that an actual rubber duck doesn’t do anything. You solve the problem you have by talking, and in doing so, thinking, to come up with a solution, an idea, or gain better understanding.

After that you either implement something yourself or have learned something.

With an LLM you offload all of that, the only thing you still do is tell it what the problem is. The agentic duck does the rest and you look at the output.

Even if you have to argue, you argue without having gone through the steps to gain anything yourself.

philipportner··on Don't Paste the AI, please
Linked a the bottom of the post is the angry version https://dontpastetheai.com/angry/
philipportner··on A third world engineer responds to “RISC-V: They should have known better”
How do you keep up with such information? Any sources you could recommend? Closest I know would be SemiAnalysis
philipportner··on Accelerating GPT-5.6 Sol Ultrafast
Good point, thanks! I haven't been keeping up with most of the new model internals.
philipportner··on Accelerating GPT-5.6 Sol Ultrafast
You'd need hundreds of GB alone for the KV cache of each user. For something like LLama 3 405B you need ~67GB at ~130k tokens. A single CS-3 has 44GB on-chip sram.

So, afaik, Cerebras are optimizing for ultra-low latency batch=1 inference.

https://newsletter.semianalysis.com/p/cerebras-faster-tokens... goes quite in-depth.

philipportner··on Show HN: I spent 2 years designing a mechanical Magic Keyboard
Congrats on the great job with the Altar II. If I didn't already have too many keyboards, I'd hop on the Kickstarter! FWIW, I fully agree with your opinion on including a `half working` fingerprint sensor.
philipportner··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
if you assume that training requires about 3x the compute of inference (one forward pass, one backward pass, parameter updates), and we take DeepSeek-V3 since their numbers are public.

they used ~14.8 trillion tokens with about 2.66 million GPU hours. 14.8 * 3 = 44.4 t inference tokens.

obviously, this is back of the envelope math, but at 100t/s you would need like ~14k years. scale this to >100k GPUs and your in the hours to a couple days range.

philipportner··on How do you use Vim in the era of AI?
Hasn't changed at all since AI agents became a thing. tmux, nvim with a few plugins, mainly fzf and LSP support. If I do use an AI agent, I just run it in another tmux window.
philipportner··on Grit: Rewriting Git in Rust with agents
> I'm not sure you can prompt a full, accurate, copy of a nontrivial codebase out of them. Even with zero temperature their accuracy is just not that high.

Granted, these are some of the most widely spread texts, and not codebases, but just fyi: https://arxiv.org/pdf/2601.02671

> For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer’s Stone and 1984 (Section 4).

philipportner··on I was recently diagnosed with anti-NMDA receptor encephalitis
> My favorite side effect is that I now love all foods. Prior to this, I was a rather picky eater. Now I love everything!

I feel like there's a burntsushi joke hiding in there somewhere.

All the best Andrew.

philipportner··on AI Agent Guidelines for CS336 at Stanford
They reference the gist of 1cg in the honor code section of CS336.

https://cs336.stanford.edu/

philipportner··on Ladybird adopts Rust, with help from AI
FYI: Claude has output styles, one of them is called `learning`. Instead of writing the code itself, it will add `TODO(human)` and comments to explain how to. Also adds `Insights` explaining concepts to you in its output.

This link also has a comparison to Skills further down.

https://code.claude.com/docs/en/output-styles#built-in-outpu...

philipportner··on Consistency diffusion language models: Up to 14x faster, no quality loss
Did you publish anything you could link wrt. query rewriting?
philipportner··on We tasked Opus 4.6 using agent teams to build a C Compiler
Granted, these are some of the most widely spread texts, but just fyi:

https://arxiv.org/pdf/2601.02671

> For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer’s Stone and 1984 (Section 4).

philipportner··on We tasked Opus 4.6 using agent teams to build a C Compiler
This seems related, it may not be a codebase but they are able to extract "near" verbatim books out of Claude Sonnet.

https://arxiv.org/pdf/2601.02671

> For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer’s Stone and 1984 (Section 4).

philipportner··on Advent of Compiler Optimisations 2025
There's a link to the AoCO2025 tag for his blog posts in the op.