HNHacker News
TopNewBestAskShowJobs

celrod

1,143 karma · joined February 4, 2018

SIMD and performance enthusiast. https://github.com/JuliaSIMD https://spmd.org/
submissionscomments
celrod··on MiMo v2.6
Yeah, that's what I'd been leaning towards. No mcp, but I'll see if I can reproduce and debug it, since other people don't seem to have that problem as badly as I've experienced it (and the idea of having a nasty bug like that bothers me).

No mcp support. I'll try copying deepseek harness's basic tool call formats as a starting point.

celrod··on MiMo v2.6
I tried it a few times and liked the speed, but often found it ended up looping, i.e. repeating the same token sequence (e.g. the same sequence of 5 paragraphs) over and over again until it hit the max output limit. This doesn't end up happening every session, but does every now and then.

My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results

I was using a vibe coded bare bones harness. I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.

celrod··on So you want to use OpenRouter?
I was just trying deepseek v4.1 flash on OpenRouter. After running reasonably well for a while, I return to the window and see that my scrollback is nothing but this repeated over and over again:

``` OK.

Let me write.

Let me go.

OK.

Let me write the script.

Let me go. ```

I'm not sure to what extant this is a model problem, vs some providers being fairly broken. If I chose a single provider, I could know how to blame and to avoid them. With OpenRouter, I don't know which provider I was on when this happened.

celrod··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
If Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.
celrod··on GPT-6 Astra
Mercor, Tacit Labs, Handshake AI... I suspect companies like these play a big part in model improvements, generating high quality benchmark/task-focused data for training.

However, these do require educated, white collar, workers.

celrod··on Claude Fable 5.1 and Claude Mythos 5.1
> A narrow set of frontier LLM development tasks, such as distributed training infrastructure, ML accelerator design, and kernel development for certain non-standard chips.

https://support.claude.com/en/articles/15363606

My work is mostly on the Nvidia b200, which apparently gets flagged as non-standard.

Opus 5 works, but sometimes I do wonder if it's surreptitiously trying to sabotage the efforts -- possibly deliberately, but more likely by something like Fable's initial launch, which did come with secretly degraded performance when detecting kernel work. Anthropic was open at the time that such a mechanism existed, but disabled it due to backlash. More likely than not, this is just paranoia on my end...

celrod··on Claude Fable 5.1 and Claude Mythos 5.1
I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.
celrod··on Qwen3.8 27B scores 52 on Artificial Analysis
I think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them.
celrod··on Delta flight hit by firework while landing at Midway Airport on Fourth of July
A few years ago, on July 4th one of my mom's dogs freaked out, somehow managed to escape, and got hit by a car before my mom found her. She loved that dog, regularly attending nose work competitions with it. One of your pets getting a seizure must be harrowing for both you and the dog.

I don't light fireworks.

celrod··on Memorizing session transcripts isn't useful
When claude goes down a wrong path, I tend to clear context and write a new prompt that helps guide it down the correct path. Whatever thinking or context that led it there has inertia and tends to be sticky, otherwise.

Pretty annoying when it brings those up again later from memory...

celrod··on The Token Compression Illusion: Why I'm Skeptical of RTK
I found this, which has some: https://arxiv.org/pdf/2605.28876 TLDR: RTK does not look good according to the author's benchmark.
celrod··on GLM-5.2 is the new leading open weights model on Artificial Analysis
That list also places Sonnet 4.6 above Opus 4.6, which doesn't match my experience.
celrod··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Exactly. How can "we" develop and encourage benchmarks for multi-turn user assistance? That is what I want. I feel like the models and harnesses push much too hard against this workflow -- that they push you towards letting go and vibe coding, with only your discipline (and desire for a quality and maintainable product) holding it back.
celrod··on Running local models is good now
What quant do you run it at? 32GB seems like cutting it close on the rtx 5090 if going 8b, but other commenters are saying 4b lobotomizes the model.
celrod··on Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
If performance is the concern, ugrep will get you most of the way there relative to gnu grep, and should be fully grep compatible in terms of syntax:

https://github.com/Genivia/ugrep#aliases

Claude Code may ship with ugrep already.

celrod··on Every Byte Matters
Yes. I think one of the big advantages of SoA is that you only pay for the fields you're currently using. If you need a field somewhere, you can add it and only pay the cost of iterating it where you need it.
celrod··on Linux Terminal Memory Usage
foot also offers a client/server architecture. If you start a foot server (e.g. with a systemd service), you can use `footclient -N`. This may reduce the memory pressure of running many terminals.

This is similar to the `kitty --singleinstance` mentioned in another comment by amarshall.

celrod··on Making Julia as Fast as C++ (2019)
Yeah, for now. I'd like it to be open, but I also want to potentially be able to make money/a living off of it. My dream would be that it can be open while hardware vendors pay me to optimize for their hardware. For how, being closed gives me more options. It's a lot easier to open in the future than to close, so it's just keeping options open.

I've thought a lot more about the engineering than any sort of marketing or businesses plan, so I just want to defer those.

celrod··on Making Julia as Fast as C++ (2019)
I'm still working on it. I'm currently working on a cache tile-size optimization algorithm that should (a) handle trees (a set of loops can be merged at some cache levels and split at others, e.g. in an MLP it may carry an output through the L3 cache, while doing sub-operations in the L2/L1/registers) (b) converge reasonably quickly so compile times are acceptable.

This is the last step before I move to code generation and then generating a ton of test cases/debugging.

My goal is some form of release by the end of the year.

celrod··on Making Julia as Fast as C++ (2019)
I was once a bit of a Julia performance expert, but moved toward c++ for hobby projects even while still using Julia professionally.

I wrote a blog post at the time with exactly that punchline (not explicitly stated, but just look at the code!): https://spmd.org/posts/multithreadedallocations/ The example was similar to a real production-critical hot path from work.

Maybe things changed since I left Julia, but that was December 2023, for years after this blog post.

celrod··on Granite 4.1: IBM's 8B Model Matching 32B MoE
Fellow kakoune user here. I'm curious about your use case/ what you're doing with it!
celrod··on C++26 is done: ISO C++ standards meeting Trip Report
In my experience, llms don't reason well about expected states, contracts, invariants, etc. Partly because that don't have long term memory and are often forced to reason about code in isolation. Maybe this means all invariants should go into AGENTS.md/CLAUDE.md files, or into doc strings so a new human reader will quickly understand assumptions.

Regardless, I think a habit of putting contracts to make pre- and post-conditions clear could help an AI reason about code.

Maybe instead of suggesting a patch to cover up a symptom, an AI may reason that a post-condition somewhere was violated, and will dig towards the root cause.

This applies just as well to asserts, too. Contracts/asserts actually need to be added to tell a reader something.

celrod··on SIMD programming in pure Rust
I netted huge performance wins out of AVX512 on my Skylake-X chips all the time. I'm excited about less downclocking and smarter throttling algorithms, but AVX512 was great even without them -- mostly just hampered by poor hardware availability, poor adoption in software, and some FUD.
celrod··on Framework Laptop 13 gets ARM processor with 12 cores via upgrade kit
They're 8x A720 + 4x M520, not Snapdragon X.
celrod··on Ghostty is now non-profit
I use niri and footclient -N, so builtin window and tab completion don't appeal to be.

Foot feels fast, but I've not actually measured the latency. It also seems to use less CPU than GPU accelerated terminals (which it isn't) from just glancing at btop. So I'm not sold on GPU-acceleration as a feature unless I see benchmarks demonstrating the value in improved latency and reduced CPU use compared to foot

I love that foot's scrollback search, selection expansive, and copy can be entirely keyboard driven. Huge QoL feature for me that often seems neglected to me in other terminals.

celrod··on Notes on switching to Helix from Vim
I use kakoune, and don't understand why helix seems to be taking off while kakoune (which predated and inspired helix) remains niche.

Kakoune fully embraces the unix philosophy, even going so far as relying on OS (or terminal-multiplexer, e.g. kitty or tmux) for window management (via client/sever, so each kakoune instance can still share state like open buffers).

A comparison going into the differences (and embracing of the unix philosophy by kakoune) by someone who uses both kakoune and helix: https://phaazon.net/blog/more-hindsight-vim-helix-kakoune

Sensible defaults and easy setup are a big deal. No one wants to fiddle with setting up their lsp and tree-sitter. There's probably more to their differences in popularity than just this, though.

celrod··on New nanotherapy clears amyloid-β, reversing symptoms of Alzheimer's in mice
I think they're arguing

Cause -> cognitive disease Cause -> plaques

That is, that the same cause is behind both.

There may be some arrows from plaque to disease as well (i.e., that plaques also increase disease).

I dont know the truth, but just trying to understand/follow Alzheimers news and reading comments.

celrod··on Helix Editor 25.07
Is including batteries the main reason helix seems to have started taking off, while kakoune hasn't?

I use kakoune, because I like the client/server architecture for managing multiple windows, which helix can't do. The less configuring I do the better, but I've hardly done any in the past year. It's nice to have the option.

I do use kakoune-lsp and kak-tree-sitter.

celrod··on Why does C++ think my class is copy-constructible when it can't be?
I agree.

The problem with eager diagnostics and templates is that the program could define a `Base<int>` specialization that has a working copy constructor later. [0]

I think if you define an explicit instantiation definition, it should type check at that point and error. [1] I find myself sometimes defining explicit instantiations to make clangd useful (can also help avoid repeated instantiations if you use explicit declarations in other TUs).

[0] https://en.cppreference.com/w/cpp/language/template_speciali...

[1] https://en.cppreference.com/w/cpp/language/class_template.ht...

celrod··on Suckless.org: software that sucks less
I use Wshadow personally. I highly recommend it. I think code that violates it (even if correct) is harder to understand.
Page 1 of 13Next →