HNHacker News
TopNewBestAskShowJobs

sigbottle

686 karma · joined August 11, 2022

submissionscomments
sigbottle··on Clef: Open-source decision models, and new RL fine-tuning platform
What even are these new "decision models?" Take an existing LLM, feed it a prompt, force it to pick a choice; decode is 1 token (or rather, the whole logit set for only that last token; token implies selecting one logit) so you made a choice. That's it?
sigbottle··on How to speed up the Rust compiler in September 2026
Is there any experimentation with new ABIs? I mean certain linker flags like --f-lto literally hijack the linker protocol to dump an AST into the backend.

At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.

sigbottle··on Commit Description as a Thinking Tool
the mechanism is self documenting; context is not unless you pollute all your files with an ADR's worth of alternatives.
sigbottle··on The AI Race Just Got Awkward
This is insanely cool, what the hell.

How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?

I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.

This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol

sigbottle··on Livenerf: Has Opus 5.5 been nerfed yet?
But as we also see in this thread, all evidence is dismissed and there's no good faith discussion. All data from the other side is lies and contamination. In that environment, it's power who decides who wins.
sigbottle··on A Staff Engineer's Guide to Inventing Work
I wonder how this applies to personal projects. I find it hard to motivate myself to just "study a textbook". I do do so, but it almost always feels useless compared to actually doing something. It doesn't have to be grand, but it needs to be something I can at least "trick" myself into believing it's useful.

Well, right now, I'm pretty happy and have a personal project that I'm very eager to have done and polished and see the result of. Hopefully life keeps throwing more of them at me. It's been a constant issue for me though.

sigbottle··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
> It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

Yes, still running into this, but surprised about this

> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.

But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.

For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.

sigbottle··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.

I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".

I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).

sigbottle··on Claude partial outage
I think the reason that intelligence is so scary is how contextual, stateful, path-dependent, and coupled it is.

Which doesn't seem like it. The ideal is to keep removing couplings. But in practice, I find that focusing on the statefulness, the distribution-specificness, these "accidentals", is how you actually get leverage.

The obvious social consequence is now the tendency towards centralization, tyranny, "you people can't think for yourselves, not because you have some kind of 'inherent intelligence defect', but because the world is so complicated and interconnected that only coherent systems thinking (conveniently administered from one central point) is what can save you". In the Enlightenment world, the ideal was that every man could in principle, think for themselves.

A good ending might be that everybody benefits each other by distributing intelligence, but in practice centralization always seems to win.

sigbottle··on I don't want to read what you didn't write
I never liked information theory because information theory as Shannon envisioned it fundamentally did not deal with semantics.

AIT tried solving it? But AFAIK it's a lot of pretty results with not much real application.

A better approximation is something of a "shared model"; then you can actually state things like, the transfer of information sometimes is "trivial" because, well, it's right there in your compressor/decompressor.

sigbottle··on Advisory Group on Mathematics and Artificial Intelligence
Alrighty man, if your worldview is, "Everything within my knowledge is the objectively correct amount of knowledge to learn, and everything outside of my knowledge is useless abstract garbage", you can live like that.
sigbottle··on Advisory Group on Mathematics and Artificial Intelligence
20th century physics is heavily based off of abstract algebra (symmetry) and pretty much all of CS is based off of discrete mathematics.

Not sure what you're trying to say. Any time someone wants to actually make a system and not a pile of spaghetti or inactionable philosophy (for all that I love philosophy), they reach for mathematics in some way, shape, or form.

sigbottle··on Fable 5 – Median thinking declined in August
> I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

The implication is that humans are unreliable and shouldn't be trusted.

Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.

OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.

Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.

sigbottle··on AX – Google’s Open Agentic Orchestrator
Could someone explain to me what the general workflow is now that people are converging to? I haven't really been catching up with the AI ecosystem but I was looking into agent sandboxes and VM's recently and there's a ton of these startups and tools now. Is giving the agent a temporary scratchbox really that valuable?

I've been still just like, making VM's with proxmox, then putting my agent in the machine and letting it run free (with my dotfiles setup script making dev env pretty much free, though I could also just make a VM snapshot). What's wrong with that? Is that not the scalable solution for enterprise rn?

sigbottle··on I built non-autoregressive decision models with RL a year ago
Furthermore, it's not about the current innovation right now - if you sell yourself on a broader mission, your core product can evolve and change with it, and you're more selling yourself as the guy who will make that abstract vision possible no matter what.

No matter how much we pretend, that's how a lot of abstractions work. Things that touch the real world can change; there's a risk that the change could be as something as simple as a bugfix to changing the underlying implementation but preserving a higher level goal; you generally want a human in the loop to make sure the semantics work out and everybody's agreeing.

sigbottle··on If math is more than proof, we need to better celebrate the rest of it
Well, it's knowing when to push and when to not. You probably have an intuition for, I don't know, abstract algebra objects (I don't know your field of specialty :P), without needing to symbolically manipulate all of it, but you developed a deep intuition for them through many proofs and attempts at proofs with them.
sigbottle··on C++26: Trivial infinite loops are no longer undefined behaviour
__asm__ __volatile("hlt"); when doing quick and hacky debugging could work
sigbottle··on OpenJev
I really want to create a nosology of common generic memes that can be applied to literally anything without context. Saying that the evaluators are just stupid and arbitrarily chasing the fashion of the week instead of the evaluators possibly actually latching onto some structure is a tale as old as time.
sigbottle··on Bend 2 and the Vibe-Coding Trap
I really wish there were search harnesses, actually. My LLMs are lazy as hell and seem to want to just report the first thing they find on google. I know they can return truly niche and useful results, but it takes a lot more prompting to get them there than I would like.
sigbottle··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
aww man. I remember following victor in college. I mean pivots gotta pivot, and this is probably a better one for business, but always thought the interaction combinator framework was cool
sigbottle··on The Relation Between Mathematics and Physics by Paul Dirac (1939)
Well there are counting arguments to say that a higher level intelligence that somehow achieves say, 100000× brain efficiency of humans, still can't do that much more work than humans, if humans found the "best abstraction". Think about computability for example - that means you could feasibly solve problem instances of n+20 relative to what a human can solve.

Of course, I think putting numbers to wishy washy meta-quantities like "how efficiently does a certain conceptual scheme help you" are super loaded and hard to properly talk about (incommensurability). I've been toying with trying to make a repository of all the possible "moves" one can make in this kind of abstract analysis - constrain the problem statement, argue something like "the system is what it does", dissolving, etc. but even that seems hard

sigbottle··on Anecdotally, programmers dislike "reduce"
I do know what a monoid is, but a monad in the category of endofunctors is the scary word for me :sob:
sigbottle··on Anecdotally, Programmers Dislike "Reduce"
Sorry I changed problems a bit and started talking about me trying to understand matrices lol
sigbottle··on Anecdotally, Programmers Dislike "Reduce"
It always messes with me: reducing across a specific axis always takes O(whole tensor) time, because there's no difference between "iterate over all dims, then collapse the final one" versus "iterate versus the first dim and do some cursed tensor accum" (and likewise for between)

Maybe there's just a better way to think about it and I'm still thinking about it way too much like a programmer

sigbottle··on Anecdotally, programmers dislike "reduce"
Isn't reduce usually used for monoidal operations? Or do people implicitly absue ordering?

If the algortihm doesn't work the same forward, backwards, and with a tree scan, it ain't reduce (as a first approximation not IFF)

sigbottle··on Vectorized and performance-portable Quicksort (2022)
Honestly part of me feels that way about way too many things in retrospect about my own life and interests.
sigbottle··on Ask HN: What are you working on? (September 2026)
Upskilling in both linux kernel dev work and also LLM inference systems work.

It's kind of hard to just "get into" these though, as they're sufficiently foreign that I'm spending more learning about the "accidentals" rather than the core dev loop that say, an employer would care about.

So on top of that, for a more "traditional" software project, I'm working on making a game with a custom networking protocol, but with a focus on the networking protocol itself. I do think a decent amount of games can be subsumed under this paradigm but it requires careful design to a necessary and sufficient protocol. Of course, whether the game itself is fun is a different story, but I think it can be made fun.

sigbottle··on Why don't machine learning research agents overfit?
Compression in this modern day and age is so slop.

Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and it's an optimal compressor. It's completely vacuous though. You can always hide the "heavy work" in your mappings and descriptions. It's ironic that a kolomogorov complexity minimizer is so loaded that it's vacuous.

This is pretty much why I roll my eyes at this point at all the compression is intelligence memes.

I wonder when intervention and causality will hit the mainstream. These tools were designed specifically to counteract purely predictive theories. But your average compression dude will hold tight to their paradigms and slogans, not realize their internal contradictions (that their own field has brought up), and then whenever a new paradigm suddenly becomes visible and mainstream, they'll latch onto that. It's not principled at all.

And to be clear - I do think intelligence is some amount of compression, and I am well aware of formal results such as the arithmetic decoding theoretical and empricial result. Just annoyed. It's literally no different than the whole Bayesianism meme. If you're not actually practicing that type of intelligence as a basis, then you don't get to go around beating the drum about how it's the ultimate reality. You're just spouting dogma to feel like part of an in-group.

sigbottle··on Claude is a Contrarian
Again, I will posit the hypothesis that it's a learned behavior from training.

Distinctions, you generally "only pay for" in computational cost, by needing to search twice over an axis you may not need to split.

Similarities, if you wrongly assume two things are similar, means you're just wrong.

Of course, we know from computer science that doing more computation isn't free either.

I find myself often being more and more pedantic the more I want correctness - but of course this comes with the tradeoff of losing the high level abstract picture.

Saying what you're not going to do is also good design hygiene.

I will say that I'm annoyed by this behavior too. It feels like the models are writing their state of mind directly to output that should be clean. Often times, I will push back, and then it will... do the correction, and write the push back into the damn output. "Claude, I want burgers, not fries". The button text now changes to "Fries (NOT BURGERS)". Like, what?

Distinctions are powerful local reasoning tools, but a component of "real" reasoning is synthesis. Which they clearly can do sometimes - but not every time and not even remotely a probable amount of times.

sigbottle··on How to write an effective software design document
It's mixed for me because there are certain things that I clearly think are needed. For example I'm building a custom network architecture and it's to the point where I'm using frontier models to reverse engineer game clients (while battling against the cyber safety system) for the sole purpose of validating that the network architecture I'm making is, if not "useful" (cause there's the game itself), at least different and superior. And to me the designs out there clearly are evidence of things not designed well and thought through ahead of time and instead a patchwork of hacks.

But then there are aspects I'm missing because while I've thought about the network protocol deeply, I'm not, say, a game developer who's ever gone through the whole game dev lifecycle. There's common patterns with software dev but it ain't it. There are probably so many things I have not thought about w.r.t. the whole deployment process that I'm not sure if letting AI vibe design+code it out is good or if I need to sit down and deeply work out the things I don't even know I don't know.

It's always a set of tradeoffs between things.

Page 1 of 13Next →