686 karma · joined August 11, 2022
At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.
How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?
I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.
This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol
Well, right now, I'm pretty happy and have a personal project that I'm very eager to have done and polished and see the result of. Hopefully life keeps throwing more of them at me. It's been a constant issue for me though.
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
Which doesn't seem like it. The ideal is to keep removing couplings. But in practice, I find that focusing on the statefulness, the distribution-specificness, these "accidentals", is how you actually get leverage.
The obvious social consequence is now the tendency towards centralization, tyranny, "you people can't think for yourselves, not because you have some kind of 'inherent intelligence defect', but because the world is so complicated and interconnected that only coherent systems thinking (conveniently administered from one central point) is what can save you". In the Enlightenment world, the ideal was that every man could in principle, think for themselves.
A good ending might be that everybody benefits each other by distributing intelligence, but in practice centralization always seems to win.
AIT tried solving it? But AFAIK it's a lot of pretty results with not much real application.
A better approximation is something of a "shared model"; then you can actually state things like, the transfer of information sometimes is "trivial" because, well, it's right there in your compressor/decompressor.
Not sure what you're trying to say. Any time someone wants to actually make a system and not a pile of spaghetti or inactionable philosophy (for all that I love philosophy), they reach for mathematics in some way, shape, or form.
The implication is that humans are unreliable and shouldn't be trusted.
Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.
OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.
Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.
I've been still just like, making VM's with proxmox, then putting my agent in the machine and letting it run free (with my dotfiles setup script making dev env pretty much free, though I could also just make a VM snapshot). What's wrong with that? Is that not the scalable solution for enterprise rn?
No matter how much we pretend, that's how a lot of abstractions work. Things that touch the real world can change; there's a risk that the change could be as something as simple as a bugfix to changing the underlying implementation but preserving a higher level goal; you generally want a human in the loop to make sure the semantics work out and everybody's agreeing.
Of course, I think putting numbers to wishy washy meta-quantities like "how efficiently does a certain conceptual scheme help you" are super loaded and hard to properly talk about (incommensurability). I've been toying with trying to make a repository of all the possible "moves" one can make in this kind of abstract analysis - constrain the problem statement, argue something like "the system is what it does", dissolving, etc. but even that seems hard
Maybe there's just a better way to think about it and I'm still thinking about it way too much like a programmer
If the algortihm doesn't work the same forward, backwards, and with a tree scan, it ain't reduce (as a first approximation not IFF)
It's kind of hard to just "get into" these though, as they're sufficiently foreign that I'm spending more learning about the "accidentals" rather than the core dev loop that say, an employer would care about.
So on top of that, for a more "traditional" software project, I'm working on making a game with a custom networking protocol, but with a focus on the networking protocol itself. I do think a decent amount of games can be subsumed under this paradigm but it requires careful design to a necessary and sufficient protocol. Of course, whether the game itself is fun is a different story, but I think it can be made fun.
Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and it's an optimal compressor. It's completely vacuous though. You can always hide the "heavy work" in your mappings and descriptions. It's ironic that a kolomogorov complexity minimizer is so loaded that it's vacuous.
This is pretty much why I roll my eyes at this point at all the compression is intelligence memes.
I wonder when intervention and causality will hit the mainstream. These tools were designed specifically to counteract purely predictive theories. But your average compression dude will hold tight to their paradigms and slogans, not realize their internal contradictions (that their own field has brought up), and then whenever a new paradigm suddenly becomes visible and mainstream, they'll latch onto that. It's not principled at all.
And to be clear - I do think intelligence is some amount of compression, and I am well aware of formal results such as the arithmetic decoding theoretical and empricial result. Just annoyed. It's literally no different than the whole Bayesianism meme. If you're not actually practicing that type of intelligence as a basis, then you don't get to go around beating the drum about how it's the ultimate reality. You're just spouting dogma to feel like part of an in-group.
Distinctions, you generally "only pay for" in computational cost, by needing to search twice over an axis you may not need to split.
Similarities, if you wrongly assume two things are similar, means you're just wrong.
Of course, we know from computer science that doing more computation isn't free either.
I find myself often being more and more pedantic the more I want correctness - but of course this comes with the tradeoff of losing the high level abstract picture.
Saying what you're not going to do is also good design hygiene.
I will say that I'm annoyed by this behavior too. It feels like the models are writing their state of mind directly to output that should be clean. Often times, I will push back, and then it will... do the correction, and write the push back into the damn output. "Claude, I want burgers, not fries". The button text now changes to "Fries (NOT BURGERS)". Like, what?
Distinctions are powerful local reasoning tools, but a component of "real" reasoning is synthesis. Which they clearly can do sometimes - but not every time and not even remotely a probable amount of times.
But then there are aspects I'm missing because while I've thought about the network protocol deeply, I'm not, say, a game developer who's ever gone through the whole game dev lifecycle. There's common patterns with software dev but it ain't it. There are probably so many things I have not thought about w.r.t. the whole deployment process that I'm not sure if letting AI vibe design+code it out is good or if I need to sit down and deeply work out the things I don't even know I don't know.
It's always a set of tradeoffs between things.