(For instance, 58 pages, 80 pages ... I don't think of these as normal academic length if this is the main body content ...)
1,509 karma · joined July 20, 2020
https://github.com/femtomc
(For instance, 58 pages, 80 pages ... I don't think of these as normal academic length if this is the main body content ...)
I mean, I feel like I'm going crazy -- but I was struck by this jarring blending of experiences ... shitting out some growth plots followed by autopilot on your wedding. Nice OpenAI. The only thing missing is a moment of self-reflection where I contemplate where exactly I lost what makes me ... me.
Is this what SV wants the world to look like? Mixing fucking cake batter while a bot shows me a regression to the mean website? Pretending like I have any sort of intentionality in my life, while a nameless entity (given quirky form) sort of walks me through my life?
I'm not sure why it gave me this impression, but strikes me as vaguely reminiscent of soma (from Brave New World).
Kind of sad, because the tech is actually incredible: who are they hiring to storyboard these commercials?
If I'm wrong here, I'd love reference links. I think of these companies as trying to inspire the idea that Claude (or GPT) are these special alien entities, in a sense?
Here's my issue with this post:
> True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use.
Okay, billions of dollars have been poured into these agentic LMs, right? Each training run to get the next increment is costing millions of dollars?
This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.
Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
It's sort of a litmus test for how little someone understands how learning in humans works or AI itself even. "Agentic" AI acts as a mirror: the more insight have yourself, the more powerfully you can pull the distribution ... the more you get out of it. (there's a bloodborne-inspired parable in here somewhere ...)
I think it's why the software from AI companies has "garbage" qualities. They are assuming that AI is going to fill in skills ... when it is a reflective amplifier.
But back to the discussion of "things we don't need to learn anymore": do we go to the gym, and say ... oh I don't need to do pull ups anymore, because there's an assisted pull up machine and I can just set it to my weight ...
Utterly idiotic.
- Use the code that your agents write in anger.
There you go. Do I know when my agents fuck up? Yes, I absolutely do -- because I'm a user of the code I have my agents write, and I ask things like "why is it taking 50 ms to start this program ..." and then I go in and find stupidity, and excise it. I do this over and over again.
Is it faster than writing it out by hand? Maybe! It's definitely a different perspective.
Start behaving like a baby "why, why, why" and then do a bit of reading, and you'll be fine.
A lot of these blog posts seem like they're aimed at software written by B2B companies who don't even use their own software ...
But taking Claude Code as the product of this style of thinking -- who is Claude Code for? Is it for everyone in the world? Well, if you look at the feature velocity, it seems like the answer is intended to be yes ... Claude Code is trying to solve every problem in software development in the world, all at the same time.
So I question the "user model" here.
Here's another thing that is true about Claude Code: it's among the most inconsistent and buggy pieces of software I've ever encountered.
- You can move the cursor with the mouse in the composer, but not in AskUserQuestion?
- When agents spawn subagents, the model name is inherited from the main agent, and seemingly none of the (4! yes, 4!) subagent tools seem to get this right (except for Explore, which seems to be fixed to a weaker model)
- Sometimes, when my usage limit halts, my agents will pick up when it refreshes (within ~2 hours or something) ... other times, nope -- even within the usage limit?
This is a sampling of my own experiences using this thing frequently. Are these sorts of details not important? Maybe not: I'm not at the level of this team, and may never be.
But I think it's a reflection of agentic engineering ... a somewhat embarrassing one, from my perspective. It paints a picture of a team who can't quite get the details right, even with the assistance of purported extremely powerful AI tools, even internal ones which we don't have access to?
I think when people look back on 2025 -- Boris is going to have his name right there in the books ... Claude Code, coding agents -- Anthropic (& Boris + team) made the first move.
But now it's 2026, and people know how harnesses work, and heavy lies the crown.
The answer is that these technologies are not universal in trade off space.
Of course, if you can have an LLM prove something by dispatch to an SMT solver ... you want to do that (and, indeed, LLMs do sometimes do this even when working on Lean proofs)
But the space of statements that you can prove using verification with SMT is smaller than dependently typed systems.
Also, Ada/SPARK are very good and robust -- they've been around for ~50 years ... but there's been a significant amount of work in formal verification since then, right? I would be more clear that Ada/SPARK are state-of-the-art for "mission critical software" (and have been for many years), but there's been a lot of research work in formal verification of low level programming since then (Low*, for instance)
Your post is setting up a strawman between automatic formal verification and formal verification using interactive theorem provers ... obviously there is a spectrum, and Ada/SPARK are navigating the space to try and automate much of the work required to automatically dispatch with obligations to prove (computable) properties about programs.
Bend2 is a QTT -- it's dependently typed, and comes from the lineage of systems which are focused on being expressive enough to formalize mathematics.
Of course you need to build a somewhat significant "standard library" of theorems, tactics (as metaprograms), etc ... to approach what is built into the compiler in Ada. These are different approaches with different trade offs.
Your post isn't clear, you don't go into any of these details ... why did you post this? Do you think this is clear writing?
- this Bend is not really related to the old Bend (only in name)
- this Bend doesn't really have anything to do with interaction combinators
- this Bend is a QTT, with a change to affinity which enforces a good performance property for GPUs
- the "higher order at comptime" is neat, reminds me of Andras Kovacs' work on 2ltt and staging in dependently typed languages.
- this Bend is likely to be good at "balanced recursive computations on ADT", and can parallelize them ... but won't be as good as CUDA or e.g. Futhark on dense rectangular array computations
- performance needs improvement in the scheduler, to possibly help with balanced work (looking at the n queens and symbolic regression numbers)?
How are you going to handle search or synthesis over irregular structures (SupaGen)?
Not many professions are formally trained to be cognizant of this lack of understanding, and how to confront it.
Usage of AI in collaborative settings is an amplifier of these issues, especially if someone doesn’t realize they don’t understand: they couldn’t teach or explain the concepts they use, or be forced to work with them malleably in a way that an expert or researcher would.
If you are cognizant of your lack of understanding, you can remedy it by slowing down and teaching yourself. This is required to make better use of AI in the domain of interest!
But you can’t have all things at once: you can’t move at speed with AI, collaborate effectively, and understand what is going on as an expert would. It is not physically possible for a human brain.
Also shout out @tonyg for Syndicate! I keep coming back to it -- a beautiful piece of PL work.
I’m working on hard things, it is very noticeable when it is hums through something and then falls over on something it should not
I can tell by analyzing my own prompts to look at when I get frustrated ;)
No, it is a shocking level of incompetence given the conveyed seriousness of the work by these labs.
So yes, models are getting better. Ask yourself: if you know that to be true, would you act the same way that the teams did in the public post mortems?
I'm only saying that that I don't think one can automate the whole thing. After all, agents don't actually possess ... agency -- they require us to breath live into them. So at some part of the stack, there has to be a human, with an opinion, thinking about what they want.
Also: I don't aspire to be a worker, I aspire to seize the means of production.
Agent swarms, self learning, Ralph loops, execution DAGs, spending hours trying to convey my preferences into skills, yada yada — the reality is that no matter how long I spend trying to close the amortization gap between my expertise and (insert autonomous agent loop) … I always come back to see shit that I have opinions about and need to polish.
If you’re keeping on top of your craft, I just don’t think general purpose pretraining / RL is going to give you something which perfectly copies you. It can get closer and closer! But I’ve never seen agent work where I didn’t immediately start thinking of refinements …
I guess many things are okay if you’re okay with averages
Reverse engineering can take several forms for me: it started with writing parts of systems from scratch, and now (due to academic training) it oftens looks like "thinking about how to explain the system to a group of colleagues", which is one degree of separation from writing the system from scratch, but is much faster - and seems to work well when working with LM agents on code (because you have to confront your ignorance)
part of the training in academia is how to know when you don't understood the thing, so this ends up working well for me now.
One trick: you can force yourself to understand something better by making a presentation and imagining you're going to be grilled by someone smarter than you. This is kind of the feeling of pressure that academic paper reviews give (in an ideal world) and it translates when you think about explaining something to someone else.
I started practicing this for the stuff I'm interested in ... and the result is that I generate a lot of ideas. 99% of them are garbage.
Before AI, a conversation with colleagues would help me weed out the garbage, and possibly give me a new perspective ("here's something surprising about what you just said!")
Now, AI plays a helpful role in helping me weed out the garbage for the subset of ideas that I can test empirically -- and sometimes it does give me a new perspective, but with significantly more noise than my colleagues, who would take me right to the point that would be useful for me to see.
I've never felt like I needed a second brain ... because if I had such a thing, I think I would be keeping the 99% garbage around -- instead I just let biology GC it.
RIP to those poor humans. I can't imagine having your brain melted by Fablespeak as your FTJ.
Of course, all I’m saying is that you need to refine your sample! For instance: the allocation architecture is not correct, and one has to run a bunch of performance investigations and resolve it.
My responses are intending to convey that I don’t believe this is possible, no matter how good LMs get — and it seems like we are in agreement.
I don't believe the problem is "ah, you didn't write the spec clearly enough" -- which is why I'm asking about your own experiences.
> If you instead spend a day writing a proper specification, then ask the agent to spend a week implementing that, you'll need zero tools and skills afterwards to clean it up, because there won't be any misunderstandings, assumptions or other things, just code fulfilling what the specification says.
Are you claiming that you reliably get good software out of agents using this process?
My experiential claim is that I've followed this process for complex pieces of software (and, often, taking significantly longer than a day -- more like several weeks, as if I was writing an academic paper) and they still get it wrong ... even with an "academic paper level spec"
Models: only the best, Fable 5 max, Sol 5.6 max, yada yada -- I'm not convinced this process works for "serious endeavors" (for instance: let's build a novel accelerator compiler or let's build a new incremental linker) where "serious" here means "something where I'll likely be spending multiple weeks working on it"
What is the most complicated thing you’ve built with LM agents? Have you done it with a single spec? How novel was it?
This comment is so laughably “you’re holding it wrong” I can’t respond to you seriously.
> If you instead spend a day writing a proper specification, then ask the agent to spend a week implementing that, you'll need zero tools and skills afterwards to clean it up, because there won't be any misunderstandings, assumptions or other things
The set of software that has followed this process is measure zero.