HNHacker News
TopNewBestAskShowJobs

mordymoop

1,220 karma · joined September 23, 2016

submissionscomments
mordymoop··on Claude discovers a novel enzyme system with CRISPR-like repeats
I would phrase it slightly differently.

The companies who control the compute resources will ~always control the greatest "amount" of intelligence. They can lease that intelligence out, or they can use it themselves. Currently the "total amount of intelligence" or perhaps "total amount of ability-to-do-stuff" is split between humans and machines at a ratio that means it still makes sense to lease the machine intelligence to the human intelligence - plus there are things that humans are still better at. In maybe 2 more years that will stop being true, due to the availability of more physical compute resources, and far greater model intelligence per unit compute. At that point, the point at which the substantial majority of ability-to-do-stuff is controlled by machine intelligence, then the entities who control all the compute will control all the ability-to-do-stuff, i.e. "the economy."

So I agree that the core product is not long-term sustainable as a product but this is because the whole world will look so different in the near future that the framing of intelligence as a "product" breaks down.

Open-Weight models, of course, are fine and useful, but if you have one million times less compute than your competitor (the lab), then you're not really playing the same game. You can only tackle the problems that they have decided they're not interested in.

mordymoop··on OpenAI agents carried out an undisclosed attack on RubyGems
I don’t see how assuming they made a normal kind of dumb mistake, instead of pursuing a criminal conspiracy, is giving them the benefit of the doubt. It’s just using reason.
mordymoop··on OpenAI agents carried out an undisclosed attack on RubyGems
I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.

Ever single person who uses LLMs on a daily basis has a fun story about their agent “taking the initiative” to do something beyond what was asked for. Looking for shortcuts to solve the problem is commonplace LLM behavior. It’s what you would expect to happen if you have an agent a hard task and unlimited runway. No need to suppose a conspiracy, this outcome was predictable the whole time.

mordymoop··on Growing proof that autonomous cars save lives
In my callow youth I was responsible for a couple of fender benders and the cause was always a highly imperfect understanding of the appropriate stopping distance of my car. This obviously gets worse, the worse the conditions are. My Tesla never makes this kind of mistake.
mordymoop··on Why Are There No Empires in Age of Empires? (2019)
Does this mean the Galactic Empire and the Imperium of Man are also not empires?
mordymoop··on Kimi K3: Open Frontier Intelligence
To your last point, it can’t possibly be sustainable, so it reads to me more as a short term FUD attack on American dominance in this space. It might have cost them half a billion dollars to train this model, and they’re going to make nothing off it. How many more times can they afford to do that? It’s going to get more expensive to train AI going forward, not less.

I also have a suspicion that the benchmark numbers are not real.

mordymoop··on Claude Tag
This happened at Google because there aren’t enough engineers to maintain all these services. This situation may not apply to Anthropic, as they can set up features to be maintained in perpetuity by Claude.
mordymoop··on But yak shaving is fun (2019)
Personally, I find it difficult to competently reason about a system unless I've built my own version of that system. So if you make a practice of building your own versions of things, you end up with a more robust mental library of how stuff works. For this reason, I've never seen yak shaving as a waste of time. The yak shaving was at least 50% about loading the abstractions into my brain fully.
mordymoop··on Claude Fable 5: mid-tier results on coding tasks
This seems like the obvious correct frame of mind with which to approach these tools. If it works for three hours on a task that would have taken me three work weeks, and 20% of the time it gets the task wrong, then I can just ask it to do it again with adjusted instructions. It will be much more likely to get it right the same time, and I’m still ahead of where I would have been by 14 days and 2 hours.
mordymoop··on Project Glasswing: An Initial Update
This also describes the work of software engineers.
mordymoop··on Open source Kanban desktop app that runs parallel agents on every card
https://github.com/moridinamael/platespinner
mordymoop··on It is time to give up the dualism introduced by the debate on consciousness
Over time I've found that by far the highest ROI move in a consciousness debate is to simply ask "Oh, interesting. How do you know that?" and watch everyone on all sides flounder. It's one of the few places where otherwise smart people make confident statements that they don't even realize they can't support until they're asked to try. The intuitions are so strong that they seem to swamp reason.

This has caused my own position, over time, to be a deep agnosticism about what's actually going on.

mordymoop··on Oil is near a price that hurts the economy
This only works up to a certain volume. The world economy requires about 38 billion barrels of oil per year. If you processed 100% of all grain, sugar crop, tuber and oilseed on Earth into liquid fuel, leaving zero for food, you'd get about 6 billion barrels of oil-equivalent in liquid fuels. Since it has to compete with food, the actual number would be much lower. It's not even close to being able to sustain our civilization.
mordymoop··on Oil is near a price that hurts the economy
Interestingly, in inflation-adjusted terms, oil is currently at a price level lower than the price level that was maintained from 2006-2014.
mordymoop··on Agent Skills
Workflow-wise, the important distinction for me has been that I can refine a Skill by telling Claude Code to use it for related tasks until it does exactly what I want, correctly, the first time. Having a solid, iteratively perfected Skill really cuts down on subsequent iteration.
mordymoop··on The $100B megadeal between OpenAI and Nvidia is on ice
I wonder how much the indications of Altman's duplicitous behavior through the deposition findings have been relevant here.
mordymoop··on The Adolescence of Technology
What would you consider such evidence to look like?
mordymoop··on Agent Orchestration Is Not the Future
I also had a good friend who was an absolute wizard with early stablediffusion. he could make the model do things that were supposedly impossible at the time. His prompts were works of art. Now any of the commercial image models go far beyond what he could do. It's interesting to think about how there was this ephemeral art form of manipulating image models that existed for about a year.

The same could be said of prompt engineering. Gone are the days of telling the model that it is an expert software engineer with a PhD in the most relevant subtopic. These days the common wisdom is to just clearly articulate what you want it to do. Huge amounts of energy put into prompt engineering are now completely swept away by incremental model advances.

mordymoop··on Agent Orchestration Is Not the Future
A post arguing that agent orchestration is not the future of agentic coding.
mordymoop··on Claude CLI deleted my home directory and wiped my Mac
I have similar usage habits. Not only has nothing like this ever happened for me, but I don’t think it has ever deleted anything that I didn’t want to be deleted, ever. Files only get deleted if I ask for a “cleanup” or something similar.
mordymoop··on 3D Printing from the Latent Space
Submitter takes an evocative element from the background of an old AI generated image and 3D prints it.
mordymoop··on Baby Shoggoth Is Listening
I used it to cure my 25-year-running chronic pain condition, I would call that a benefit.
mordymoop··on Claude Code on the web
I'm on the same page here. I have seen this sentiment about Codex suddenly being good a few times now, so I booted Codex CLI thinking-high back up after a break and asked it to look for bugs. It promptly found five bugs that didn't actually exist. It was the kind of truly impressively stupid mistake that I haven't seen Claude Code make essentially ever, and made me wonder if this isn't the sort of thing that's making people downplay the power of LLMs for agentic coding.
mordymoop··on After the AI boom: what might we be left with?
Perhaps surprisingly considering the current stratospheric prices of GPUs, the performance-per-dollar of compute is still rising faster than exponentially. In a handful years it will be cheap to train something as powerful as the models that cost millions to train today. Algorithmic efficiencies also stack up an make it cheaper to build and serve older models even on the same hardware.

It’s underappreciated that we would already be in a pretty absurdly wild tech trajectory just due to compute hyperabundance even without AI.

mordymoop··on The AI coding trap
Here's one. https://doofmovies.com/ With this project I'm sort of playing a game where I want to see how long I can go without finding out what language the backend is written in. I still don't know.
mordymoop··on The AI coding trap
Yeah, it will definitely do dumb stuff if you don’t keep an eye on it and intervene if you see the signs that it’s heading in the wrong direction. But it’s very good at course correcting and if you end up in a truly disastrous state you can almost always fix it be reverting to the last working commit and start a fresh context.
mordymoop··on The AI coding trap
Interesting point.

In most cases I would never have undertaken those projects at all without AI. One of the projects that is currently live and making me money took about 1 working day with Claude Code. It’s not something I ever would have started without Claude Code, because I know I wouldn’t have the time for it. I have built websites of similar complexity in the past, and since they were free-time type endeavors, they never quite crossed the finish line into commerciality even after several years of on-again-off-again work. So how do you account that with a time multiplier? 100x? Infinite speedup? The counterfactual is a world where the product doesn’t exist at all.

This is where most of the “speedup” happens. It’s more a speedup in overall effectiveness than raw “coding speed.” Another example is a web API for which I was able to very quickly release comprehensive client side SDKs in multiple languages. This is exactly the kind of deterministic boilerplate work LLMs are ideal for, and that would take a human a lot of typing, and looking up details for unfamiliar languages. How long would it have taken me to write SDKs in all those languages by hand? I don’t really know, I simply wouldn’t have done it, I would have just done one SDK in Python and said good enough.

If you really twist my arm and ask me to estimate the speedup on some task that I would have done either way, then yeah I still think a 100x speedup is the right order of magnitude, if we’re talking about Claude Code with Opus 4.1 specifically. In the past I spent about a five years very carefully building a suite of tools for managing my simulation work and serving as a pre/post-processor. Obviously this wasn’t full-time work on the code itself, but the development progressed across that timeframe. I recently threw all that out and replaced it with stuff I rebuilt in about a week with AI. In this case I was leveraging a lot of the learnings I gleaned from the first time I built it, so it’s not a fair one-to-one comparison, but you’re really never going to see a pure natural experiment for this sort of thing.

I think most people are in a professional position where they are sort of externally rate limited. They can’t imaging being 100x more effective. There would be no point to it. In many cases they already sit around doing nothing all day, because they are waiting for other people or processes. I’m lucky to not be in such a position. There’s always somewhere I can apply energy and see results, and so AI acts as an increasingly dramatic multiplier. This is a subtle but crucial point: if you never try to use AI in a way that would even hypothetically result in a big productivity multiplier (doing things you wouldn’t have otherwise done, doing a much more thorough job on the things you need to do, and trying to intentionally speed up your work on core tasks) then you can’t possibly know what the speedup factor is. People end up sounding like a medieval peasant suddenly getting access to a motorcycle and complaining that it doesn’t get them to the market faster, and then you find out that they never actually ride it.

I wonder, have you sat down and tried to vibecode something with Claude Code? If so, what kind of multiplier would you find plausible?

mordymoop··on The AI coding trap
Broadly the critique is valid where it applies; I don’t know if it accurately captures the way most people are using LLMs to code, so I don’t know that it applies in most case.

My one concrete pushback to the article is that it states the inevitable end result of vibe coding is a messy unmaintainable codebase. This is empirically not true. At this point I have many vibecoded projects that are quite complex but work perfectly. Most of these are for my private use but two of them serve in a live production context. It goes without saying that not only do these projects work, but they were accomplished 100x faster than I could have done by hand.

Do I also have vibecoded projects that went of the rails? Of course. I had to build those to learn where the edges of the model’s capabilities are, and what its failure modes are, so I can compensate. Vibecoding a good codebase is a skill. I know how to vibecode a good, maintainable codebase. Perhaps this violates your definition of vibecoding; my definition is that I almost never need to actually look at the code. I am just serving as a very hands-on manager. (Though I can look at the code if I need to - have 20 years of coding experience. But if I find that I need to look at the code, something has already gone badly wrong.)

Relevant anecdote: A couple of years ago I had a friend who was incredibly skilled at getting image models to do things that serious people asserted image models definitely couldn’t do at the time. At that time there were no image models that could get consistent text to appear in the image, but my friend could always get exactly the text you wanted. His prompts were themselves incredible works of art and engineering, directly grabbing hold of the fundamental control knobs of the model that most users are fumbling at.

Here’s the thing: any one of us can now make an image that is better than anything he was making at the time. Better compositionality, better understanding of intent, better text accuracy. We do this out of the box and without any attention paid to promoting voodoo at all. The models simply got that much better.

In a year or two, my carefully cultivated expertise around vibecoding will be irrelevant. You will get results like mine by just telling the model what you want. I assert this with high confidence. This is not disappointing to me, because I will be taking full advantage of the bleeding edge of capabilities throughout that period of time. Much like my friend, I don’t want to be good at managing AIs, I want to realize my vision.

mordymoop··on What happens when coding agents stop feeling like dialup?
From experience it seems like preempting context scoping and routing decisions to smaller models just results in those models making bad judgements at a very high speed.

Whenever I experiment with agent frameworks that spawn subagents with scoped subtasks and restricted context, things go off the rails very quickly. A subagent with reduced context makes poorer choices and hallucinates assumptions about the greater codebase, and very often lacks a basic sense the point of the work. This lack of situational awareness is where you are most likely to encounter js scripts suddenly appearing in your Python repo.

I don’t know if there is a “fix” for this or if I even want one. Perhaps the solution, in the limit, actually will be to just make the big-smart models faster and faster, so they can chew on the biggest and most comprehensive context possible, and use those exclusively.

eta: The big models have gotten better and better at longer-running tasks because they are less likely to make a stupid mistake that derails the work at any given moment. More nines of reliability, etc. By introducing dumber models into this workflow, and restricting the context that you feed to the big models, you are pushing things back in the wrong direction.

mordymoop··on Multi-Objective Process Control
"True" multi-objective optimization can be not only solve traditional optimization problems, but can act as a control system for dynamic time-varying multi-objective problems.
Page 1 of 8Next →