HNHacker News
TopNewBestAskShowJobs

AstroBen

1,795 karma · joined April 6, 2025

submissionscomments
AstroBen··on System Card: Claude Mythos Preview [pdf]
Right, and that's why it's only part of the job. The benchmarks they're currently doing compose of the AI being handed a detailed spec + tests to make pass which isn't really what developing a feature looks like.

Going from fuzzy under-defined spec to something well defined isn't solved.

Going from well defined spec to verification criteria also isn't.

Once those are in place though, we get https://vinext.io - which from what I understand they largely vibe-coded by using NextJS's test suite.

> First one that comes to mind is that 100% code coverage in tests means that software is perfect

I agree.. but I'm also not sure if software needs to be perfect

AstroBen··on System Card: Claude Mythos Preview [pdf]
> If it can replace SWEs, then there's no reason why it can't replace say, a lawyer

SWE is unique in that for part of the job it's possible to set up automated verification for correct output - so you can train a model to be better at it. I don't think that exists in law or even most other work.

AstroBen··on System Card: Claude Mythos Preview [pdf]
Everyone wouldn't starve in a few months. There is more than enough food and I have faith it'd be given out. The starvation we see today in a world where most genuinely have a chance to get out of it is nothing like a world in which people can't earn an income.

The government only has as much power as they are given and can defend, and the only way I could see that happening is via automated weapons controlled by a few- which at this point aren't enough to stop everyone. What army is going to purge their own people? Most humans aren't psychopaths.

I think it'd end in a painful transition period of "take care of the people in a just system or we'll destroy your infrastructure".

AstroBen··on System Card: Claude Mythos Preview [pdf]
It seems inevitable that costs will come down over time. Expensive models today will be cheap models in a few years.
AstroBen··on System Card: Claude Mythos Preview [pdf]
Of course it's what they're going for. If they could do it they'd replace all human labor - unfortunately it's looking like SWE might be the easiest of the bunch.

The weirdest thing to me is how many working SWEs are actively supporting them in the mission.

AstroBen··on Sam Altman may control our future – can he be trusted?
Our DNA does contain our pre-training, though. It's not true that we're an entirely blank slate.
AstroBen··on The cult of vibe coding is dogfooding run amok
It's not subjective at all. It's not art.

Code quality = less bugs long term.

Code quality = faster iteration and easier maintenance.

If things are bad enough it becomes borderline impossible to add features.

Users absolutely care about these things.

AstroBen··on The cult of vibe coding is dogfooding run amok
The difference here is that everyone else in this product category are also sprinting full steam ahead trying to get as many users as they can

If they DIDN'T heavily vibe-code it they might fall behind. Speed of implementation short term might beat out long-term maintenance and iteration they'd get from quality code

They're just taking on massive tech debt

AstroBen··on The cult of vibe coding is dogfooding run amok
> it doesn't really matter in the end

if you have one of the top models in a disruptive new product category where everyone else is sprinting also, sure..

AstroBen··on The cult of vibe coding is dogfooding run amok
99.999999% of products can't get away with what Anthropic is able to - this is a one in a billion disruptive product with minimal competition, and its success so far is mostly due to Claude the model, not the agent harness
AstroBen··on Ask HN: Where are all the disruptive software that AI promised?
Strange, even
AstroBen··on I Won't Download Your App. The Web Version Is A-OK
Do you really think developers are going through the hellish pain of dealing with Google and Apple for no reason? Real world users prefer and expect apps as opposed to web versions for many product categories.
AstroBen··on Codex pricing to align with API token usage, instead of per-message
Kimi K2.5 (as an example) is an open model with 1T params. I don't see a reason it has to be local for most use cases- the fact that it's open is what's important.
AstroBen··on Codex pricing to align with API token usage, instead of per-message
Unfortunately the fools holding the bag are going to be those who own index funds when these companies are inserted into them.
AstroBen··on Codex pricing to align with API token usage, instead of per-message
Things must be bad if they're doing this before their IPO
AstroBen··on The threat is comfortable drift toward not understanding what you're doing
By working in this way you're proactively de-skilling yourself. Do it long enough and you're now replaceable by anyone that can type a prompt.
AstroBen··on I'm 60 years old. Claude Code killed a passion
Fun fact: the person who wrote that original "Claude Code has re-ignited a passion" never commented or posted again.

In fact that was their first and only contribution.

Weird.

AstroBen··on Speed at the cost of quality: Study of use of Cursor AI in open source projects (2025)
"Notably, increases in codebase size are a major determinant of increases in static analysis warnings and code complexity, and absorb most variance in the two outcome variables. However, even with strong controls for codebase size dynamics, the adoption of Cursor still has a significant effect on code complexity, leading to a 9% baseline increase on average compared to projects in similar dynamics but not using Cursor."
AstroBen··on Speed at the cost of quality: Study of use of Cursor AI in open source projects (2025)
They're measuring development speed through lines of code. To show that's true they'd need to first show that AI and humans use the same number of lines to solve the same problem. That hasn't been my experience at all. AI is incredibly verbose.

Then there's the question of if LoC is a reliable proxy for velocity at all? The common belief amongst developers is that it's not.

AstroBen··on Speed at the cost of quality: Study of use of Cursor AI in open source projects
> On average, Cursor adoption has a modestly significant positive impact on development velocity, particularly in terms of code production volume: Lines added increase by about 28.6% (Table 2). There is no statistically significant effect for the volume of commits.

This doesn't equate to a faster development speed in my eyes? We know that AI code is incredibly verbose.

More lines of code doesn't equate to faster development - even more so when you're comparing apples (human written) to oranges (AI written)

AstroBen··on US Job Market Visualizer
Uh huh.. but the data in Andrej's visualizer is showing software development growth outlook is at 15% (much faster than average)

Over the past year (where Opus has supposedly changed the game), we're seeing ~10% more job postings for software developers compared to this time last year [1,2]

A huge amount of our work is not easily verifiable, therefore it's extremely hard to actually train an LLM to be better at it. It doesn't magically get better across the board.

AI HAS WON. SURF OR DROWN. YOU DONT KNOW WHATS COMING!!!?!?!

Stop with this doomer drivel. It's sick. It's not based in reality and all it does is stress innocent people out for no reason.

1: https://fred.stlouisfed.org/series/IHLIDXUSTPSOFTDEVE

2: https://trueup.io/job-trend

AstroBen··on How I write software with LLMs
I do too, but it comes from a bang-for-your-buck and not a test coverage standpoint. Test coverage goes up in importance as you lean more on AI to do the implementation IMO.
AstroBen··on How I write software with LLMs
This is fantasy completely disconnected from reality.

Have you ever tried writing tests for spaghetti code? It's hell compared to testing good code. LLMs require a very strong test harness or they're going to break things.

Have you tried reading and understanding spaghetti code? How do you verify it does what you want, and none of what you don't want?

Many code design techniques were created to make things easy for humans to understand. That understanding needs to be there whether you're modifying it yourself or reviewing the code.

Developers are struggling because they know what happens when you have 100k lines of slop.

If things keep speeding in this direction we're going to wake up to a world of pain in 3 years and AI isn't going to get us out of it.

AstroBen··on Ask HN: How is AI-assisted coding going for you professionally?
If an LLM could be profitable trading why wouldn't the creators use it themselves and not release it? It'd be by far the most profitable thing they could do.
AstroBen··on What is agentic engineering?
Yes I'm with you. I spent the last 2 months heavily doing "agentic engineering" and I don't think it's optimal to work like that as a default.

LLMs are for sure useful and a productivity boost but generating 99% of your code with it is way overdoing it.

AstroBen··on Ask HN: How is AI-assisted coding going for you professionally?
You're trusting AI to trade with your real money?
AstroBen··on Allow me to get to know you, mistakes and all
I'm stating it here before anyone accuses me of being an LLM that I love using fanfic drama dots and I loved them before AI started with it.
AstroBen··on Allow me to get to know you, mistakes and all
The problem with AI writing isn't its style, it's the content.

It's full of fluff. Analogies that sound like something a 12 year old would make, but make no sense when you stop to think about them.

It's full of baloney that the author didn't even intend to communicate.

That's where the "soulless" part comes from. There's no consistent mind behind the writing with opinions of its own, formulated into one understandable framework it's trying to convey. It's just a mishmash of BS that only superficially resembles it, made to trick us.

AstroBen··on The 100 hour gap between a vibecoded prototype and a working product
Because now that website is fully cross-platform and sandboxed with no practical downside
AstroBen··on The 100 hour gap between a vibecoded prototype and a working product
The average user doesn't even know what a file is
← PreviousPage 2 of 24Next →