HNHacker News
TopNewBestAskShowJobs

adam_arthur

4,083 karma · joined December 23, 2020

Worked for a Silicon Valley startup from early days through successful acquisition.
submissionscomments
adam_arthur··on GPT-6 Sol and Luna
You can now start to add automations on top of typical dev flows.

There are a ton of use cases that open up with cheaper models.

E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc

adam_arthur··on Jev Ultrafast: A browser agent with a dynamic, indexed action space
Surprised by these takes.

Are people not getting that this (Jev) can do classification, programmatic branching, real time decision making (e.g. applicable to robotics) an order of magnitude faster and cheaper?

adam_arthur··on What will our economic future look like?
When wages decline, cost of producing the service/good does down too. And when the costs go down, relative demand goes up.

When relative demand goes up, the total labor needed expands.

How many Uber rides would you take at $5 vs $20?

How many times would you eat out at $10 vs $50?

Would you pay for a house cleaner at $25/visit over $150/visit?

A live in nurse at 50k/year vs 200k?

The costs of all these things will decline materially when the labor pool expands meaningfully due to displacement.

Technology has always been deflationary in the long run throughout history.

adam_arthur··on >10x More Efficient Pretraining
Certainly there's still a business there, I'm not saying they won't exist. But it's not going to be a monopoly-esque business with so many players in the ring, OpenAI, Anthropic, Google, Meta, Deepseek, Alibaba, GLM, Kimi etc. It will be cutthroat and a race to the bottom on price. And the difference from today -> 6 months ago intelligence will not be very meaningful.

Investors are largely treating these as future monopolies though.

We can already do so much with existing models. Harness improvements are probably more meaningful at this point.

e.g. say most image recognition can get saturated by a model of size xB parameters, so your tool for that can handoff to a smaller model. Document text extraction can use a model of size yB parameters. A model of size zB for summarizing text.

We are starting to get to a point where you can reasonably scope out an upper bound of required size/effort for many common tasks, and if you string these together, the frontier will largely act as an intelligent invoker of more efficient models.

Up until now there have been meaningful gains to each of those types of workstreams by using newer models, but that is starting to no longer be the case.

Yes, I do believe token consumption will rise exponentially from here in the near term. But cost of switching is low, and substantial profitability will be difficult.

adam_arthur··on >10x More Efficient Pretraining
If the output is commoditized, how much can you afford to pay for the input?
adam_arthur··on >10x More Efficient Pretraining
Yes, agree that token consumption will increase exponentially for the next while.

Disagree that the frontier model is where the economic gains will be realized.

The smaller the relative gap between frontier and non-frontier/open weights, the less pricing power.

This gap has shown only to shrink over time, not expand.

Businesses will pay more for frontier, but not meaningfully more to justify the economics. It's always going to be a low margin business, perhaps outside of cyber security, warfare/intelligence and perhaps drug discovery.

Though the expensive and time consuming part of drugs is doing the trials and getting approval, not coming up with ideas

adam_arthur··on >10x More Efficient Pretraining
There are an enormous number of tasks that can get by on good enough.

If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.

And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.

Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.

5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?

The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.

I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.

adam_arthur··on What will our economic future look like?
Nominal wage doesn't really matter, just real wages+purchasing power.

Goods and services prices will fall far more than wages for most people.

Why?

Because there will be far more efficiency and competition than pre-AI. It's plainly evident from the nature of the technology.

If one business can do something cheaper, other businesses will do it cheaper too, and have to cut prices to compete.

But the shock will be sudden so it won't feel good for our gen. Future generations will benefit without the drama.

Don't get me wrong, there will be big losers in some fields, and the short term will be painful due to retraining and loss of purpose/emotional toll.

Many of us built careers doing things that may not be relevant anymore. Or relevant in a different, perhaps diminished way.

But the same has happened to many professions throughout history and it's always led to general improvement of the broader public's welfare.

adam_arthur··on What will our economic future look like?
It seems obvious that labor displacement due to AI will create labor surpluses in existing fields.

Which will feel quite bad for many of us in the present, but net-net be good for society in the long run.

Wages will go down, but cost of goods and services will decline even more.

Markets will be more competitive than ever, which limits the moat/margin capability of many businesses.

It's only really particularly bad for those earning very high wages where there likely won't be any way to materially fill the gap.

adam_arthur··on GPT-6 Astra
I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.

I think some combination of:

1) Using 1 thread for everything

2) Reviving old threads which are no longer in cache

3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.

4) Non-coding workflow which is more output than input heavy

5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR

Just a guess. I think 3 is likely the primary reason.

You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME

adam_arthur··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
Codex (5.6 Sol) is extremely to the point and direct in its responses.

Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.

No idea why people are still using Claude models.

My impression is they started on Claude Code and never tried Codex or other harness+model combos

adam_arthur··on AI didn't erase the junior engineer's value, it increased it it
I see the complete opposite becoming the case.

In a world where code can be generated rapidly, it's super critical that you have a core few set of people who really understand the macro design of the codebase and can continue to factor it well and iterate quickly.

Adding more people and contributors just increases the probability that nobody really understands the structure of the codebase, it degrades into DRY and unfactored slop.

The cost of reviewing other people's code is almost too high to be worthwhile now... It's much easier to just cut them out and do it yourself.

A core set of very skilled people can just implement whatever change you are doing, but better, cleaner and faster.

I built a fairly large and complex project with Codex and had to spend about 50% of the time factoring things down as I went into well contained modules, had a full understanding of the architecture at a high level. It would have been pretty difficult to do this if bringing in other contributors.

Too many people comment about AI from the perspective of throwing feature A or B over the wall at the workplace, but anybody who has built a huge project from scratch will see how important good design is in regard to iteration speed and result quality.

That being said, there are still areas where changes should be sized reasonably and human reviewed e.g. foundational or very mature software

adam_arthur··on The Mysterious Syndrome Destroying Endurance Athletes
Yes, given all the recent data coming out, it seems to me HIIT is a much better bet for general health purposes than LISS style cardio.

Curious what the consensus will be a few years out.

adam_arthur··on 30-year Treasury yield tops 5.31%, the highest in 19 years
30y is keyed to inflation expectations.

If fed hiked to 5% tomorrow, 30y would invert and yield would go down.

It's not as simple as hikes lead to higher 30y yields.

adam_arthur··on The Mysterious Syndrome Destroying Endurance Athletes
Additionally recent studies have shown high duration/volumes of cardio increases plaque buildup in the arteries.

"Indeed, many studies have shown a relationship between endurance sports and higher volumes of coronary calcified plaque as determined by computed tomography."

https://pmc.ncbi.nlm.nih.gov/articles/PMC11395881/

adam_arthur··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Yes, if you set reasoning to none you can force the granularity of the thinking.

It will actually adhere to your request for e.g. 3 sentences max.

Thinking mode will override any instructions in the prompt (at least for other models in my experience).

Of course this will probably hurt performance, but works great for easy tasks that you know are trivial. Tons of pipeline, image recognition etc use cases where this works well.

I'd be curious to see Qwen 3.8 27B low thinking benchmarks though.

adam_arthur··on Software Engineering fundamentals matter more
There are tons of core principles that can be learned that largely apply across fields.

One prime example, single source of truth for data/concepts. To be violated only when performance is meaningfully improved (denormalized databases). But when you do so, you should definitely recognize you're opening up out of sync issues for that performance gain.

Though for the majority of code, there is no performance benefit to adding multiple sources of truth. Yet it's the most common error I see re: quality.

The sad thing is that software engineering fundamentals and best practices never became widespread or widely taught in school prior to LLMs

adam_arthur··on Abdominal fat predicts heart disease risk better than BMI
Fat is inherently unhealthy for you beyond some low baseline level.

Additional muscle is positive for health, but only up to some reasonable threshold. There is no health benefit to having very high levels of muscle, and in fact it may be negative for your health at extreme levels. E.g. many bodybuilders have trouble breathing, sleep apnea etc.

Both people in the example have 60lbs of fat. 1lb of muscle doesn't cancel out the negative health effect of 1lb of fat

adam_arthur··on Abdominal fat predicts heart disease risk better than BMI
I believe evidence points towards absolute level of fat being more important than % for overall health.

E.g. 20% bodyfat at 250lbs is still a lot of fat.

Of course it's difficult to ever get very high on absolute fat if at 15% or below.

adam_arthur··on Semaglutide linked to lower predicted dementia risk
It tends to happen when:

1) You must have a crowded optical disc as a precursor (low cup to disc ratio). Genetic and can be tested

2) You have stiff veins/arteries due to poor metabolic/cardio health

3) when blood pressure drops too much, blood flow can get cut off to the optical nerve temporarily due to reliance on high blood pressure due to 2

4) Due to loss of blood flow, optic nerve swells and closes blood flow into the eye due to 1

adam_arthur··on Semaglutide linked to lower predicted dementia risk
It's also true that having a crowded optical disc is pretty much a required precursor condition to suffering NAION.

Which you can get checked for via a $50 optical scan.

No crowded disc, very little risk.

(Cup to Disc ratio of 0.2ish or less starts to present risk)

adam_arthur··on Rails Is Built for AI
Abstractions allow for better business logic density and fewer tokens per concept/area of code. This is still very meaningful for today's LLMs.

Strictly speaking, if codebase A is doing the same thing as codebase B, but is 5-10x the LoC, you will be limited in how broadly you can effectively prompt the LLM. Queries will take longer, more technical debt+anti-patterns will creep in.

Anyone building large projects from scratch with LLMs will see plainly that they start to perform worse and slow down meaningfully as the codebase grows if you aren't pruning and compartmentalizing the code as you go.

Agree that abstractions which have a large performance cost are net-net not so worthwhile anymore. But many abstractions can be done with minimal performance hit (e.g. Rust)

adam_arthur··on uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook
You can already do this with smaller LLMs, like 10B range. The computer vision/image interpretation has been sufficient at this magnitude for a few months now.

Can run on most decently powered consumer devices.

Latency not great vs ad block though, of course.

Perhaps a fine tuned model explicit to this purpose could be even more compact.

adam_arthur··on Claude Opus 5
The first chart in the blog post shows a similar $/performance curve to GPT 5.6.

Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.

There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)

If they had a more efficient model at coding they would lead with that chart.

adam_arthur··on Claude Opus 5
GPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5.

Where are you getting cheaper per dollar?

adam_arthur··on Moonshot AI suspends new subscriptions due to Kimi K3 demand
Exactly.

I've found 6 minutes or so the sweet spot for upper bound with 5.6 Sol.

And it sounds like the OPs query above requires scanning throughout a large portion of the codebase, which will inherently consume a large number of input tokens. No locality to it.

adam_arthur··on Newly retired couples may lose $16,900/year in Social Security in 2033
The hole can be filled overnight by just raising the retirement age.

And if you normalize for longer lifespans, it's perfectly reasonable.

adam_arthur··on Kimi K3: Open Frontier Intelligence
When they can sell that for tens of Billions a year against competition, they might have a financial case.

In reality, there will be many clones of Claude Design, especially if it gains big revenue traction.

Doubly so if you believe the narrative that coding and apps will be "free" and instant to create in the future.

adam_arthur··on Kimi K3: Open Frontier Intelligence
Then it's hard to switch?

It's irrelevant the reason why, the margin any business can take will be constrained by the cost to switch to a cheaper, good enough, competitor.

You are saying it's difficult to switch due to compliance and admin issues. Ok!

LLMs can be swapped easily and open weight models can be hosted to adhere to whatever legal, uptime requirements etc are needed trivially.

The big cloud providers already host and resell open weight models.

The two don't event compare

adam_arthur··on Kimi K3: Open Frontier Intelligence
Completely different levels of stickiness.

The OS runs everything; the LLM can be swapped in a second.

(But yes, you will have to tweak prompts+tuning anytime you change models)

Page 1 of 34Next →