There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc
4,083 karma · joined December 23, 2020
There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc
Are people not getting that this (Jev) can do classification, programmatic branching, real time decision making (e.g. applicable to robotics) an order of magnitude faster and cheaper?
When relative demand goes up, the total labor needed expands.
How many Uber rides would you take at $5 vs $20?
How many times would you eat out at $10 vs $50?
Would you pay for a house cleaner at $25/visit over $150/visit?
A live in nurse at 50k/year vs 200k?
The costs of all these things will decline materially when the labor pool expands meaningfully due to displacement.
Technology has always been deflationary in the long run throughout history.
Investors are largely treating these as future monopolies though.
We can already do so much with existing models. Harness improvements are probably more meaningful at this point.
e.g. say most image recognition can get saturated by a model of size xB parameters, so your tool for that can handoff to a smaller model. Document text extraction can use a model of size yB parameters. A model of size zB for summarizing text.
We are starting to get to a point where you can reasonably scope out an upper bound of required size/effort for many common tasks, and if you string these together, the frontier will largely act as an intelligent invoker of more efficient models.
Up until now there have been meaningful gains to each of those types of workstreams by using newer models, but that is starting to no longer be the case.
Yes, I do believe token consumption will rise exponentially from here in the near term. But cost of switching is low, and substantial profitability will be difficult.
Disagree that the frontier model is where the economic gains will be realized.
The smaller the relative gap between frontier and non-frontier/open weights, the less pricing power.
This gap has shown only to shrink over time, not expand.
Businesses will pay more for frontier, but not meaningfully more to justify the economics. It's always going to be a low margin business, perhaps outside of cyber security, warfare/intelligence and perhaps drug discovery.
Though the expensive and time consuming part of drugs is doing the trials and getting approval, not coming up with ideas
If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.
And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.
Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.
5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?
The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.
I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.
Goods and services prices will fall far more than wages for most people.
Why?
Because there will be far more efficiency and competition than pre-AI. It's plainly evident from the nature of the technology.
If one business can do something cheaper, other businesses will do it cheaper too, and have to cut prices to compete.
But the shock will be sudden so it won't feel good for our gen. Future generations will benefit without the drama.
Don't get me wrong, there will be big losers in some fields, and the short term will be painful due to retraining and loss of purpose/emotional toll.
Many of us built careers doing things that may not be relevant anymore. Or relevant in a different, perhaps diminished way.
But the same has happened to many professions throughout history and it's always led to general improvement of the broader public's welfare.
Which will feel quite bad for many of us in the present, but net-net be good for society in the long run.
Wages will go down, but cost of goods and services will decline even more.
Markets will be more competitive than ever, which limits the moat/margin capability of many businesses.
It's only really particularly bad for those earning very high wages where there likely won't be any way to materially fill the gap.
I think some combination of:
1) Using 1 thread for everything
2) Reviving old threads which are no longer in cache
3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.
4) Non-coding workflow which is more output than input heavy
5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR
Just a guess. I think 3 is likely the primary reason.
You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME
Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.
No idea why people are still using Claude models.
My impression is they started on Claude Code and never tried Codex or other harness+model combos
In a world where code can be generated rapidly, it's super critical that you have a core few set of people who really understand the macro design of the codebase and can continue to factor it well and iterate quickly.
Adding more people and contributors just increases the probability that nobody really understands the structure of the codebase, it degrades into DRY and unfactored slop.
The cost of reviewing other people's code is almost too high to be worthwhile now... It's much easier to just cut them out and do it yourself.
A core set of very skilled people can just implement whatever change you are doing, but better, cleaner and faster.
I built a fairly large and complex project with Codex and had to spend about 50% of the time factoring things down as I went into well contained modules, had a full understanding of the architecture at a high level. It would have been pretty difficult to do this if bringing in other contributors.
Too many people comment about AI from the perspective of throwing feature A or B over the wall at the workplace, but anybody who has built a huge project from scratch will see how important good design is in regard to iteration speed and result quality.
That being said, there are still areas where changes should be sized reasonably and human reviewed e.g. foundational or very mature software
Curious what the consensus will be a few years out.
If fed hiked to 5% tomorrow, 30y would invert and yield would go down.
It's not as simple as hikes lead to higher 30y yields.
"Indeed, many studies have shown a relationship between endurance sports and higher volumes of coronary calcified plaque as determined by computed tomography."
It will actually adhere to your request for e.g. 3 sentences max.
Thinking mode will override any instructions in the prompt (at least for other models in my experience).
Of course this will probably hurt performance, but works great for easy tasks that you know are trivial. Tons of pipeline, image recognition etc use cases where this works well.
I'd be curious to see Qwen 3.8 27B low thinking benchmarks though.
One prime example, single source of truth for data/concepts. To be violated only when performance is meaningfully improved (denormalized databases). But when you do so, you should definitely recognize you're opening up out of sync issues for that performance gain.
Though for the majority of code, there is no performance benefit to adding multiple sources of truth. Yet it's the most common error I see re: quality.
The sad thing is that software engineering fundamentals and best practices never became widespread or widely taught in school prior to LLMs
Additional muscle is positive for health, but only up to some reasonable threshold. There is no health benefit to having very high levels of muscle, and in fact it may be negative for your health at extreme levels. E.g. many bodybuilders have trouble breathing, sleep apnea etc.
Both people in the example have 60lbs of fat. 1lb of muscle doesn't cancel out the negative health effect of 1lb of fat
E.g. 20% bodyfat at 250lbs is still a lot of fat.
Of course it's difficult to ever get very high on absolute fat if at 15% or below.
1) You must have a crowded optical disc as a precursor (low cup to disc ratio). Genetic and can be tested
2) You have stiff veins/arteries due to poor metabolic/cardio health
3) when blood pressure drops too much, blood flow can get cut off to the optical nerve temporarily due to reliance on high blood pressure due to 2
4) Due to loss of blood flow, optic nerve swells and closes blood flow into the eye due to 1
Which you can get checked for via a $50 optical scan.
No crowded disc, very little risk.
(Cup to Disc ratio of 0.2ish or less starts to present risk)
Strictly speaking, if codebase A is doing the same thing as codebase B, but is 5-10x the LoC, you will be limited in how broadly you can effectively prompt the LLM. Queries will take longer, more technical debt+anti-patterns will creep in.
Anyone building large projects from scratch with LLMs will see plainly that they start to perform worse and slow down meaningfully as the codebase grows if you aren't pruning and compartmentalizing the code as you go.
Agree that abstractions which have a large performance cost are net-net not so worthwhile anymore. But many abstractions can be done with minimal performance hit (e.g. Rust)
Can run on most decently powered consumer devices.
Latency not great vs ad block though, of course.
Perhaps a fine tuned model explicit to this purpose could be even more compact.
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
Where are you getting cheaper per dollar?
I've found 6 minutes or so the sweet spot for upper bound with 5.6 Sol.
And it sounds like the OPs query above requires scanning throughout a large portion of the codebase, which will inherently consume a large number of input tokens. No locality to it.
And if you normalize for longer lifespans, it's perfectly reasonable.
In reality, there will be many clones of Claude Design, especially if it gains big revenue traction.
Doubly so if you believe the narrative that coding and apps will be "free" and instant to create in the future.
It's irrelevant the reason why, the margin any business can take will be constrained by the cost to switch to a cheaper, good enough, competitor.
You are saying it's difficult to switch due to compliance and admin issues. Ok!
LLMs can be swapped easily and open weight models can be hosted to adhere to whatever legal, uptime requirements etc are needed trivially.
The big cloud providers already host and resell open weight models.
The two don't event compare
The OS runs everything; the LLM can be swapped in a second.
(But yes, you will have to tweak prompts+tuning anytime you change models)