But the larger problem is sound, and the answer is something jointly optimized (idk how they do the routing) but it’s hard to shoehorn it into the current paradigm.
364 karma · joined December 29, 2025
But the larger problem is sound, and the answer is something jointly optimized (idk how they do the routing) but it’s hard to shoehorn it into the current paradigm.
> I think the last month has resulted in a lot of people saying “f u I would rather fail in life than serve you”.
I agree a lot of people are saying this and doing this but this won't last long. People do not want to suffer or fall behind their peers. At some point not using AI for anything will not be a feasible option. Unless you operate from an incredibly uncommon and privileged financial position, you don't really even have a choice.
I don't actually have a use case in mind where you would _need_ or get a substantial competitive advantage of the absolute frontier in some sort of workflow or whatever load bearing DAG there is for your business. But if there are cases like this, that will suck.
I agree with this, and I think you are right about a lot of this being really just not market-relevant factors that are driving a lot of the downsides here. All of these should be regulated, and even the ones where regulating will have impact to markets and consumers.
And I agree -- effective advertising targeting necessitates "reward hacking" in that I can get you to buy something if I prey on your insecurities or otherwise manipulate you in a way that goes beyond "oh I know what watwut would love as a birthday gift".
But you really should be aware that effective advertising really does disproportionately benefit smaller businesses and this is generally a Good Thing. Really my only point here is that there is a baby in this bathwater and we should be careful.
This is why, because this completely misses the point. Every power law has diminishing marginal returns, that is by construction. The point is that they diminish predictably and smoothly across seven orders of magnitude of compute. There is no plateau anywhere in the observed range, and in fact larger models more sample efficient than expected.
You can't just cherry pick one sentence containing the phrase "diminishing returns" and ignore the actual paper. The whole reason the labs spent the GDP of a small country on GPUs is that the curve is boring and reliable. No one just was like "oops we missed that line in the Kaplan paper, shit".
Lots of arguments as to why things may slow down or hit fundamental limits requiring a paradigm shift. But at several points in the past 5 years many major players have assumed this to be the case and have been continually proven wrong. Not to mention the plethora of socioeconomic horrors coming out of this Pandora's box, which I believe to be a large colocated B300 compute cluster somewhere in the midwest.
A point that would have been consistent with Kaplan was them finding alpha ~ 0.05 on compute means halving the reducible loss is a million times more expensive. But this was invalidated by Chinchilla who found it to be ~0.36 (100x more compute, not 1million). And loss has a non-linear relationship with capabilities, but roughly halving pretraining loss is ~doubling the METR task horizon. We also haven't factored in something incredibly important which is: algorithmic / architectural / training pipeline gains. These are substantial too!
Regardless of any export controls (or really anything else) China would be aggressively ramping up chip production capabilities as fast as they possibly could. Export controls maybe help slightly here but it's worth it. China will succeed too, but it takes time. China has been pursuing this for decades.
> Let's face it - all bans were dumb.
Predictably, this is oversimplifying.
> They just gave China the legal (per WTO rules) justification to start producing everything domestically.
It just doesn't matter, reliance on chips from an adversary is an untenable place to be. They don't need any legal excuses.
> he bans work as a reverse tariff, as a protectionist measure that actually protects your competitor.
Again untrue -- tariffs are devastating economically but BOTH the US and China want independence from another, it's just a very painful cost economically to do the divorce. It's far less efficient and horrifies economists but geopolitics is multifaceted.
> If China did those, others could bring China to court at the WTO. But the US did that, so nobody can sue China.
Hmmmmm would this really work? Would this have any meaningful impact whatsoever on their strategic initiatives? I highly doubt that.
But also I think people forget: this is not cut and dried. This is not simple. This is not just "these companies are evil and should be stopped". It is: there is a market pressure to do these things. "I enjoy the product and use it a lot" and "I am addicted" is blurry and market pressure is not going to recognize that limit because it does not care about human suffering unless that suffering meaningfully impacts the bottom line.
If these companies hit regulations that effectively cap their advertising revenue per user (i.e. the "addictiveness"), they are dead. That may be totally fine, and I'm sure majority of people would rejoice hearing this. But remember: advertising dollars are earned, especially at tech company scale, by the effectiveness of the targeting to get get more $ / DAU since DAU cannot grow beyond the Earth's population and that is the scale that these companies have already achieved.
If you cap advertising dollars, you cap advertising effectiveness. You cap the ability for small companies to connect quickly with prospective customers without being locked out because they have to spend too much to find them. Yes you also cap scammers and other nefarious actors too, but thats arguably a different issue. The impact of reducing advertising effectiveness is disproportionately concentrated on small business where cheap and effective advertising is so important.
My hope is that there is a way to do both and I don't have to be constantly horrified when I look at my screen time hours.
But
- Plenty of businesses are stuck between a rock and a hard place: make 3p models load bearing, or sacrifice performance to the competitors that are willing to swallow that risk
- I just cannot imagine a viable equilibrium where OSS models compete on capabilities with the frontier without a hidden payer and shaky economics. Quant funds, some sovereign AI effort, cloud business, sanctioned distilled frontier models; all of these are demonstrably viable vehicles for OSS development. But the optimal position for these purposes is not to be the best, it's to be good enough (which I think is the position they find themselves in). I can imagine temporary points where open models pull ahead but not sustainably.
I agree there are many many use cases where the optimal choice is to reach for open weight models (I assume yours is one of them).
But the economics of frontier model development necessitates that these models are behind and thats a problem for other cases.
Nowhere near the value of having access to chips, at any cost. They have extremely deep pockets. They already pay 6x the cost per FLOP.
> Instead, the US banned China from chips and lithography machines, giving China the legal excuse to start producing them domestically without violating WTO rules. Now China produces cheap chips and uses them with cheap electricity.
You think without export restrictions China wouldn't be doing the exact same thing? China needs absolutely zero legal excuse. I mean sure they have compute available on grey market / domestically but at 6x the cost per FLOP. Access to NVIDIA chips would make it dramatically cheaper for them. Yes you get chip income but that is not even close to what you lose. The strategy is doing what it was always supposed to do: slow them down, bleed their resources to force them to spin their wheels catching up. China is doing a great job with this but they are fundamentally constrained by these export controls.
You are right that this greases the wheels, they are further along than they would have been without export restrictions, but they are still delayed even with the reduced friction. The alternative is that they move slightly slower _while having the same compute infrastructure available_ and at dramatically lower energy costs. That is a far worse position for the US to be in.
> This was a dumb move by the US. Brought upon it by dumbf-ck aristocratic elites who grew up in isolated mansions and then received law degrees, with absolutely no understanding of technology and technology ecosystems. They thought they'd just make the rules and everybody would have to obey. It turns out in technology, they don't have to...
I think this is too cynical. Neither one of us is in the room to actually observe the real decision making, but export restrictions as a strategy are not some "dumbf-ck aristocratic elite" thing. They are perfectly rational from a strategic standpoint and arguably doing what they're supposed to do.
The tradeoff is worth it. They’re even publishing papers which blows me away — their efficiency gains quickly become incorporated into frontier models because they are open sourcing them. They would be aggressively pursuing the same chip pipeline strategy as they are today.
US is lagging in efficiency work because the ROI is better elsewhere for us. We have the same tier of talent, once the script flips so can the research.
Of course they have ways around this -- you can get black market GPUs and also API costs are SUPER cheap there -- they hack the subscription model, bundle a bunch of user accounts, and route API requests through them.
And yes they are getting to parity with US technology and will get there in a few years, they have decent chips but still not the quality of NVIDIA.
It's really a very complex situation
Im not sure exactly what you’re saying here, is it that you trade off coding performance and performance on other tasks like communication? If so (correct if not) this isn’t true — generalization happens. Doing good on coding lifts all boats.
> Somehow the hype over the successes with coding in the last year or so made everyone forget the intrinsic limit posed by the exhaustion of real human text output, which is absolutely inescapable
You’re absolutely right that we’re quickly running out of human text data but that isn’t at all the limitation you think it is. No one has “forgotten” this — coding agent performance is primarily from reinforcement learning on synthetic data traces with verifiable rewards, though pretraining is still important.
Also don’t forget: there is a world of multimodal data (video, audio, 3D maps, etc) that is incredibly rich.
- scaling laws exist
- downstream perf trends also exist (epoch capability index)
- gpt4 to gpt5 leap in capabilities every 16-18 months
- actual adoption and retention and engagement numbers are out of this world
- RL with verifiable rewards will get you to super human performance even with poor sample efficiency
- zero evidence of a plateau
I don’t know the future I am just skeptically looking at the data.
What data are you looking at?
Is there a level of abstraction where human involvement will always be necessary? If so where?
Also, frontier token prices have remained roughly constant:
3.5 sonnet: $3/$15 3.7 sonnet: $3/$15 Opus 4: $15/$75 (opus tier) opus 4.1: same Opus 4.5: $5/$25 Opus 4.6 (same) 4.7 (same) 4.8 (Same) Fable: $10/$50
So Fable is cheaper than Opus 4 was at launch.
One thing that has increased quite significantly? Spending and adoption.
> Consciousness isn’t really guessing the next thing to say-
I don't know what consciousness is either and these debates are a dumpster fire when they happen, but it sounds like you're pulling forward this "LLMs are just predicting the next token" (true by construction) implies that they can't learn or reason or be conscious (2/3 are wrong, the last one isn't falsifiable without a useful definition).
> they require the ability to abstract in a way that the question itself does not immediately suggest
yes, yet there are multitudes of other measurements of the same kind where LLMs reason perfectly well and better in many cases than a human could.
> Math problems are different in that they are described in terms of art that are closely related to certain patterns of manipulation (that is, the paper texts tend to contain both in close proximity to one another).
Is your logic really that math problems are actually easier to answer without reasoning and just by blending together closely related papers? I would definitely suggest reading the literature a bit more on this topic.
That models cannot do ALL logic problems does not mean that they cannot properly use logic...they can write Lean-verified theorems. How is that not logic?
> They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature.
What does their "grasp[ing] their true nature" have anything to do with what they can do?
> In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?" would be answered correctly and reliably.
Again, just because you have interesting failure modes or brittleness does not mean they do not reason.
Since you refuse to actually define what you consider to be reasoning let me at least put one out there: a system exhibits reasoning when an answer depends on nontrivial intermediate computation over the problem. If you find problems with this, fine, but just make an effort to contribute an alternative.
If you increase test time compute you get better performance. If the model was just "interpolating" this wouldn't really work would it? Models can do FrontierMath expert problems (unpublished, expert authored, peer reviewed math problems) that require an insane amount of compositional reasoning. If they were regurgitating training data, that wouldn't really work would it? Chain of thought, while not always faithful to internal computation, improves performance. If the models were just regurgitating information, it wouldn't work that well would it?
"regurgitating training data" is also of course misleading. Yea they can memorize parts of the training data, but they generalize very well.