Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.
Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.
For large tasks like a web browser or a compiler, even expensive swarms of frontier LLMs have not been shown capable of producing codebases that actually work. (Anthropic built a C compiler with Opus 4.6 but it lacked optimizations and apparently hit a complexity wall.)
I also want to use LLMs for reverse engineering, but apparently it's pretty hit-or-miss, especially if you're forced to use open-source models to avoid restrictions.
It's also interesting because, while coding agents are important and are a notable success, they are never going to be a multi trillion dollar business. And are there any other domains where LLMs have such a large impact?
Seriously something feels really off about Opus 5. I hope they correct it before 4.6 is removed.
“Very few people actually require a Pentium workstation, a 486 is perfectly adequate for the majority”
The logical fallacy is taking an extant distribution of “product capability” that is priced to fit what the market will bear and assuming the “next upgrade” simply tacks on a little bit more to the right hand rail of that curve.
No!
It shifts the entire curve!
Everything for everyone gets better and the top 1% of the most demanding users will continue to pay the same-ish premium.
“Nothing” will change.
Look at it this way: you can buy a $200 laptop for your kid or a $20,000 Mac with an M5 Ultra processor.
BOTH are vastly more powerful than either a $200 PC or a $20,000 “workstation” from 20+ years ago.
Look at: https://arena.ai/leaderboard/text?q=openai&utm_source=chatgp...
The “budget” 5.5 Instant model beats o1 and o3 which were “pro” models at the time of their release!
In other words. PC users didn't figure out that they could buy super powerful PCs and play games on them, that was a carefully managed market transition.
What is going to do the same for LLMs?
It wasn't "Intel" that found new uses for PCs, it was everybody who found new uses for them. Billions of people and millions of companies found uses for "more computer power".
It was only the journalists with limited imaginations (and no industry experience) who struggled to come up with potential uses.
> carefully managed market transition.
You make it sound like a conspiracy! It wasn't. It was simple capitalist competition. If Intel hadn't improved their products, their competitors would have left them behind.
That very nearly happened ten years ago because Intel become stuck on the 14nm process and their products stagnated while Apple, ARM, and AMD lapped them repeatedly.
> What is going to do the same for LLMs?
Everybody.
Are you saying that unless you're "carefully managed" by some third-party, you could not find any use for "unlimited intelligence on tap"?
The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.
Both animated and live action results would be acceptable.
Unfortunately most existing LLMs lack the capability to maintain context across tens of thousands of frames.
I think, also, like in the traditional film makers career, this process should be built iteratively, start with a fast food commercial, then do a music video, then you can probably do a short film. Continue to improve the process, and one day I’m sure the LLM film studio can make you any movie you want, provided you have enough tokens.
EDIT: Your username doesn't help, either.
Sort of, but I want it to be relatively low on human effort. I feel burned by spending lots of time in 2023 learning image generation pipelines (using control net etc) only for that to be rendered trivial by the next generation of LLMs.
This movie would be only for personal consumption and I’m okay with waiting for model improvements.
I've done this sort of with comfyui/same agent factory stuff, but the verification loop only works for models like fable as planner/writer, with gemini as verifier for like a very short movie. Sub 3-5 mins. After that you burn through million tokens.
Can't go too low fidelity audio/video or it craps out. Too long video and it loses consistency. Look at only snippets, it lacks global consistency, etc.
There are two more points in favor of this kind of AI movie project: there's zero chance that anyone would greenlight a Hollywood budget for the Silmarillion, and it is beyond human capability to write that screenplay.
Plus, the token costs involved should be pretty low! (Other costs may not be.)
Realistic, scientifically useful simulations still require tuning all sorts of parameters based on physical intuition and understanding of the system being simulated. Both Sol and Fable/Opus 5 fail at it and either blow up the computation cost to levels that can't be processed realistically, or they invent a justification for a visibly unphysical result.
Super computers keep getting better but most people don't need them for most things.
In 10 or 20 years maybe we'll all be running AI models that are currently considered "frontier" on smart phone type devices.
The frayed edges on what I have slopped together as unreasonably ambitious, ludicrous projects with fucktons of tokens from models 6-9 months ago mostly look like situations where a capable-enough-to-be-dangerous developer tries to muscle through problems that explode in width & depth but keep digging (so, a tier below stopping early to do more design, two below recognizing the need for more planning from the outset). The primitives are there, most major things work well enough, but the remaining functionality and performance is inaccessible. At a cost of multiples of >1/8th of a $200/mo subscription.
Right now I can put $10 into DSv4 Pro/Flash or Qwen 3.8 Max/Flash, hand it a project and all of its unfinished forks in a state I barely remember, tell it that I want the things these forks have been working towards, and 8 hours later it has ie an working, tested, benchmarked multicore car physics simulation fabric with all of the forks evaluated, the gains merged in, the remaining work documented. It only needed a few hundred more lines of code but Codex 5.5 was never going to see that.
9 months ago I was saying developers are not being ambitious enough with these things, that's only more true now. They lend themselves to digging far deeper than they should: 200kloc god files, dozens of forks. Let it happen, you don't need to read it, it's for them, later. The only time you make them clean up is when it has a severe adverse effect on how long builds/lints/tests/benchmarks take. Spend a whole week having it do nothing but dig up published papers in relevant fields with cutting edge techniques and translating them into feature specs. Pick whatever state of the art is and try to crush it, throw everything at it, leave it looping on vague but wildly ambitious goals. When it modularizes and refactors it all down you might 'do a breakthrough', or maybe it happens in an hour over christmas break when you're trying the next one.
I feel bad for him as a human, but as a cycling fan I'm glad that we'll have an interesting WC in Canada