I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.
But in the last few days something seems to have happened that made Codex's models massively stupider (for what I am doing).
Really weirdly, it suddenly refused to even run tests it previously wrote itself (and previously ran), because of some false positive about cybersecurity.
That by itself is not evidence of stupidity. Trying to make a 200+ file PR full of research notes is, and the PR didn't even solve the problem I asked it to.
For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine.
Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious
Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.
Just download claude code or codex and ask it to give suggestions about where to integrate agents into your workstream.
I couldn't imagine being so presumptuous as to know that my workflow fits all sizes, and all others are just holding it wrong – or worse, they're not doing real work. It would take a bigger ego on my part, or maybe less social awareness, to presume this.
> but if you think it outright doesn't have any benefits over Opus 4.8 then your workflow is probably not making good use of the tools.
I don't even use claude, I give exactly zero shits about fable or opus or bingus bongus.
And what even are these ambitious companies and people one shotting and building with Fable? AI has been around for almost 3 years now. Tell me one app or software you use which has gotten significantly better and has amazing new useful features landing on a weekly basis? If anything, every single software product I use has gotten worse.
I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value for me.
Mediocre means average / middle of the pack. It sounds like its doing exactly what you would expect nothing more. Why would you stop it? Why would you need exceptional?
https://www.merriam-webster.com/dictionary/mediocre
Clearly they were using the word to mean low quality. Why would you ask this odd question?
Mediocre means of only ordinary or moderate quality—neither very good nor very bad, and often slightly disappointing
The quality is average but expectations of high quality are not met. He expected more but got what he asked for. We overuse top models because of this.
Opus delivered mediocre results. Not garbage, but would have required me to do lot's of things myself. Fable did not needed my supervision with this task.
i swear they trained in on threejs in particular so those idiots on twitter could spam their garbage demos
For coding it's a little harder to tell, but at least the prose feels a little better.
The reality is, it doesnt matter if LLMs keep getting more powerful because they still need a human to steer it. Without the human providing inputs to the LLM it just sits there and does nothing.
You can, for example, hook it up to a logging system and have it fix errors as they occur on your platform.
I’d be curious about:
- your setup. How it all works - The types of errors it fixed and how quickly - Any regressions or issues it caused - The cost
Thanks!
It works surprisingly well. The errors fixed are both genuine errors in the harness itself, but increasingly so upstream bugs (in the underlying agent apps like Codex, or in Herdr, which is used to expose uniform programmatic access to all those different apps) for which it needs to come up with workarounds. No regressions so far.
The cost is hard to judge on a subscription, especially when you're running really heavy tasks otherwise that dwarf any harness work.
I think it's worth acknowledging that the power of LLMs at this point is not really so much in the smarts, but in the coordination and the surrounding harness tech. "Written english" turning into sequences of commands[0]. The whole agentic "stuff" in general. Tools + coordination is the superpower. The reasoning... it doesn't have to be _that_ good for the rest of the stuff to work. On good codebases and infra, at least.
And I say this as someone who really would rather most of this stuff disappear!
[0]: programming is obviously text to commands, but there's a loooooooot of futziness that LLM reasoning has let us remove in some flows
You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX Spark. And that’s just an architecture preview. The Qwen 4 family is expected to arrive this fall.
Do you know what kinda throughput you’re getting on that kinda setup?
(I have a secondary problem of being “locked into” Claude Code by it being good enough for me, I’d probably need to investigate the other harnesses… my impression is other harnesses are a bit more aggressively OK with nuking your setup from orbit)
The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth.
I think these numbers may improve because the model is pretty new and optimizations aren’t done.
The only reason to run locally is privacy.
Renting tokens from open model providers is cheaper but it incurs the same issues: unexpected changes in model quality, inconsistent speeds, service outages.
Things get cheaper at scale but that's where the provider's margins come in!
I do think there's also an interesting idea: you buy a box like this and run it at a fixed-ish cost (well, electricity). Your demand goes up but your supply is fixed... and that back pressure means that you still have good cost control.
With cloud providers it's a _biiiiit_ too easy to just increase spend.
Sometimes it's OK for things to just be slow.
But the basic single-NN frontier capability has been pretty stationary since Opus 4.8. Kimi K3 is almost as good as that with open weights, which has the frontier labs terrified.
The only big thing on the horizon is if we can get diffusion models working reliably; that would be a big step forward. Inception's Mercury is AFAICT the leader here. It's stupifyingly fast but has obedience/hallucination problems that the autoregressives solved ~2 years ago. So it's not ready yet but improving.
Also, FFS why is Grok the only model that knows how to do parallel tool calls? Such a useful ability and nobody else trains it in. Or if they do it just doesn't work.
Right now the barrier is data and compute.
Quality data can be created synthetically at an exponential rate as models improve. Humans are actively feeding them with private IP.
Compute advancements will begin to skyrocket as we unlock photonic computing and materials science advancements and scale up chip fabs. This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc.
It's a big self-accelerating feedback loop. There is no plateau.
No it can't? Every time the labs try this we see model collapse, e.g. shoving goblins into every conversation.
And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.
The latest studies demonstrate model collapse is not a given and synthetic data can be used just fine. The latest models are proof of that, they're all trained on large swathes of synthetic data. It can't be used as the -only- data source of course, but that's not how it is being used. This is an obvious conclusion, too, because there's no difference between synthetic data and the data people can create, the difference is whether that data is revealing new information about the thing the model is trying to learn. If the synthetic data is just teaching the model the same thing over and over again it results in overfitting, so it needs to be done intelligently.
For example, if I have an example of a puzzle, I can generalize that example and create thousands of synthetic data examples, with different rotations/perspectives, rather than having to find the data naturally. It's not that the models are just generating data out of thin air, they're generating the synthetic data on top of real world data. The smarter the models get, the better they are at generating quality synthetic variations and finding valid synthetic variations.
> And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.
It is accelerating how quickly researchers and engineers can do their jobs.
https://news.mit.edu/2026/ai-helps-design-new-materials-that...
This is only the beginning, too... Look ahead a year or two.
Which studies? [edit: I'll assume you mean these two given by @dorolow: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515]
> It can't be used as the -only- data source of course, but that's not how it is being used
Right, so human data creation would also have to scale up exponentially, and that's not gonna happen.
> because there's no difference between synthetic data and the data people can create
I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there.
> It is accelerating how quickly researchers and engineers can do their jobs. > https://news.mit.edu/2026/ai-helps-design-new-materials-that...
That's pretty clearly a hype article, the headline even says "The CrysVCD tool developed at MIT COULD cut the huge amounts of time and money spent". I'm asking for empirical measurements of timelines, not hypotheticals.
> This is only the beginning, too... Look ahead a year or two.
Lol that excuse is getting really old
It doesn't need to. We're not even close to exhausting the useful synthetic data within the human data we have, let alone all of the new data that is being created.
> I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there.
It's not. It's just bytes of information. A machine and a human can write the same bytes (and often do). Like I already said, model collapse happens when you are overfitting on data without useful, fresh training signals. That's the key difference between the data. The data itself isn't in some way "special", some unique configuration of bytes that imbues special powers, it's that the useful information in it has already been exhausted by the model. You can get the same phenomena by having a poor distribution of human training samples as well. I think you're confusing LLM generated data with synthetic data. Synthetic data doesn't need to be created by an LLM, although an LLM can assist in the creation.
Wiki:
> In early model collapse, the model begins losing information about the tails of the distribution – mostly affecting minority data. Later work highlighted that early model collapse is hard to notice, since overall performance may appear to improve, while the model loses performance on minority data.[11] In late model collapse, the model loses a significant proportion of its performance, confusing concepts and losing most of its variance.[10][12][13]
As models retrain on outputs sampled disproportionately from the higher-probability center of the distribution, rare words and uncommon syntactic constructions are among the first features to disappear.[25] Statistical analysis of recursive next-token prediction training has shown that, when language models are trained recursively on synthetic data, the learned conditional distributions concentrate probability mass on a small subset of highly predictable continuations (a phenomenon characterized as "total collapse")
> That's pretty clearly a hype article
It was just the first article I saw on a quick google search, there are thousands of these stories. It's easy to dismiss anything that doesn't align with your worldview as hype, but you're the one lacking evidence now.
> I'm asking for empirical measurements of timelines, not hypotheticals.
Go and find it then? You haven't bothered looking.
> Lol that excuse is getting really old
You're doing the same thing people have been doing for years, comparing this very second in time and failing to extrapolate. HackerNews was full of developers who said that AI would never be useful for programming, it can't do x, y, z. Now these same people don't write code by hand anymore and haven't looked at their codebases in months.
You had people in mathematics saying the same thing, now you have Terrence Tao posting articles about how AI is stealing their job.
You had artists, designers and photographers saying the same thing, now they can't tell the difference between something human created or AI created.
Edit: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515
There are plenty of research papers on synthetic data that show its value, do a search on arxiv for "synthetic data". There are plenty of open-source post-training pipelines that incorporate synthetic data.
As for the claim about accelerating the progress of hardware or materials science, I've seen quite a number of news articles from teams at universities using AI in their work with high quality outcomes, and they're becoming more frequent.
https://openai.com/index/jalapeno-first-results/
> We used AI to design the chip, and designed the chip so AI could program it AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule.
https://www.anl.gov/article/scientists-deploy-ai-agents-to-a...
> An AI-driven system automates a powerful simulation method used to discover new materials. The system can potentially reduce discovery time from months or years to just days.
Those are pretty significant barriers seeing as we're closed to/have exhausted all the data on the internet and most of those compute bottlenecks are a castle of sand of dodgy finance deals that are getting blocked by community action.
You say "synthetic data" but that's still vaporware right now in terms of being useful for model training. The good synthetic data uses are still grounded in real data and it's a coin flip on if it works well or not.
Depends on defnition of "plateaued" and "ceiling". I am not impressed with 2026 consumer models at all.
> This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc.
Yet it does accelerate - so is does Twitter. But does it to any substantial degree, esp. in AI theory? All the modern LLMs are the same old tired 2017 paper.