What eval is tracking that? It seems like it's potentially the most imporatnt metric for real-world software engineering and not one-shot vibe prayers.
What eval is tracking that? It seems like it's potentially the most imporatnt metric for real-world software engineering and not one-shot vibe prayers.
https://charlielabs.ai/research/gpt-5
Often, our tasks take 30-45 minutes and can handle massive context threads in Linear or Github without getting tripped up by things like changes in direction part of the way through the thread.
While 10 issues isn't crazy comprehensive, we found it to be directionally very impressive and we'll likely build upon it to better understand performance going forward.
For better accessibility and a safer experience[1] I would recommend not animating the background, or at least making it easily togglable.
[1] https://developer.mozilla.org/en-US/docs/Web/Accessibility/G...
Edited to add: I am, in fact, photosensitive (due to a genetic retinal condition), and for my eyes, your site as it is very easy to read, and the visualizations look great.
Love that you included the judge prompts in your article.
I'm not sold on the efficacy of AI and I share your reservations about having to scrutinise their output, but I see great value in being able to offload a long-running task to someone/something else and only have to check back later. In the meantime, I can be doing something else - like sitting in those planning meetings we all enjoy!
This is exactly right. We've adapted our workflow to kick off a task and then kick off the next one and the next. Then we review the work of each as they come through. It's just CPU pipelining for human workflow.
The process is far from perfect but the throughput is very high. The limiting factor is review. I spend most of my time doing line-by-line review of AI output and asking questions about things I'm unsure of. It's a very different job from the way I historically operated, which involved tight code -> verify loops of manually written code.
For whatever reason Github's Copilot is treated like the redheaded stepchild of coding assistants. Even through there are Anthropic, OpenAI, and Google models to choose from. And there is a "spaces"[0] website feature that may be close to what you are looking for.
I got better results for testing some larger task using that than I did through the IDE version. But have not used it much. Maybe others have more experience with it. Trying to gather all the context and then review the results was taking longer than doing it myself; having the context gathered already or building it up over time is probably where its value is.
For my use cases, this is mostly needing to be really home in on relevant code files, issues, discussions, PRs. I'm hopeful that GPT5 will be a step forward in this regard that isn't fully captured in the benchmark results. It's certainly promising that it can achieve similar results more cheaply than e.g. Opus.
If there's no substantial difference in software development expertise then GPT-5 absolutely blows Opus out of the water due to being almost 10x cheaper.
Because if not, I'd still go with Opus + Claude Code. I'd rather be able to tell my employer, "this will cost you $200/month" than "this might cost you less than $200/month, but we really don't know because it's based on usage"
> Availability and access > GPT‑5 is starting to roll out today to all Plus, Pro, Team, and Free users, with access for Enterprise and Edu coming in one week. Pro, Plus, and Team users can also start coding with GPT‑5 in the Codex CLI (opens in a new window) by signing in with ChatGPT.
But is it really 272k even if the output was say 10k? Cause it does say “max output” in the docs, so I wonder
I found 100k was barely enough for a single project without spillover, so 4x allows for linking more adjacent codebases for large scale analysis.
To get great results, it's still very important to manage context well. It doesn't matter if the model allows a very large context window, you can't just throw in the kitchen sink and expect good results
>"GPT‑5 is the strongest coding model we’ve ever released. It outperforms o3 across coding benchmarks and real-world use cases, and has been fine-tuned to shine in agentic coding products like Cursor, Windsurf, GitHub Copilot, and Codex CLI. GPT‑5 impressed our alpha testers, setting records on many of their private internal evals."
However, at leas for me there is lots of "small enough context" boilerplate that the context can deal with.
Clearly this is not a tool in the sense it's predictable.
I'm using it mostly for C#, WPF and OpenTK. The type system seems to help a lot.
The UI logic it recommends is mostly god awful. But at least for me when it's given a pattern it can apply, it does so pretty well.
But GPT-5 is substantially cheaper[0].
[0] https://simonwillison.net/2025/Aug/7/gpt-5/#pricing-is-aggre...
The power of these models has peaked and simply arn't going to manage the type of awareness being promised.
don’t have long-running tasks, llms or not. break the problem down into small manageable chunks and then assemble it. neither humans nor llms are good at long-running tasks.
If LLMs are going to act as agents, they need to maintain context across these chunks.
That's a wild comparison to make. I can easily work for an hour. Cursor can hardly work for a continuous pomodoro. "Long-running" is not a fixed size.
LLMs multiply errors over time.
Claud always misunderstands how API exported by my service works and every compaction it forgets all over and commits "oh api has changed since last time I've used, let me use different query parameters", my brother Christ nothing has changed, and you are the one who made this API.
Yes, I can also tell any agent to commit more often, but that's again not what I'm saying. I'm saying version control can be integrated way deeper into agent workflow.
Kiro creates tasks from spec document, and you can revise tasks either by prompting or editing tasks.md file.
My thing works a bit differently, but essentially the same.
The long running task, at it's core, is composed of many smaller tasks and you mostly focus on one task at a time per brain part. It's why you cannot read two streams of text simultaneously even if both are in your visual focus field.
I think the plan is not just words, if it was, you could read a book on how to ride a bike.
Because we communicate in language and because code output is also a language we think that the process is also language based, but I think it's not, especially when doing hard stuff.
I know for certain in my case it isn't -- when tracking a hard problem for a junior after 2 hours of pair programming the other week, I had to tell him to commit everything and just let me do some deep thinking/debugging and I solved the problem myself. Sure I explained my process to him in language the best I could, but it's clear it was not language, it was not liniar, I did not think it step by step.
I wish I could explain it, but when figuring out a hard problem, for me it takes some time to take it all in, get used to the moving parts, play with them. I'm sure there are actual neurons/synapses formed then, actual new wires sprawling about in the brain, that's why it takes time. I think the solution is a hardware one, not a software one.
That's why we can sleep on it and get better the next day and that's why we feel the problem. There are actual multiple paralel "threads" of thinking going at the same time in our heads and we can FEEL the solution as almost there.
I think it simply is that hard problems can occur in a combination of code, state, models that simply cannot be solved incrementally and big jumps are necessary.
I'm not saying the problem cannot be solved incrementally, but it's possible that by going in small steps, you either reach the solution or a blocker that requires a big jump.