SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
cognition.com
cognition.com
What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.
I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.
They did not.
I actually started typing the same point that the chances are actually high because of train/eval overlap then realised you answered your own question with that same observation.
It is interesting though!
Perhaps in some way this means we should decide which eval set aligns best with our taste?
Back to the blog post. This is an excellent write up of an excellent technical achievement.
I have a lot of respect for the Cognition/Devin (always "Windsurf" to me) and Cursor teams.
I found it interesting - but justified - that they referred to themselves as a foundation lab rather than a dev tools company.
…that are my own private internal suite on my own code bases where I can judge the output properly
I also measure wall clock time to completion which has been a surprising separator in practice.
But here, both Kimi 2.7 and its derivative SWE-1.7 are ahead of GLM 5.2. This tells me the benchmarks they use are cherry-picked.
Which benchmarks would you have chosen instead, and why?
The ideal way to run these benchmarks would be to give a 3rd party the model to run in an isolated environment so the prompts don't make their way back to the AI engineers.
That seems doable for open weight models, but not for private models.
Set up a script that launches the harness for each model, prompts them to implement one of the tasks, let it churn until either tests pass or it hits some budget limit.
Then, most importantly, read the transcript and output and judge subjectively - I don't think this actually can be narrowed down to a score, although tokens burned to fix, whether it actually got the tests green etc are all good signals.
(I've done this, but so far only on a codebase that was too complicated with models that were too weak because I didn't want to spend more than a few dollars - results were inconclusive, planning on iterating on my personal benchmark in future)
RL environments building on top of each other will get these models there
needs people doing software development lifecycles to figure it out and implement
I remember them saying a few years ago that, they didn't think it was worth specializing models for code, because their general purpose models kept beating them. I guess they changed their mind? Since they did start making codex models again.
Today my "coding" sessions often enough begin with real life problems, where I discuss domain or inter-domain things, ranging from business, economics, psychology, etc. Being able to do all of that with one model is something I am willing to pay a premium for.
Of course not having to pay the premium, because the routing is smart or whatever, would be great. I just don't want to have to think about it.
intuition is that your sessions consists of 10% of domain related reasoning, and 90% of code plumbing. Those 90% could be moved to cheap and efficient specialized and focused model.
Regardless, it's fairly obvious to me that none of what I do now will require "frontier models" for much longer. Models are getting better more quickly than my problems are getting harder.
most agentic coding app can use powerful model for planning/reasoning then use "budget" model to do ground work
I've had terrible success using budget models to do ground work. The justifications that the budget models will use and document, polluting the rest of the session, are sometimes just insane. Like making code compatible with a bug that was implemented within the same session, not handling errors due to precedence in the code it just implemented, etc. I DO have success using the heavy models with lower effort, and using budget models on relatively changes post ground work. But major planning and initial ground work, I just get absolutely slop if I use a budget model.
If you're doing web stuffs, or GUI, then the budget models seem fine.
Taking a good model like GLM5.2 and just fine tuning it on coding can decrease real world performance due to mechanics like catastrophic forgetting. There is also other interesting behaviors were training on a broad training set can improve coding performance because there is positive transfer.
There is 100% an effort to make solid coding focused models, but it is very hard to do that without including capabilities across a broad set of adjacent tasks.
This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technical jargon...
* Based on the first comment in the link that claims to summarize the video.
Could you expand on this?
Remember when AGI was going to replace all jobs in 6 months? It's always been like that.
I want to work in the AI space on actual AI research, at any part of the stack. Even if I'm developing training infra - as long as people are advancing knowledge of what intelligence could be.
But it seems like either it's big labs or grifters, that's it, and even the big labs, at least publicly, seem very grifty at times. Not like I have the technical chops probably, but still.
I'm an OpenCode user, but I'll fall back to Claude Code if I want to use Opus end to end for something, given my company has a subscription. But I'm not using yet another tool and subscription for a model that isn't even winning.
Apparently 'free' on the $20/mo Devin plan (presumably within some quota still)
and that is "via Cerebras at 1000 TPS" according to the announcement
I live on Opus 4.8 High and their benchmark scores SWE-1.7 slightly higher ... if at all realistic that sounds like a great deal ... too good to be true?
But the normal speed one seems to be free or with very generous limits.
That company truly subsidized its user base to the extreme before, the $15/mo subscription was the best value on Earth paired with weekly deals reducing credits for premium models. Now it's barely any messages for paid models, completely watered down.
So far, Cursor provides the best value for their subscription, but I have to imagine they're basically lighting money on fire. There's no way their current pricing is sustainable.
I'm currently experimenting with OpenSpec[0] as the "framework" and using different subscriptions for different parts of the spec-driven process: Opus via Claude Code for exploration, Devin SWE for building, and GLM 5.2 via the Z.ai Coding Plan for verification. I don't love having to mix and match harnesses, but in practice it's barely more effort than switching models.
I like Cerabras, but I really wish they would make more of their hosted models generally available.
And yes, with Fable, the chance of that is higher than with SWE/Composer, but in my experience it's not so much higher that the extra time and cost is worth it. But it certainly depends on your goals and what you're building.
Faster iteration means i mentally checkout less and am more involved with the code being created.
My hope is that in the far far future, we can get LLMs so fast that i can work in my IDE like normal and the LLM will just be an extension of autocomplete. I can state a goal, rough out functions, code, etc, and it'll just work around me like a very fast pair programmer / autocomplete.
The chat interface is an intermediate step that frankly i hate. The faster it is the less i wait.
Now for vibe-slop i'm making on the side, yea i don't care about speed. But that's not something i'm employed to do or anything i truly care about. It's a different workflow entirely.
> Faster iteration means i mentally checkout less and am more involved with the code being created.
This is a good point I didn't consider and you're right. More interaction brings you closer to the code.
I still think that this is the opposite of what I personally want. Either I write the code (or a large majority of it), and be fully involved; or be more disconnected but more free to focus on other things. The middle ground removes me from the equation, but also requires me to babysit.
theres a lot of cases where a prof forced their students to put them first even if they had an advisor role, or even credit someone for zero real work because they threatened to block submission and prevent the students from getting their degree.
At least with low level programming languages. They're all very good for webdev stuff.
For hobby projects I've completely switched to DeepSeek v4 pro. I spend less than on a $10 Claude plan and am not subjected to quota limits (when I have time and motivation, the last thing I want is a 5 hour quota running out). And the difference in model performance is fine for those smaller projects, most of which will end up abandoned or in a state of "good enough" anyways
And for utility tasks, those 30b models are also great. I'm a big fan of gemma4
For context, I'm paying under $30/year and get GLM-5.2. An extra $2300/year isn't going to get me much better outcomes.
Put another way, what I get for my under $3/mo is better than what you were getting 3-5 months ago paying $200/mo. So you're paying a lot just to be ahead by a few paltry months.
As a (former) Windsurf user I'm pretty happy with the progress of the Cognition/Devin ecosystem after they took over Windsurf, now known as Devin Desktop.
Time to support it in my agent IDE just like Cursor's...
review the top stash and tell me what's in it (grouped appropriately)
1.6 does this fine nearly instantly.
1.7 tried for 17s before I killed it
It will work until they IPO.
Imagine how far community might have pushed if 2 past versions of 'morally superior' Anthropic and 'completely Open AI' open sourced their models for the community to build on top of them
Should as in "would it be nice?" - yeah. Should as in they have to? No.
> Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so
You can do pretty much anything you want with an MIT license.