Streamer discovers major vulnerability in Cognition's Devin live on air
twitter.com
twitter.com
This situation is both embarrassing and difficult to excuse given the severity. It’s especially puzzling since they’ve been testing this for about a year. Something of this nature should have been caught and feels amateurish. That said, I’ll give them credit for reacting quickly to address the issue and can understand that, given startup pressures, mistakes can happen and bigger organizations with more experience have made similar mistake. Not to justify, but I feel providing context is fair.
But to provide more context, let us remember their months of dubious claims about their "models" supposed superiority. They’ve repeatedly asserted that Devin outperforms all other LLM implementation for coding tasks purely in quality and ability, actively identifies bugs, and independently fixes them. In the public access announcement thread [0], several early access users highlighted this as a key difference between Devin and tools like Cursor/VSCode plugins for LLMs. They claimed Devin could be "told" to independently "search for issues" and then create pull requests, effectively functioning like a "virtual Junior Dev" assigned to tasks.
Clearly, that’s not the case if they themselves don't seem to rely on Devin?
The crux of the problem lies beyond their past credibility issues or the tragic "Leetcode genius dropout" theatrics (something that should have gone out of style long before SBF started playing League of Legends during investor calls). My primary criticism of Devin, purely as a product, is the misleading presentation. They try so hard to position themselves as superior in raw coding performance, not just implementation of other models with their toolset, yet the code examples and pull requests don’t support that claim at all.
If, as I suspect, Devin is just a set of tools to give OpenAIs API more access to ones codebase, that is totally fine, less then promised, but still something useful and can add significant value, as seen with tools like Cursor. While, whether LLMs at this stage should operate with the level of autonomy Devin seems to target is debatable [1], I can see the concept behind Devin having a future. If Devins team simply stated their aim was to create a solid foundation for LLMs to interact with code beyond traditional IDEs, priced it appropriately and avoided any notion that they aren't reselling existing models, I feel they'd do a lot better in the long run and I'd be a customer.
There is no shame in relying on another companies API something said company doesn't, as seen with Curor. Devin should position themselves in the same manner as Cursor. Acting like the model part of the equation is what makes Devin superior, that is what opens them to criticism.
Well, that and their previous fake demos, lack of benchmarks or whitepapers, charging more than all competitors combined, having rather embarrassing security issues and how they approached the media up to this point.
Alternatively, maybe I am wrong and those previous assertions of being superior purely in coding performance are credible. In that case, they could silence any doubts by proving their system isn’t merely an API wrapper by releasing transparent performance metrics [2].
[0] https://news.ycombinator.com/item?id=42378994
[1] I feel LLMs in coding at this stage are roughly at the same point driver assistance in modern vehicles tends to be, just good enough that users could be lead into a false sense of security, stop paying attention and end up with issues. A solution like Devin or some "FSD" solution could be reliable for multiple hours, then suddenly produce faulty code or disengage, making it harder to react straight away. In both cases, I feel that UX at this stage should lean into ensuring consistent verification, something that the current crop of VSCode plugins or Cursor do a decent job of by consistently highlighting diffs.
[2] Example how this is done: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bb...