OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
arxiv.org
arxiv.org
I gave it one example and then asked it to do the work for the other files.
It was able to do about half the files correctly. But it ended up taking an hour, costing >$50 in OpenAI credits, and took me longer to debug, fix, and verify the work than it would have to do the work manually.
My take: good glimpse of the future after a few more Moore’s Law doublings and model improvement cycles make it 10x better, 10x faster, and 10x cheaper. But probably not yet worth trying to use for real work vs playing with it for curiosity, learning, and understanding.
Edit: writing the tests in this PR given the code + one test as an example was the task: https://github.com/roboflow/inference/pull/533
This commit was the manual example: https://github.com/roboflow/inference/pull/533/commits/93165...
This commit adds the partially OpenDevin written ones: https://github.com/roboflow/inference/pull/533/commits/65f51...
I have found it immensely useful for a handful of one-off tasks, but it's not yet a mission-critical part of my workflow (the way e.g. Copilot is).
Core model improvements (better, faster, cheaper) will definitely be a tailwind for us. But there are also many things we can do in the abstraction layer _above_ the LLM to drive these things forward. And there's also a lot we can do from a UX perspective (e.g. IDE integrations, better human-in-the-loop experiences, etc)
So even if models never get better (doubtful!) I'd continue to watch this space--it's getting better every day.
Aider wrote 61% of the new code in its last release. It’s been averaging about 50% since the new Sonnet came out.
Data and graphs about aider’s contribution to its own code base:
For a project like yours I guess you should be given free credits. I hope that happens, but so far nobody has even given Karpathy a good standalone mic.
I can’t get anything useful out of these AI tools for my tasks and I’d really like to see what someone who can does.
I’d like to know if it’s me or my tasks that aren’t working for the llm.
It's also very sensitive, unsurprisingly, to development documentation that is moving quickly, e.g., most AI APIs right now. A lot of manual intervention is still required here because of out-of-date references to imports, etc.
It only knows statistically what is the most likely sequence of words to match your query.
For rarer datasets e.g. I had Claude/OpenAI help out with an IntelliJ plugin it would continually invent methods for classes that never existed. And could never articulate why.
You can auto- lint and test code before you set eyes on it, then re-run the prompt with either more context or an altered prompt. With local models there are options like steering vectors, fine-tuning, and constrained decoding as well.
There's also evidence that multiple models of different lineages, when their outputs are rated and you take the best one at each input step, can surpass the performance of better models. So if one model knows something the others don't you can automatically fail over to the one that can actually handle the problem, and typically once the knowledge is in the chat the other models will pick it up.
Not saying we have the solution to your specific problem in any readily available software, but that there are approaches specific to your problem that go beyond current methods.
It’s the new pop psych.
I asked Claude and OpenAI models over 30x times to generate code. Both failed every time.
Which is the elephant in the room.
There is no roadmap for any of these to happen and a strong possibility that we will start to see diminishing returns with the current LLM implementation and available datasets. At which point all of the hype and money will come out of the industry. Which in turn will cause a lull in research until the next big breakthrough and the cycle repeats.
b) Synthetic data sets have been shown to not be a substitute.
c) I have no idea why you are linking Moore's Law with AI. Especially when it has never applied to GPUs and we are in a situation where we have a single vendor not subject to normal competition.
While Moore's Law probably doesn't strictly apply to GPUs, it's not far off. See [1] where they find "We find that FLOP/s per dollar for ML GPUs double every 2.07 years (95% CI: 1.54 to 3.13 years) compared to 2.46 years for all GPUs." (Moore's law predicts doubling every 2 years)
https://epochai.org/blog/trends-in-gpu-price-performance#tre...
That incentive doesn’t invalidate research, but AI results are so easy to nudge in any direction that it’s hard to ignore.
So while we're not 10x-ing everything, it's not like there's no significant improvements in many places.
There have been a lot of people making these sorts of claims for years, and they nearly never end up accurately predicting what will actually happen. That's what makes observing what happens exciting.
Maybe there’s some wins to be had on the software side still.
The "Browsing agent" is a bit worrisome. That can reach outside the sandboxed environment. "At each step, the agent prompts the LLM with the task description, browsing action space description, current observation of the browser using accessibility tree, previous actions, and an action prediction example with chain-of-thought reasoning. The expected response from the LLM will contain chain-of-thought reasoning plus the predicted next actions, including the option to finish the task and convey the result to the user."
How much can that do? Is it smart enough to navigate login and signup pages? Can it sign up for a social media account? Buy things on Amazon?
you want something unique but not too unique as to be weird.
I work with like 6 Matts.
Your interestingly different ire would be better-directed at the original project.
https://www.cognition.ai/blog/introducing-devin
Previous discussions on that fwiw include:
dont build a platform for software on something inherently unreliable. if there is one lesson i have learnt, it is that, systems and abstractions are built on interfaces which are reliable and deterministic.
focus on llm usecases where accuracy is not paramount - there are tons of them. ocr, summarization, reporting, recommendations.
software system is like legos. they form a system of dependencies. each component in the chain has interfaces which other components depend on. 99% reliability doesnt cut it for software components.
The word "need" is an extreme overstatement here. The vast majority of software out there is unreliable. If anything, I believe it is AI that can finally bring formally verified software into the industry, because us regular human devs definitely aren't doing that.
but why should the solution involve AI (thats just the latest bandwagon)? formal verification of software has a long history which has nothing to do with AI.
Because AI is able to produce lots of results, covering a wide range of domains, and it can do so cheaply.
Sure, there are so quality issues. But that is the case for most software.
One case, the other dev refused to allow a commit (fine) because some function had known flaws and was should no longer be used for new code (good reason), this fact wasn't documented anywhere (raising flags) so I tried to add a deprecation tag as well as changing the thing, they refused to allow any deprecation tags "because committed code should not generate warnings" (putting the cart before the horse) — and even refused accept that such a warning might be a useful thing for anyone. So, they became a human compiler in the mode of all-warnings-are-errors… but only they knew what the warnings were because they refused to allow them to be entered into code. No sense of irony. And of course, they didn't like it when someone else approved a commit before they could get in and say "no, because ${thing nobody else knew}".
A different case, years after Apple had switched ObjC to use ARC, the other dev was refusing to update despite the semi-automated tool Apple provided to help with the ARC transition. The C++ parts of their codebase were even worse, as they didn't know anything about smart pointers and were using raw pointers, new, delete everywhere — I still don't count myself as a C++ despite having occasionally used it in a few workplaces, and yet I knew about it even then.
And, I'm sure like everyone here has experience of, I've seen a few too many places that rely on manual testing.
I have a suspicion that there's a "best design pattern" and "best architecture" for getting the most out of existing LLMs (and some equivalents for non-software usage of LLMs and also non-LLM AI), but I'm not sure it's worth the trouble to find out what that is rather than just wait for AI models to get better.
It really seems like the tech-bro space hates humans so much that their motivation in working on these products is replacing them to never have to work with a human again.
Sure, but then humanity was denigrated the first time a calculator was used to compute a sum instead of asking John Q Human to do it.
I'd argue that the more we find ways to replace humans with AI, we're more clearly defining what humanity is. Not about denigration or elevation, just truth.
Are you sure we live in the same world? The world where there is Crowdstrike and a new zero day every week?
Software engineering is beautifully chaotic, I like it like that.
So much of the stuff being built on LLMs in general seems fixated on making that illusion more believable.
I prefer to think of agents as _feedback loops_, with an LLM as the engine. An agent takes an action in the world, sees the results, then takes another action. This is what makes them so much more powerful than a raw LLM.
It can't self-correct its own reasoning: https://arxiv.org/abs/2310.01798
An LLM in a loop creates agency much like a car rolling downhill is self driving.
It was a bit inscrutable what it did, but worked no problem. Much like chat gpt interpreter looping on python errors until it has a working solution, including pip installing the right libs, and reading the docs of the lib for usage errors.
N of 1 and a small freestanding task I had done myself already but I was impressed.
The best way to use arxiv.org is to find a paper you want to read from a "real" publication and get the pdf from arxiv.org so you can read it without the publication subscription.
That is not to say arxiv.org is all horseshit though. Plenty of good stuff gets added there; you just need to keep your bullshit radar active when reading. Even some stuff published in Nature or IEEE smells like unwashed feet once you read them, let alone what arxiv.org accepts.
Good citation count and decent writing are often better indicators than a reputable publication.
Aider is more understandable to me, doing small chunks of work, but it won't do a google search to find usage, etc. It depends on you to choose which files to put in context and so on.
I wish aider had a bit more of the self directedness of this, but API calls and token usage would be greatly increased.
Edit: or maybe an agency loop like this steering aider based on a larger goal would be useful?
I still think a tool like aider is where AI is heading, these "agents" are built upon running systems that are 15% error prone and just compound errors with little ability to actually correct them.
To what end anyway? This is massively resource heavy, and the end goal seems to be to build a program that would end your career. Please work on something that will actually make coding easier and safer rather than building tools to run roughshod over civilization.
- fix A please
- hmm, ok A fixed, B broken; fix B please
- hmm, ok B fixed, A now a bit broken, fix A please
- A & B working
But when you check the code, you often see that it wrote code for A that broke B, then it fixed B while leaving the code for A, now basically dead code but not necessarily detectable. Then it wrote code for A, again, after the code of B and the user thinks all is fine as it works. And this happens 1000x / day in normal projects.
I see it everywhere. Good for me (my company troubleshoots and fixes code/systems), but not for the world.