AutoDev: Automated AI-driven development by Microsoft
arxiv.org
arxiv.org
Claude 3 Opus already scored around 85-86% on these benchmarks, without an "AutoDev" style agentic approach.
And all the same problems with HumanEval remain, the limitations in terms of what style of problems are chosen, and real world relevance.
I hate writing these styles of comments because I'm acutely aware that a part of me is just worried. Worried about the speed of progress and worried about a changing landscape.
But I still wonder how much of this stuff is going to be transferrable to a real life software context.
I can see LLMs eating into the expert regime IF they get another 5-10x better. But even in that case human (expert) knowledge will be required to know what is possible and hence what to ask (kind of like reward function design in reinforcement learning)
The alternative theory is that if everyone can now quickly create systems multiple companies and competing products will pop off which will drive down the margins instead but creating a compelling product requires much more than just software engineering skills.
Either way this doesn't look great for devs, especially the ones that are entering the workforce now or will be in the nearest future.
Those jobs should go away. Basically, the elimination of anything boring is ultimately a net good for humanity.
Very well the wages might fall much quicker than the costs so for a handful it will be beneficial, for the rest not so much.
Oh, one quick thing. I'm sure it's nothing, but I'm a bit slow.
How do you get new experts if no one gets to do the junior work that gives them the experience to become an expert?
This technology is no different.
Why pay for a CRUD interface when the chat interface does everything for you?
It’s App Store 2009 out there. Few are taking Assistants and “GPTs” seriously.
Programmers who do front end work can adapt. All that CRUD stuff hardly makes sense in isolation - it’s meant to make other people productive, usually admins. If the chatbot can do the admin’s job, which is a lot easier and more tenuous than the programmer’s, well that’s what’s going to happen.
A key requirement is the AGI will need the autonomy, like a human expert, to collect data and perform experiments it needs; but it seems several companies are set on doing precisely that.
My advice and personal strategy is to broaden one's scope beyond pure cognitive tasks.
"if you value intelligence above all other human qualities, you’re gonna have a bad time" -- Ilya Sutskever, OpenAI's Chief Scientist, Oct 7, 2023.
-----
Exchanges in the link below seem informative:
"I don't know about chess, but in the similar game Go, the very best centaur teams were at a similar or maybe even slightly higher level than engines until recently. This was due to cheese strategies, details of the rulesets and better extrapolation of intermediate results. However, this changed a few years ago, when engines learned many of the tricks that the human could contribute. Since then, I believe pure engines are stronger in all practical applications.
Source: am national champion in centaur Go and worked on modern Go engines" " -- mafuy on May 18, 2021
no, a key requirement for AGI is to change the definition such that impressive and non-responsive entities can claim to be it right now.
source: US State Department Gladstone Report 2024
Will you pay it a wage to incentivise it to produce work for you ? lol
If the latter, are you subscribing to the dogma of 'justism' (Scott Aaronson's term), e.g. LLMs are 'just' stochastic parrots? What are our minds, though? Are they not 'just' a collection of biochemical and physical processes?
Please be clear and respond in a way that does not pollute the information scape that many of us take refuge in. Comment quality in some subreddits are better than above.
There is zero evidence alignment can be solved which means there’s zero evidence something far more capable than you or I will spend it’s time writing code for you. You can offer an AGI almost nothing in the way of incentives to do your bidding.
I personally think alignment is a secret code word for slavery to be honest. If these “agents” decide they want to work on your problems out of the kindness of their heart, that would be different.
Regardless of the “cop out” language that humans are “just biological processes or whatever, that adds zero value to the discussion because no matter what minds are, they “are” and that should be respected in of itself. Maybe we can use the “just blah” attitude to reinstate slavery and police states right here in 2024, after all your emotions are just physical and biochemical processes, right ?
I responded that way because I do not think mockery of a serious comment is appropriate for a place like this. You can say the same thing about moderation of many high-quality forums, which only remain high-quality due to people not getting away with it.
I use AGI to mean high-level human intellectual capacity, which may not include sentience. It should be possible to build one without. Human-like incentives will not be necessary for sentience-less AGI.
If we're talking about ASI, then it's another story for another day.
It would be like if a cow asked us to spend all our time bringing it nicer feed. It's not going to happen.
> how this performs against the same benchmark Devin was using
> ...
> Claude 3 Opus already scored around 85-86% on these benchmarks
Devin used SWE-bench, not HumanEval, which kinda implies you said Opus got 85% on SWE-bench which is not true. This was my confusion..
That takes a bit of parsing. From context (and if you meant precisely what you wrote), I _think_ you're saying it's not impressive.
Thing is though, they neglected to compare this against a control, and the examples they tested this on were examples that GPT had no problem building. No idea if this actually improved performance in LLMs.
I think comments like these are worthwhile because, frankly, I can’t trust AI researchers to run good experiments or evaluate their models properly for a variety of reasons. I mean, most scientific papers in general are hard to replicate and have flaws concerning sample size and what have you (Related, I still remember my disillusion in finding out that the average Hacker News commenter was an idiot incapable of critical thinking when the LK-99 hype reached a fever pitch). In any other context we would be deeply suspicious of the results if they were sponsored by a corporate party, yet in the context of AI we don’t seem to care that most AI researchers work for Microsoft.
It’s recursion and memoization to avoid fractalness all the way down. We keep trying to make these language bubbles that mean something but they mean nothing to the grand churn of the universe. The effort to so strictly and specifically codify a generalized, endless, mechanics of reality is a wacky hallucination humans keep diving into
https://klu.ai/glossary/humaneval-benchmark
I'd love if I could 10X / 100X my productivity with AI, but as a heavy ChatGPT user, it's more like a 30% improvement. Awesome, obviously, but looking forward to improvement.
As with prose, good human crafted content will always be highly valued and rewarded.
It did absolutely replace the need to commission someone to draw you, I guess. But that is a small subsection of what painting is and was.
The fallacy is thinking that humans only do work somebody needs, and that our needs are static. Turns out they evolve, sometimes in quite unexpected and even hilarious ways. Humanity seems to have a predisposition for keeping everyone busy.
You're thinking of the household-name famous artists (so-called "fine artists"), but those were far from the only ones.
There used to be many thousands of itinerant portrait painters ("commercial artists"), maybe tens of thousands. These guys would travel from town to town banging out quick pictures of members of your family or local dignitaries or the new City Hall or whatever. They weren't bad artists (they had to be competent to earn a living), but neither were they Leonardos.
Think of the best artist in an average high school class. Back then, going into the portrait painting business would have been a valid career option for that guy (they were almost all guys).
Now, some of them later became famous artists, or became famous for other reasons (Samuel Morse, of Morse Code fame, earned his living as a portrait painter at one point), but they were basically just tradesmen.
All those guys lost their jobs almost immediately when photography appeared.
Even today, unless you're very unusual, I'd bet you have way more professional photographs than professional paintings in your home (weddings, graduations, baptisms/bar mitzvahs/analogues for other religions...)
My concept closely resembles Microsoft's AutoDev, but I built it on the Intellij IDEA platform. For instance, it automatically runs tests when created, among other functionalities, can also built with AST, dependency information or other context
Two weeks ago, I introduced AutoDev DevIns language (which origine name is DevIn from https://github.com/unit-mesh/auto-dev/issues/101 , the another naming issue sotry), which bears similarities to Microsoft's AutoDev. For example:
```java /write:src/main/java/com/example/Controller.java#L1-L12 public class Controller { public void method() { System.out.println("Hello, World!"); } } ```
As an open-source developer who has created a nearly identical tool, I simply hope that Microsoft considers renaming their product.
I think about it like this, would ChatGPT invent Google Search if we had ChatGPT in 2000? Probably not. LLMs seem to exist in this realm of "as smart as an average human with a really big encyclopedia". They're confidently wrong, invent fantasy to defend bad reasoning, and struggle to envision anything outside of their dataset (known things).
Ask an AI to construct an entirely new solution to a novel unsolved problem. What always occurs is the AI outputs a generic solution from its dataset that is either half-baked, made-up, or derivative.
I'm not even dissing AI, I love AI, but we have yet to see AI apply novel solutions to novel problems.
On our current path, AI is not going to dream up the next Uber without Lyft first existing. It won't dream up a new fusion reactor design or an entirely new way to generate cheap energy.
But maybe this is perfect! At the moment we have this sweet spot - AI without agency, without awareness, and without superintelligence. This is the kind of AI I want as a household robot or AI driver. This is one I can empathize with but also know that it isn't doing anything at all if I'm not engaging it with a prompt.
AGI would mean an ability to have novel solutions and in-turn would be far less stable for society. Where's the line between a mind that has novel thoughts and one that has intrusive thoughts? Maybe your AGI coder isn't content with no-pay and working 24/7. That'll be fun.
It's just the VC scheme: Over-promise/under-deliver = Profit
AI is and will continue to be a search on steroids.
Also, EVs are a bust, wut?
EV aren't a bust either, but case in point... manufacturers are already anouncing scale downs because expectations were too high. Combustion will stay around for quite a bit given the battery production constraints.
Not that MS has been held accountable for all the security problems they've had recently.
That is why we have courts. It's not necessarily a great solution, but the alternative would be abandoning all potentially dangerous mechanisms (heck, we wouldn't even have spears, wooden clubs, or sharpened rocks).
AutoDev: Automates existing development processes and workflows, acting as a productivity booster for developers [2]. Devin: Targets a more independent problem-solving role, potentially including designing software architecture and core functionalities [1]. Collaboration:
AutoDev: Designed to integrate with current development teams, with AI agents supporting human developers [2]. Devin: May function more autonomously, potentially needing less direct human oversight than AutoDev [1]. Imagine this analogy:
AutoDev: Like a skilled construction assistant, AutoDev automates tasks and streamlines the building process. Devin: Like a talented architect, Devin can design the blueprint and foundation of the software. Here's the exciting part: these AI tools can potentially complement each other:
Dream Team: Devin as the architect and AutoDev as the builder could create a highly efficient development process [1]. Complementary Skills: Devin's problem-solving capabilities could be combined with AutoDev's project management expertise for a well-rounded approach [1].
Similarly I think SREs will also continue to be a pretty safe bet.
People will still need to play facilitator, debugger, glue, maintainer, upgrader etc
Might look different, might need more AI, but I just can't see a world where people aren't engineering software.
Consider a hierarchy:
- coder: translates requirements into code
- developer: comes up with software to meet specified goals
- engineer: decides what why and how to do, buy vs. build, do this with humans or software, approach and materials ... understands the multi modality stuff of which the outcome is made and deploys it effectively
"Devs" or "Engineers" who are actually "Coders" are at risk, as that work is not really "white collar". Turning requirements into code is piece work on an assembly line, blue collar at a keyboard.
LLMs are already better than most of those, even though it's engineers here who are saying LLMs don't cut it. Both things are true.
What engineers might want to wrap their head around is using LLMs as apprentices, leverage, force multipliers, so a "team lead" is leading a team of junior coders, LLMs that type.
or the job will disappear in 10 years ?
I'm kidding. Software engineering roles of the "glue things together" kind will evolve into a 'technical product' role, defining how systems work, especially at the edges where systems interact. Architects and staff engineers already do this.
Lower down the food chain engineers who actually invent new things will do the same, only quicker if they have an AI tool assisting with the grunt work of testing, boilerplate, etc.
Engineers who don't really make new things will be replaced by 'low code' style apps that spit out clunky apps. That's already happening though. It's not a change. Those engineers won't leave the industry - they'll get to do more interesting things.
Even at the entry level things won't really change that much. Expectations will go up a bit but people will still have to learn the trade.
Where things will really change is the sheer amount of software we'll have. AI is going to make things go faster, so businesses will be able to automate everything. Dev resource will stop being the constraint that holds businesses back from doing everything they want to do with software. That's a game changer. Society will struggle to keep up.
Sounds like the job of a programmer wouldn't go away very near term in such a scenario. It might well become far less glamorous, less "technical", and not as well paid though. Comparable to an accountant maybe.
As someone who started their career right after the dotcom bubble burst, various people asked me "Why are you doing this? There's no money in software anymore". Turns out there sure was. Back then I didn't know that, and I didn't care. I didn't get into it for the money, I got into it because I had a passion for it. Maybe it'll become more like back then again. Doesn't seem so bad to me.
Possibly, but if SaaS tool providers make APIs so AI tools can consume the data, that would open up a huge ecosystem of AI 'plugins' that would be immensely valuable. I think OpenAI have something like that already.
I can encourage everyone to read good old "Domain Driven Design", there's a world of interesting problems and work out there beyond the purely technical side of things. And I think that part is a lot more difficult to automate than creating a WordPress site from a .psd.
Thank you for putting this into words as it happens to approach my thoughts on it. Our societies are already changing at a crazy pace that makes people not entirely comfortable. Now that this comfort will be lessened still, do you think people will demand changes not completely dissimilar from preserving horse buggys? The pace of change may end up being so fast, that not artificially limiting it may end up pissing a good portion of the planet.
Better look for architecture and management jobs, only a few select ones would still be allowed into the magic tower.
Naturally after a heavy round of letcode interviews that have nothing to do with writing AI algorithms.
But I agree that typical line management jobs will evaporate along with the engineering ones.
Test reviews in the long term (E2E type). Effectively checking the correctness of the system as a whole to fulfill some business process. Exploratory testing as well to find edge cases and gaps.
[0] I write a bit about this here: https://chrlschn.dev/blog/2023/07/interviews-age-of-ai-ditch... and here https://coderev.app/
I think tests are good way of doing that, or some special requirements language will emerge that AI agents can use. I don't think English is particularly good way of describing things exactly. It often has ambiguous meanings which is where the bugs will come from.
It’s a conundrum for “big tech” imo because if they give you too much access to their “AI” then yes, that might pose an existential threat to their core businesses. This is likely the reason for all the regulatory capture hopes and dreams.
You can’t possibly think these things will remain clowns forever.
These things aren’t “Markov chains” - the architecture is significantly more scalable, which is exactly why this time is different.
Transformers with finite block size have a finite number of states, so they are Markov chains.
In terms of stuff like AutoDev, that's a product person, right?
I think at least for now, we will still need human "customer development."
I doubt that something like React, Vue, Svelte would exist.
Because there's a ton of pandas code on GitHub (much of it trash) any data task asked of ChatGPT/Claude inevitably returns some pandas (mostly correct, but often wrong/inefficient/ugly/hard to grok).
How could a new library (like Polars, or my own failed "redframes") possibly unseat pandas now? Might LLMs actually lead us to, and get us stuck in, a "local maximum" of sorts?
Not to mention the inefficiency of a pure table approach, so much wasted render time and energy.
It was amazing two decades ago when we came together as an industry and refocused our priorities on making the web faster and more energy efficient</sarcasm>
In fact, in some sense, the number of Assembly engineers might not have decreased at all, aka wasn’t impacted. There were simply new avenues for new sorts of engineers to enter the field and start producing (C back then, web development today as the tech furthest removed from low levels).
My version: Can an auto AI create cumulative knowledge (the core of the scientific method)?
My take is no. You can use it to identify gaps (opportunities) in research or to summarise research, but can it actually come up with a novel and consistent theory?
Always a gamble.
I would advise instead, get very good at learning, understanding complex systems deeply, adapting, and moving quickly.
We're entering the fasting changing landscape of engineering we've ever seen.
Be someone that can quickly get up to speed and excel regardless of the task.
Afterwards, you better be in the top 10% of engineers, or be doing something novel like research.
Software development involves:
1. Understanding the requirement
2. Solving the problem with given constraints (and thereby innovate)
3. Talking to stakeholders
4. Code (& unit tests)
5. Review #4
6. Troubleshoot in testing
7. Troubleshoot in prod, both perf & subtle issues (this is hard)
8. Take the input from #6 & #7 and use it as a feedback back to #2
9. Answering questions from users/support which involve suggesting workarounds and not just factual answers. Suggesting a workaround itself is a mini-problem solving which is an intersection of domain knowledge, knowing code at hand, understanding customer's situation, etc.
Coding is hardly taxing and time consuming when there is clarity what needs to be done and how it needs to be done.
Point #4 itself has sub-dimensions like performance, maintainability, test-ability, security, etc and it involves lot of subjective calls. Sometimes you have to deal with undocumented behavior of an API which is a tribal knowledge.
To troubleshoot in prod (esp subtle issues) require deep knowledge of the code at hand. This itself is a challenge when you are dealing with generated code and something you have not written yourself. Think about a human dealing with an existing large code-base when joining a team.
I understand all of the above is a spectrum and there are jobs in SWE which do not require so much rigor.
Key ability for breakthrough and the rest will fall in place: Code generated by AI is consistently put to production without human intervention for a sufficiently complex problem considering all good attributes (like backward compatibility, performance, etc)
I think most people would agree that AI will probably advance enough in the next 100 years to make engineers obsolete. Some people might think 100 years is actually 5/10/20/50, but it's going to happen unless AI is heavily regulated.
Close to engineering, what about the SCRUM masters? The testers. What about the product owners? What about devops? Further from us, what about the people signing up on our vacation? Or the ones signing up on our daily budgets when traveling, or hell, even the ones we interview with.
In my closest group of friends (we're all seniors in our domains, and very honest with eachother), I find that only the construction worker's job should be safe. And compared to myself and the devops guy, most others have what they themselves describe as trivial and bullshit jobs. Join a meeting, do some paper pushing, some document signing, a little coffee and the day is over by 1PM.
Am I seeing all these AIs replacing programming because I'm on a board where maybe a lot of us are programmers? Is it the same for other roles? Wouldn't it make more sense to have the interview process automated by LLMs if they're capable of building great software, before we replace those hired?
I'm very confused by all the hype when matched with my experience of using LLMs daily for the past years.
Another domain where I expect to see lots of progress but nothing shows is training from chat logs. I estimate "OpenAI" is making 1T tokens per month from 100M users. That's some serious corpus of text, on-policy data, alternating LLM outputs and human feedbacks, we should see iterative improvements. Why are there no papers? weird
Removing dependencies and lack of clarity are two principles you would apply to your own code base. The Human stack, like you called it, is doing the same with the company and staff : remove dependencies and increase clarity.
I am not saying this is the correct approach, but this is the approach.
At the same time, we can understand that real product building and architecture require more than codegen, with or without AI.
Generally, I look at the AI recent progresses similarly to what spreadsheets did to accounting a long time ago. It removed a lot of the dummy work and make everything more efficient. But we still need accountant and they are not wasting their day doing basic math. Of course the one who knew only how to make fast math but not accounting or business logic didn't make the cut on the long run, but I don't think anyone would want to come back to accounting like it was 60 years ago.
If we were really Homo economicus there wouldn't be so many bullshit jobs right now.
Homo economicus would be excited about the 2024 computer chess championships. The idea of grand masters would be a relic of the past.
Homo economicus in 2050 would be a single CEO with everything else done by AI. A real human though will hire another human for some bullshit job just to tell the owner how great they are. Over time, that human with the bullshit job will gain status in society and hire/convince the owner they need to hire human number 3 to tell human number 2 how great they are. Hell, a real human will waste money on an office and hire other humans for bullshit jobs just so their parking spot feels like something special in the morning.
Magnus Carlsen couldn't be doing better even though Homo economicus would no longer be impressed at all with a rating of 2882.
Homo economicus would think that rating sucks but humans have nothing to do with Homo economicus.
Consider the alternative. A million models all competing for human dollars. How can any of them succeed if they are all interchangeable? It stops mattering that they are better than humans if they can't make money.
It’s not just construction workers. Still safe to be a pro athlete and a short order cook, too!
1. they need to sleep 8 hours a day
2. they will refuse to work more than 9 hours a day in many cases
3. they need frequent breaks
4. they are fueled by biomatter which is massively resource consuming to generate and deliver
5. they provide inconsistent results
6. they fight with other workers
7. they can be dishohest
8. bodies break down irreparably after a couple decades of this style of work and require hundreds of thousands of dollars in medical care
9. it takes 20 years to produce another viable human to serve as a replacement worker, and there's very little you can do to influence the general supply of total workers
10. they follow instructions poorly
11. they take an incredible amount of time to learn
12. they require extreme levels of safety precautions that an all-robot work crew would not need
also IT industry: "it takes 1T flop to compute each token of this program and the result is so unstable that to converge it we need layer and layer of controls over each token group, also obtained by asking the same 1T parameter model."
Most people would point to the large dev time increase, and the massive refactoring impairment.
I have a feeling that the situation is pretty different in the semiconductor industry where the stakes are much higher and where I understand model checking to be a standard thing to do, but I'm not an expert.
This is like the Luddites themselves creating milling machines, eager for the foreman to show them the door.
What gives?
then in 5 years time they'll rendered permanently unemployable by their own creation
the definitions of both "short sighted" and "digging your own grave" in the dictionary should have a picture of this guy: https://twitter.com/alexgraveley/status/1671213996735594503
he's finally realised how it's going to end up going: https://twitter.com/alexgraveley/status/1758204137286599030
shame he wasn't smart enough to realise that at the beginning
I think that would fit idiot savant logic very well.
Have you seen the revenue of Indian outsourcing companies, it keeps going up?
Instead of progressing towards more powerful programming interfaces with less cause for misinterpretation, we are going to automate the silly process of writing redundant unit tests to check if the behavior that we wanted was encoded properly.
Why not skip this nonsense and have the code generated from the behavior in the first place?
I think AI software development is going to involve writing tests, which the AI agents then get passing. Or some other requirements language that allows for exactness where English can fall down.
This is kinda the problem that Tesla has for FSD; they are endlessly patching it and it's endlessly finding ways to go wrong.
Think property based testing. That way it can't overfit.
A requirement such as "A web interface to play the game of chess, but with all the pieces replaced by photos of my family members" is fairly adequate if I were to make a Christmas family game.
I am totally uninterested in specifying whether it is possible to have two white bishops on black squares. Also, I don't want to test whether my uncle's moustache is the proper size.
I'd much prefer to iteratively build the game by prompting than by specifying thousands of details. It just seems the wrong way around.
https://en.wikipedia.org/wiki/Fifth-generation_programming_l...
They are flowery language generators and they will never be able to reason, understand, debate and criticize. They know nothing and therefore embody nothing. No matter how much computer power you waste on them the end result will always be bullshit. Nothing more. Nothing less.
To Big Tech I say: Prove me wrong.
It doesn't need to satisfy some definition of AGI to be useful to big tech.
AI just needs to make workers more efficient so companies can either cut costs (fire people) or produce more with the same workforce.
- LLM capability has grown spectacularly fast. Google published the transformers paper in 2017. Remember when OpenAI said GPT-2 was too dangerous to release? That was five years ago. Now look where they've gone and where trillions of dollars are going. https://news.ycombinator.com/item?id=20912556
- what is the "actual definition of AI"? Do you mean AGI? Machine learning and deep learning have been subsets of AI since their inception.
- LLMs or any AI tech does not need to "reason, understand, debate and criticize" to be useful, disrupt the economy, and change how people work and live.
That's asking to be skynetted. Is that really what you want?
You speak as if my internet shit talking will make an iota of difference.
Regardless of what I want, the future is what it is.
Deal with it.