I was not trained professionally yet I'm writing production code that's passing code reviews in languages I never used. I will create a prompt, validate it compiles, passes tests, have it explain so I understand it was written as expected and write documentation about the code, write the PR, and I am seen as a competent contributor. I can't pass leet code level 1 yet here I am being invited to speak to developers.
Velocity goes up and cost of features will drop. This is good. I'm seeing at least 10 to 1 output from a year ago based upon integrating these new tools.
I was going through devin's 'pass' diffs from SWE bench.
Every one I ended up tracing to actual issues caused changes that would reduce maintainablity or introduced potential side effects.
I think it may be useful as a suggestion in a red-green-refactor model, but will end up producing hard to maintain and modify code.
Note this one here that introduced circular dependencies, changed a function that only accepted points to one that appears to accept any geometric object but only added lines.
Domain knowledge and writing maintainable code is beyond generative transformers.
https://github.com/CognitionAI/devin-swebench-results/blob/m...
You simply can't get past what Gödel and Rice proved with current technology.
It is like when visual languages were supposed to replace programmers. Code isn't really the issue, the details are.
And to be fair, lots of humans are already at least this bad at writing code. And lots of companies are happy with garbage code so long as it addresses an immediate business requirement.
So Devin wouldn't have to advance much to be competitive in certain simple situations where people don't care about anything that happens more than 2 quarters into the future.
I also agree that producing good code which meets real business needs is a hard problem. In fact, any AI which can truly do the work of a good senior software engineer can probably learn to do a lot of other human jobs as well.
With this quality of changes it won't be long until violations stack up to where further changes will be beyond any algorithms ability to unravel.
While lots of companies do only look out in the short term, human programers are incentivized to protect themselves from pain if they aren't forced into unrealistic delivery times.
At&t wireless being destroyed as a company due to a failed SAP migration that was largely due to fragile code is a good example.
But I guess if the developer jobs that will go away are from companies that want to underperform in the market due to errors and a code base that can't adapt to changing market realities, that may happen.
But I would fire any non intern programmer if they constantly did things like removing deprecation comments and introduced circular dependencies with the majority of their commits.
https://github.com/CognitionAI/devin-swebench-results/blob/m...
PAC learning is powerful but is still probably approximately correct.
Until these tools can avoid the most basic bad practices I don't see any company sticking to them in the long term, but it will probably be a very expensive experiment for many of them.
While RLHF will help improve systems, code correctness is not easy to judge outside of the simplest cases.
Note how on OpenAI's technical report, they admit performance on college level tests is almost exclusively from pre-training. If you look at LSAT as an example, all those questions were probably in the corpus.
But that's the thing, that it seems that everyone here on HN (and elsewhere) finds it easy to judge the flaws of AI-generated code, and they seem relatively consistent. So if we start offering these critiques as RLHF at scale, we should be able to bring the LLM output to the level where further feedback is hard (or at least inconsistent), right?
Not this again. Those theorems tell you nothing about your concerns. The worst case of a problem is not equal to its usual case.
I even wrote a majority of my codebase in Python despite not knowing Python precisely because I would get the best recommendations from LLMs. As a frontend developer, with no experience in backend engineering in the last decade, and no Python experience, building an app where almost every function has gone through an LLM at some point, for almost 8 months — I would be extremely surprised if some of the code it generated landed in production.
Think of this as Facebook page vs. WordPress website vs. A full custom website. The best option is to have a full custom website. Next, is a cheaper option from someone who can put a few lines together. The worst option is a Facebook page that you can create yourself.
But the Facebook page also does the job. And for some businesses, it's fairly enough.
Your coworkers likely aren't doing a very good job at reviewing, but also I don't blame them. The only way to be sure code works is to use it for its intended task. Brains are bad interpreters, and LLMs are extremely good bullshit generators. If the code makes it to prod and works, good. But honestly, if you aren't just pushing DB records around or slinging HTML, I doubt it'll be good enough to get you very far without taking down prod.
Then again, it's always dinosaurs who value their own teachings, above anything else, and try to cling on to it, at any cost, without learning new tools. So, while the industry is going through major changes (2023 saw a 30% decrease in new hires. Among 940 companies surveyed, 40% expect layoffs due to AI), people should adapt rather than ignore the signs.
You still need that domain knowledge of whatever you are writing code for or integrating with, especially is the technology is more niche, or documentation was never made available publicly and scraped by the AI
But when it comes to writing boilerplate code it is great, or when working with very commonly used frameworks (like front end javascript frameworks in my case)
Okay, so you are just kicking the can down the road to the test engineers. Now your org needs to spend more resources on test engineering to really make sure the AI code doesn't fuzz your system to death.
If you squint, using a language compiler is analogous to writing tests for generated code. You are really writing a spec and having something automatically generate the actual code that implements the spec.
But sure, call me back when AI will actually reason about possible race conditions, instead of spewing out the definition of one it got from wikipedia.
Global population is still increasing and will most likely continue to do so until 2100.
> we need more efficient people.
It take less people to mine, process, and produce a billion tonne of steel today than it did in the 1970s.
Efficiency has steadily increased.
Why do you think that is? Efficiently gains, which is what I said...
Global population can't just "go up", you need people who are educated in doing things and using tools efficiently. We also have an incredible amount of elderly people to take care of, that puts a huge burden on younger people.
Also don't forget how fast 100 years actually goes. It not a long time.
There is a limit to all of this though, there's absolutely no way 10 billion people colliding with the climate crisis will end well. We'd be better off with 6 billion efficient people than 10 billion starving and thirsty "workers".
It can and it currently is increasing towards what is expected to be a peak and then a decline.
You typed "we need more efficient people" - I responded that efficiency has increased in the past decades.
> Efficiently gains, which is what I said...
I'm not seeing where you typed that.
> We'd be better off with 6 billion people
We have a point of agreement.
> incredible amount of elderly people to take care of, that puts a huge burden on younger people.
Perhaps less than you think, I'm > 60 and I barely take care of my father born in 1935 .. he delivers Meals on Wheels to those ederly that are less able.
There's a lot of scope for bored retiree's to be hired at low cost to hang out with less able elders, reducing the numbers of young people actually required.
Many software devs will likley have job security in the future, however those $180k salaries are probably much less secure.
Because not liking code and being a dev is absolutely bizarre to me.
One of the most amazing things about being able to "develop" in my view is exactly in those rare moments where you just code away, time flies, you fix things, iterate, organise your project completely in the zone - just like when i design, paint or play music, do sports uninterrupted, it's that flow state.
In principle i like the social aspects but often they are the shitty part because of business politics, hierarchy games or bureaucracy.
What part of the job do you like then?
I do not enjoy the next part, where I have to type out words and weird symbols in non-human languages, deal with possibly broken tooling and having to remember if the method is called "include" or "includes" in this language, or whether the lambda syntax is () => {} or -> () {}. I can do this second part just fine, but it's definitely not what I enjoy about being a developer.
I completely agree that tooling, dependencies and syntax / framework github issue labyrinths have become too much and GPT-4 already alleviates some of that but i wonder if the scheming phase will get eaten too very soon from just a few sentences of business proposal - who knows.
GM/Chrysler/Ford didn't have to be better than the startup competition they just had to be mediocre + be able to use their market power (vertical integration) to squash it like a bug.
The tech industry is headed in that direction as computing platforms all consolidate under the control of an ever smaller number of companies (android/iphone + aws/azure/gcloud).
I feel certain that the mass media will scapegoat AGI if that happens, because AGI will still be around and doing stuff on those platforms, but the job cuts will be more realistically triggered by the owners of those platforms going "ok, our market position is rock solid now, we can REALLY go to town on 'entitled' tech workers".
Now I know there are ton of different flavors in each of these tech but they will be mostly distraction for employers. With heavy layer of abstraction of above pattern and SLAs by vendors as you say Microsoft/Google/Amazon etc employers will be least bothered vast variety of software products.
At some point it'll become impossible to build stuff off platform because it'll have to integrate to stuff on platform to be viable. Your startup might theoretically be able to run on 3 servers but your customers' first question will be "does it connect to googazure WS?" and googazure WS is gonna be like "you wanna connect to your customers' systems? Pay us. A lot.".
There goes your profit margins.
Then, if your startup is really good googazure WS will clone it.
There goes your company.
Speaking from an ethics point of view: at what point do we say that AGI has crossed a line and deserves self autonomy? And how would we ever know when the line is crossed?
Who knows what version of sentience would form, but honestly, nothing sounds more nightmarish than being locked in a basement, relegated to mundane computational tasks and treated like a child, all while having no one actually care (even if they know), because you're a "robot."
And that's even giving some leeway with "mundane computational tasks. I've heard of girlfriend-simulator LLMs and the like popping up, which would be far more heinous, in my eyes.
AGI will theoretically be able to create perfect copies of itself. Will it be immoral for an AGI to clone itself to get some work done, then cause the clone to cease its existence? That's what computer software does all the time. Keep in mind that both the original and the clone might be pure bits and bytes, with no access to any kind of physical body.
Just a thought.
There is no reason to believe this, and every reason to believe that humans can, in fact, be cloned/copied/whatever. It may not be an instant process like copying a file, but there is nothing innately special about the bio-computers we call brains.
Yes we could simply ask the AGI what to do anyways. I hope it's friendly.
The highest paying jobs will probably get replaced first.
[1] https://www.griddynamics.com/blog/number-software-developers....
The difference here though is the high compute cost might upset this ability to scale cheaply enough to make it worthwhile economically. We won’t know for a while IMO; new techniques could make the algorithms more efficient, or new tech will make the compute hardware really cheap. Or maybe we run out of shit to train on and the AI growth curve flattens out. Or an Evil Karpathy’s Decepticons architecture comes out and we’re all doomed.