The black box that lets non-technical folk type in business requirements and generate an end to end application is still very much an open research question. Getting 70% on SWEBench is an absolute accomplishment, but have you seen the problems? 1. These aren't open-ended requests like here's a codebase, implement x feature and fix y bug. They're issues with detailed descriptions written by engineers evaluated against a set of unit tests. Who writes those descriptions? Who writes the unit tests to verify whatever the LLM generated? Software engineers. 2. OpenAI had a hand in designing the benchmark and part of the changes they made included improving the issue description and refining the test sets [1]. Do you know who made these improvements? "professional software developers" 3. Issues were pulled from public, popular open source Python projects on Github. These repos have almost certainly found their way into the model's training set. It doesn't strike me as unlikely that the issues and their solutions ended up in the training set too. I'm a lot more curious about how well o3 performs on the Konwinski Prize which tries to fix the dataset tainting problem.
The proposed solution to this is just throw another AI system that can convert ambiguous business requirements/bug reports into a formal spec and write unit tests to act as a verifier. This is a non-trivial problem and reasoning-style models like o1 degrade in performance when given imperfect verifiers [2]. I can't find any widely used benchmarks that check how good LLMs are at verifying LLM-generated code. I also can't find any that check e2e performance of prompt -> app problems I'm guessing because that would require a lot of subjective human feedback you can't automate like unit tests.
LLMs (augmented by RL and test time compute like o3) are getting better at the things I think they're already pretty good at: given a well-defined problem that can be iteratively reviewed/verified by an expert (you), come up with a solution. They're not necessarily getting better at everything that would be necessary to fully automate knowledge jobs. There could be a breakthrough tomorrow using AI for verification/spec generation (in which case we and most everyone else are well and truly screwed) but until that happens the current trajectory seems to be AI-assisted coding.
Software engineering will be about using your knowledge of computing to translate vague/ambiguous feature requests/bug reports from clients/management into detailed, well-specified problem statements and design tests to act as ground truth for an AI system like o3 (and beyond) to solve. Basically test-driven development on steroids :) There may indeed still be layoffs or maybe we run into Jevons paradox and there's another explosion in the amount of software that gets built necessitating engineers good at using LLMs to solve problems.
However, if the worst comes to pass my plan is finding an apprenticeship as quickly as possible. I've seen the point that an overall reduction in white collar work would result in a reduction of opportunities for the trades but I doubt mass layoffs would occur in every sector all at once. Other industries are more highly regulated and have protectionist tendencies that would slow AI automation adoption (law, health, etc). Software has the distinct disadvantage of being low-regulation (there aren't many laws that require high-quality code outside of specialized domains like medtech) and a culture of individualism that would deter attempts at collective bargaining. We also literally put our novel IP up on a free, public, easily indexable platform under permissive licenses. We probably couldn't make it easier for companies to replace us.
So while at least some knowledge workers have their jobs, there's an opportunity to put food on the table by doing their wiring, pipes, etc. The other counterargument is improvements in embodied AI i.e. robotics will render manual labor redundant. The question isn't will we have the tech (we will), it's whether the average person is going to be happy letting a robot armed with power tools into their home and how long it will take for said robot to be cheaper than a human.
[1] https://openai.com/index/introducing-swe-bench-verified/ [2] https://arxiv.org/abs/2411.17501