Someone who's wrong and cagey can cause a lot of damage before anyone figures out what's up. At least a talkative idiot can be managed.
Correctness will go from a binary pass/fail to a probability
What often divides the excellent engineers from the poor ones is how well they can think about corner cases, and tests are mostly about writing down the corner cases in a sustainable way (vs half-assedly writing down half of the corner cases and writing 0.5% of tests as asserting a bug.)
The main problem I saw with DevOps and QA automation was that if people don't write code all day, having them write code that gates release of software to production does not result in good outcomes. Sooner or later developers have to inject some engineering practices.
If the engineers are running AI generated code, nobody knows how to do that job, and you will get a long string of permanently damaged brand names in the aftermath.
I can particularly imagine a regulatory environment much more rigorous than software has got away with thus far, for example strict requirements around certifying that your model doesn’t exhibit X Y Z biases according to standard frameworks of evaluation.
Excellent point. The pervasiveness of neural nets will require engineers (and really, everyone) to start thinking more probabilistically and establish acceptability thresholds instead of certainty. It's the way of the future.
Is 90% accuracy good enough? Is 95%? 99%? 99.9%? No matter the answer, you have to tolerate errors. Now your stakeholders have to tolerate errors. Are they going to accept errors just because "Software 2.0" is here and that's what we all have to live with? Nope.
It may be possible in the future to have "for all practical purposes flawless" software, which might make sense for select special applications. That would be a new thing though, rather than something we have and could lose due to adopting AI development.
This is typically not that true when it comes to correctness. Most software does the correct thing in the eyes of the user, nearly 100% of the time. And when it doesn't, the bug gets fixed and that edge case is corrected for every other user going forward.
AI generated software, from what I've seen, has a wide range of errors in correctness (along with all the other errors that you mentioned all software having..which is true). Like it literally just does the wrong thing given what the user is expecting it to do. The path toward iteratively improving and getting it to an acceptable level of correctness for any given application might be there, but so far I have not seen it.
Let's say the software is good enough if it does the right thing 99.9% of the time it is used. I take it, you're saying that if an AI starts modifying it and only writes correct code in 99.9% of cases (yes, current AI is not even close, but it will improve), that makes it worse because the software might start failing completely. However, if you have proper tests and release management, such obvious flaws will quickly be detected and fixed or rolled back. For most applications that seems pretty much equivalent to what we have now.
The other case is software that is completely AI generated and cannot reasonably be modified by humans anymore. In that case, again, you have tests and a sane deployment strategy that mitigates failures to a sufficient degree depending on application.
So the only issue is when you start making a completely AI generated software and fail to ever meet requirements, or pass the human written test cases? Users at least will never be impacted by that. Even now, many human-written software projects never get to the stage where they can be used. Is this really a problem, especially if the attempt at AI generation of software is cheap?
This is really my main argument currently...
> but it will improve
With this being the "if" question. Improve, but improve to a point where we can trust it to do the things with the level of correctness actually required? Unclear so far. GPT3, the state of the art, can't be trusted to answer basic questions correctly (yet).
I also don't personally buy into the other notions in these comments that "the future is probabilistic software". I think that's wishful thinking outside of some specific domains and an attempt to bend our actual requirements to meet the capabilities of AI software, rather than the opposite.
> pass the human written test cases
I'm not super sold on this idea either. It seems reasonably possible that writing the test cases to a level of specification necessary to ensure that correctness we're after is just as much effort as just writing the code.
But, time will tell with everything.
It's usually much easier (and never harder) to specify what needs to be done than how it should be done. Whether you can trust the result (enough) depends on the application. AI driven development will be applied first wheree errors and failures are least harmful and advance from there. It might take a long time until the degree of correctness improves enough and trust is build around it. After all, some countries' railroad networks still don't use computers but instead have a human map out new routes and schedules on paper to make sure trains don't crash. Nevertheless, the speed of AI development has exceeded (at least my) expectations time and time again in recent years.
What might happen is that we eventually get a lot of software that fails more often than now but is a lot cheaper. That would still mean that the new methods are widely adopted. People accept software errors as a part of life already, so even if they get more frequent in less critical applications, we will adapt.
What is really dangerous is when the various software components get too fast and complex to understand or control and develop pathological feedback loops in situations that cause real trouble. This kind of thing is a continuum of badness which tops out at "AI taking over and wiping out humanity". Given market incentives driving adoption, which I anticipate to be strong, it's hard to imagine how such risks might be mitigated.
Yes but I'll point you back at my original comment about correctness. I've never been on a team that shipped code we knew would do the wrong thing. I ship code with known failure points all the time. But when the code runs to completion, I'm pretty darn sure it's doing the correct thing and we try very, very hard to make sure of that. With AI I am seeing that they can't discern between issues of correctness and issues of failure or availability. It's just like 95% success across all of those spectrums.
> What is really dangerous is when the various software components get too fast and complex to understand or control and develop pathological feedback loops in situations that cause real trouble.
Yea I agree, that's somewhat scary to think about.
You're confusing pleasantness with correctness. Pleasantness is the property of pleasing the user, or more often the software's owner. Correctness is the property of conforming to a specification. Since most software written has no specification, correctness is undefined. Evidently this works adequately well in the marketplace.
As pointed out, there are already formal languages that allow formal verification like B [0] notably for like-critical systems.
I would assume this actually gels very nicely with neural nets since its constantly optimizing for fitness. Hell in theory you could bake in your SLA/SLIs into the models to self correct? Give the model direct feedback that its unfit?
We already do that in manufacturing. Physical parts are imperfect and we design with such variation in mind.
While it's possible for a highly-skilled, highly-professional developer can both write code that will solve a given problem 100% correctly and write tests that will prove that it solves them for the entire domain, in practice most developers fall short on both counts. Every time you interact with a date or phone number field that chastises you for your use or non-use of punctuation, you know this.
So, for many use cases, it's possible imperfect programmers will be replaced with neural networks that are 95% accurate, perhaps with a differently-trained one checking the work of the first one.
1. Generate a 95% accurate model
2. Use it to generate test cases
3. Code the thing, with the help of the cases
4. Manually remove cases in the 5%
I image steps 1 and 2 being completed by a product owner and 3 and 4 being completed by a software engineer.
We're so horrifically bad at communicating a requirement's intent, I wonder what would happen if we tried to use AI to communicate them via their extent instead.
I mean ... can it be done?
- build a platform (ie the data we care about and are going to build some workflow over)
- have business describe what should happen in english
- How does GPT build something that will run? Can it create the infrastructure? does it speak AWS?
- then ...
OK - I am actually excited by that
So if AI can give us a mockup that's workable enough to skip the first few iterations of "no that's not what I want" and get right to the part where the engineer is asking questions about the edge cases that weren't explicit in the requirements... That's a win.
I imagine you'd still want to have it in a box of some sort re: creating infra. Like you give it a very small cluster and probably make the stack decisions "write me a postgres schema for... write a fastapi API for the schema... write a react UI for the API... write me a k8s operator that up/down's the above components... Workshop the idea with other product people...
...and only then involve the engineer like: "make this AI-generated house of cards into a fortress".
> Our software has a 99% chance to calculate your taxes correctly! And only a 1% chance of failure in which it's your fault and it's you that's committing tax fraud
1. making a mistake in this part is really common, ChatGPT makes common mistake.
2. you have uncommon situation affecting here, ChatGPT ignores and writes things that cause you to get in trouble, or it writes things that cause you to pay more than you should.
Also the longer is goes on writing things the more likely that things it writes does not hang together with the past things it wrote, when a human lies they try to make their lies at least follow a sensible pattern. ChatGPT would be likely to get you flagged for audits because you can't be sure that what it wrote on page 1 jibes with what it writes on page 3.
Family member got ripped by a government audit of what was supposed all fine by the person doing his company taxes.
For some use cases, sure. We go through painstaking efforts to ensure things like correctness, consistency, and idempotency for a reason though. Most things we want to be deterministic, and when something's not deterministic we freak out and fix it ASAP (including waking people up in the middle of the night to do so)
Assume your software has a probability to fail or have bugs or gets hit by bit flips or unreliable hardware. There’s a whole field for dealing with those kind of things that typical web devs haven’t had to worry about as much.
https://github.com/williamcotton/empirical-philosophy/blob/m...
The AI doesn't have to be perfect, but only offer a lower error rate than humans.
- Lange Clinical Neurology - 11th Edition
- Bradley's Neurology in Clinical Practice, 8th Edition
"System complexity, particularly in software systems, making SIL estimation difficult to impossible"
"The requirements of these schemes can be met either by establishing a rigorous development process, or by establishing that the device has sufficient operating history to argue that it has been proven in use."
You could prove that normal code satisfies some specs, but you can't do that with neural nets unless the number of possible inputs is tiny. So, the only way to establish that the black box neural net meets some SIL target is through "sufficient operating history".
“Don’t worry, we stuck the flight data recorder in the training set, and rebuilt the model. Should be good to go now”?
Low code and AI are going to have many of the same failure modes. Until someone combines them and then they'll have exactly the same failure modes.
Probably, but is testing to the necessary level of correctness more or less effort than writing the code ourselves?
There's still a lot of very, very sophisticated work that goes into locking down requirements that tightly.
That doesn’t make any sense. What passes the proof is the program code. You have the code, and then you construct a formal proof that the code is correct, similar to how a mathematician proves that some theorem is correct. The code is a prerequisite for the proof. When you can construct a proof for the code, you’re done.
This proof-construction process is what AIs currently aren’t good at, because it requires logical precision, and probability isn’t sufficient. They can generate code, but they can’t construct the formal proof that the code is correct (and it often isn’t).
> There's still a lot of very, very sophisticated work that goes into locking down requirements that tightly.
What’s true is that you need to know what you want to prove about the code, and that isn’t always easy.
What use is a banking app if it's only going to be correct some of the time?
Something like Godbolt for neural networks? Can’t be that far away
I’d like to see more work done to incorporate all the advances AI has brought us into our traditional software. Using the example from the article, if databases can be 70% faster and use 10x less memory by leveraging a neural net, how? If we can figure out the how, we can understand it and incorporate it in other areas. A great deal of success in various fields draws on inspiration from other fields, for example biomimicry. We even come up with mental models in areas like computer science that trivialize complex topics to simple objects a child could understand (trees, stacks, etc.). We would benefit immensely from learning from neural networks, but instead we have decided to largely ignore the how and see what they can do. Both are important, but one is severely lacking.
If AI don't care about the process and only care about results. These types of thing is going to happen everywhere. And from this moment, no code is understandable to human.
Having AI write a 90% accurate 1M line (or "parameter") codebase all at once, (which seems to be the expectation here), is the "risk" you're overlooking. No human will be able to know where to start debugging that. At least not yet. But will yet come before incredibly dangerous amounts of AI written code is pushed into critical systems everywhere, by naively optimi$tic opportuni$t$?
I'll also note that the next hardest part of programming is troubleshooting "in production" whether a Web application, in an embedded device, or running on someone's machine. Is the "AI" going to help there? Are we going to even be able to fix those problems when doing so could make the code we don't understand fail in another way or hit the wrong side of the performance tradeoff the "optimized code" entailed?
My dad was worried about whether I should go into a CS degree because there was a code generation cycle going on at the time and people thought the computers would be programming themselves. We've had a few since, and my skills are more valuable than ever.
This article is from 2017, and I don't know about you, but I increasingly feel like software is more and more broken year after year. I think that has a lot to do with "delegating complexity" and building things without really understanding what they're doing. I think the correlation between human understanding of the fine details and underlying logic and desired outcome and software quality is pretty tight. That doesn't mean "software 2.0" neural net stuff doesn't fit in, you just need a human to plug it in right that really understands its benefits and its limitations.
The author mentions the downsides, but I think they underestimate them. If you lean on AI too heavily and don't ever translate to a traditional language, you've basically liquified your logic/there's no garbage dump to salvage from when things go bad and zero understanding of implicit context. If you use it to generate ostensibly human readable code you get "documentation" with no guarantee of accuracy, which makes it worse than having no documentation (depending on how high the error bars are). While that's not a new problem either, if it's autogenerated that means it's easy to create way more of it than human generated code, which means it'll probably be an ever larger portion of what gets sucked up into later AI models. If they become too self referential they'll become increasingly detached from human judgement about whether the code is doing what it should and error bars will grow.
I'm still convinced these things are virtually all going to end up in a fancy autocomplete suggestion and compression niche after a lot of pain. But that's still a big deal/I don't think the limitations of these means they don't have a big future place. The sheer number of things you can have autosuggest for now with these AI models are amazing, and that expansion is boosting productivity and creating a large number of new products that are going to become essential tools. That being said, every time this type of thread comes up I'm like "woah woah woah, pump the breaks, these things have no understanding of what they're doing. You can't just stop thinking about stuff and let a machine do it, bad bad bad idea."
The reality is it all comes down to testing. If S1.0 has a unit test that says "Person should not get financial help", S2.0 should also have a unit test that says "Person should not get financial help", and work the same.
Of course my unit test name is designed to enrage, but let's be honest, we're writing code that makes these decisions.
I keep complaining about some NIH code we have and people say things like, "well that library didn't exist at the time, so we had to write one."
Git history says they're wrong. Time and again they were writing code that had existed for two years already.
Soon we'll have camps advocating various rhetorical paradigms to program AIs. Instead of imperative versus declarative programming, we'll have effusive versus abusive prompt engineering:
"The best way to program is to be really nice to the AI, and tell it how much you appreciate it and it will come up with the best solution itself."
"No! The best paradigm is to berate the AI and to beat it into submission to get exactly what you want!"
The more I hear people talk about programming in the future, the more it sounds like that's really where some people want to take us. I'm not excited.But, as the article points out, at least for a considerable set of problem domains, it's not.
My general rule of thumb is that anything that has a relatively small, finite, discrete set of inputs and outputs is better suited to "Software 1.0" (e.g. coding a calculator). But there are a huge number of domains, some highlighted in the article, which usually have to do with an infinite possible number of inputs and outputs, where human written code is not faster or more correct than a neural network.
Developers and machines that can't explain themselves end up out on their ear or damaging the projects they're involved in.
Similarly, I don't whether AI weights are computed all in virtual parallel, but if they are computing every node simultaneously that will be less efficient than the Neumann model in which a Program Counter (PC) acts like a "cursor" hopping arund and iterating states of the discrete code model at points throughout it. E.g. a video game with a controller and various sprites will have objects that update at various rates and the player moves with different code underwater than in the air, than on land, so different parts of the model would execute.
...
"WTF! We have so much technical debt! Fire the code monkeys and replace them with AI! We need more Software (AI)"
I'm sure this is going great.
Maybe you can use some special purpose artificial language created with the purpose of writing unambiguous texts... Like Java or Python.
then add additional specifications to clarify ambiguity when observed
That's such a clean, almost clinical way to describe, "After the software has caused several billion dollars in trading losses, bankrupting the entire company." (https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stoc...) Imagine if software with large inherent risk was developed using formal
methods, and the massive remainder of software was developed using rapid
development methods.
Imagine if we could tell the difference between use-cases with large inherent risk and the massive remainder of use-cases and avoid using software designed for low-risk situations in high-risk applications.Or just use English and do lots of acceptance testing plus add fail safes. I bet that will be more economical.
Formally verifying it will easily take more computing power than training the AI on the first place, so I don't count that one as viable.
And yes, it is a very obvious point, and that people keep missing that point on this site is unsettling. (Also, yes, this can be trivially circumvented if you just let those people program, instead of only do verification.)
You can, also obviously, replace millions of average developers with (way more) millions of extremely competent spec writers if they can use formal methods. Those will require way more computing power than it can ever exist on Earth to do their work, but they can mathematically get there.
Why do you think you need orders of magnitude more spec writers than coders rather than the other way around?
AFAIK, those are the only ones in common use, but differently from formal ones, non-formal things tends to come on a multitude of widely different types. So I wouldn't be surprised if people have invented many more.
Gonna be really fun when the first financial companies start trying to generate billing software from specs and end up blowing up entire people's bank accounts irrevocably because of lax regulation and greedy shareholders. Not to mention that any resources you trim from the developer side of things you'll have to at least quintuple in QA.
And you can also use the AI to check for correctness (of course, again not 100% accurate), potential issues, potential improvements, etc.
The fault lies with humans treating these programming AIs like the greenest of junior bootcamp devs would treat the most accomplished senior engineer.
One way or another, I want to be able to read the source. Or am I missing the point? Maybe future software will be a big neural network blackbox that we verify purely through tests?
Also if it could recommend more efficient routes for how people wrote the code.
So, what are best testing frameworks people have worked with?
It requires a perspective shift. Then again, cost efficiency is a big motivator.
90% failure rate can be lousy in one use case and can be perfectly adequate in another.