Software 2.0 (2017)
karpathy.medium.com
karpathy.medium.com
Correctness will go from a binary pass/fail to a probability
Excellent point. The pervasiveness of neural nets will require engineers (and really, everyone) to start thinking more probabilistically and establish acceptability thresholds instead of certainty. It's the way of the future.
Is 90% accuracy good enough? Is 95%? 99%? 99.9%? No matter the answer, you have to tolerate errors. Now your stakeholders have to tolerate errors. Are they going to accept errors just because "Software 2.0" is here and that's what we all have to live with? Nope.
It may be possible in the future to have "for all practical purposes flawless" software, which might make sense for select special applications. That would be a new thing though, rather than something we have and could lose due to adopting AI development.
This is typically not that true when it comes to correctness. Most software does the correct thing in the eyes of the user, nearly 100% of the time. And when it doesn't, the bug gets fixed and that edge case is corrected for every other user going forward.
AI generated software, from what I've seen, has a wide range of errors in correctness (along with all the other errors that you mentioned all software having..which is true). Like it literally just does the wrong thing given what the user is expecting it to do. The path toward iteratively improving and getting it to an acceptable level of correctness for any given application might be there, but so far I have not seen it.
Let's say the software is good enough if it does the right thing 99.9% of the time it is used. I take it, you're saying that if an AI starts modifying it and only writes correct code in 99.9% of cases (yes, current AI is not even close, but it will improve), that makes it worse because the software might start failing completely. However, if you have proper tests and release management, such obvious flaws will quickly be detected and fixed or rolled back. For most applications that seems pretty much equivalent to what we have now.
The other case is software that is completely AI generated and cannot reasonably be modified by humans anymore. In that case, again, you have tests and a sane deployment strategy that mitigates failures to a sufficient degree depending on application.
So the only issue is when you start making a completely AI generated software and fail to ever meet requirements, or pass the human written test cases? Users at least will never be impacted by that. Even now, many human-written software projects never get to the stage where they can be used. Is this really a problem, especially if the attempt at AI generation of software is cheap?
This is really my main argument currently...
> but it will improve
With this being the "if" question. Improve, but improve to a point where we can trust it to do the things with the level of correctness actually required? Unclear so far. GPT3, the state of the art, can't be trusted to answer basic questions correctly (yet).
I also don't personally buy into the other notions in these comments that "the future is probabilistic software". I think that's wishful thinking outside of some specific domains and an attempt to bend our actual requirements to meet the capabilities of AI software, rather than the opposite.
> pass the human written test cases
I'm not super sold on this idea either. It seems reasonably possible that writing the test cases to a level of specification necessary to ensure that correctness we're after is just as much effort as just writing the code.
But, time will tell with everything.
It's usually much easier (and never harder) to specify what needs to be done than how it should be done. Whether you can trust the result (enough) depends on the application. AI driven development will be applied first wheree errors and failures are least harmful and advance from there. It might take a long time until the degree of correctness improves enough and trust is build around it. After all, some countries' railroad networks still don't use computers but instead have a human map out new routes and schedules on paper to make sure trains don't crash. Nevertheless, the speed of AI development has exceeded (at least my) expectations time and time again in recent years.
What might happen is that we eventually get a lot of software that fails more often than now but is a lot cheaper. That would still mean that the new methods are widely adopted. People accept software errors as a part of life already, so even if they get more frequent in less critical applications, we will adapt.
What is really dangerous is when the various software components get too fast and complex to understand or control and develop pathological feedback loops in situations that cause real trouble. This kind of thing is a continuum of badness which tops out at "AI taking over and wiping out humanity". Given market incentives driving adoption, which I anticipate to be strong, it's hard to imagine how such risks might be mitigated.
Yes but I'll point you back at my original comment about correctness. I've never been on a team that shipped code we knew would do the wrong thing. I ship code with known failure points all the time. But when the code runs to completion, I'm pretty darn sure it's doing the correct thing and we try very, very hard to make sure of that. With AI I am seeing that they can't discern between issues of correctness and issues of failure or availability. It's just like 95% success across all of those spectrums.
> What is really dangerous is when the various software components get too fast and complex to understand or control and develop pathological feedback loops in situations that cause real trouble.
Yea I agree, that's somewhat scary to think about.
You're confusing pleasantness with correctness. Pleasantness is the property of pleasing the user, or more often the software's owner. Correctness is the property of conforming to a specification. Since most software written has no specification, correctness is undefined. Evidently this works adequately well in the marketplace.
As pointed out, there are already formal languages that allow formal verification like B [0] notably for like-critical systems.
We already do that in manufacturing. Physical parts are imperfect and we design with such variation in mind.
I would assume this actually gels very nicely with neural nets since its constantly optimizing for fitness. Hell in theory you could bake in your SLA/SLIs into the models to self correct? Give the model direct feedback that its unfit?
While it's possible for a highly-skilled, highly-professional developer can both write code that will solve a given problem 100% correctly and write tests that will prove that it solves them for the entire domain, in practice most developers fall short on both counts. Every time you interact with a date or phone number field that chastises you for your use or non-use of punctuation, you know this.
So, for many use cases, it's possible imperfect programmers will be replaced with neural networks that are 95% accurate, perhaps with a differently-trained one checking the work of the first one.
1. Generate a 95% accurate model
2. Use it to generate test cases
3. Code the thing, with the help of the cases
4. Manually remove cases in the 5%
I image steps 1 and 2 being completed by a product owner and 3 and 4 being completed by a software engineer.
We're so horrifically bad at communicating a requirement's intent, I wonder what would happen if we tried to use AI to communicate them via their extent instead.
I mean ... can it be done?
- build a platform (ie the data we care about and are going to build some workflow over)
- have business describe what should happen in english
- How does GPT build something that will run? Can it create the infrastructure? does it speak AWS?
- then ...
OK - I am actually excited by that
So if AI can give us a mockup that's workable enough to skip the first few iterations of "no that's not what I want" and get right to the part where the engineer is asking questions about the edge cases that weren't explicit in the requirements... That's a win.
I imagine you'd still want to have it in a box of some sort re: creating infra. Like you give it a very small cluster and probably make the stack decisions "write me a postgres schema for... write a fastapi API for the schema... write a react UI for the API... write me a k8s operator that up/down's the above components... Workshop the idea with other product people...
...and only then involve the engineer like: "make this AI-generated house of cards into a fortress".
> Our software has a 99% chance to calculate your taxes correctly! And only a 1% chance of failure in which it's your fault and it's you that's committing tax fraud
1. making a mistake in this part is really common, ChatGPT makes common mistake.
2. you have uncommon situation affecting here, ChatGPT ignores and writes things that cause you to get in trouble, or it writes things that cause you to pay more than you should.
Also the longer is goes on writing things the more likely that things it writes does not hang together with the past things it wrote, when a human lies they try to make their lies at least follow a sensible pattern. ChatGPT would be likely to get you flagged for audits because you can't be sure that what it wrote on page 1 jibes with what it writes on page 3.
Family member got ripped by a government audit of what was supposed all fine by the person doing his company taxes.
For some use cases, sure. We go through painstaking efforts to ensure things like correctness, consistency, and idempotency for a reason though. Most things we want to be deterministic, and when something's not deterministic we freak out and fix it ASAP (including waking people up in the middle of the night to do so)
Assume your software has a probability to fail or have bugs or gets hit by bit flips or unreliable hardware. There’s a whole field for dealing with those kind of things that typical web devs haven’t had to worry about as much.
https://github.com/williamcotton/empirical-philosophy/blob/m...
The AI doesn't have to be perfect, but only offer a lower error rate than humans.
- Lange Clinical Neurology - 11th Edition
- Bradley's Neurology in Clinical Practice, 8th Edition
“Don’t worry, we stuck the flight data recorder in the training set, and rebuilt the model. Should be good to go now”?
Low code and AI are going to have many of the same failure modes. Until someone combines them and then they'll have exactly the same failure modes.
"System complexity, particularly in software systems, making SIL estimation difficult to impossible"
"The requirements of these schemes can be met either by establishing a rigorous development process, or by establishing that the device has sufficient operating history to argue that it has been proven in use."
You could prove that normal code satisfies some specs, but you can't do that with neural nets unless the number of possible inputs is tiny. So, the only way to establish that the black box neural net meets some SIL target is through "sufficient operating history".
What use is a banking app if it's only going to be correct some of the time?
Probably, but is testing to the necessary level of correctness more or less effort than writing the code ourselves?
There's still a lot of very, very sophisticated work that goes into locking down requirements that tightly.
That doesn’t make any sense. What passes the proof is the program code. You have the code, and then you construct a formal proof that the code is correct, similar to how a mathematician proves that some theorem is correct. The code is a prerequisite for the proof. When you can construct a proof for the code, you’re done.
This proof-construction process is what AIs currently aren’t good at, because it requires logical precision, and probability isn’t sufficient. They can generate code, but they can’t construct the formal proof that the code is correct (and it often isn’t).
> There's still a lot of very, very sophisticated work that goes into locking down requirements that tightly.
What’s true is that you need to know what you want to prove about the code, and that isn’t always easy.
What often divides the excellent engineers from the poor ones is how well they can think about corner cases, and tests are mostly about writing down the corner cases in a sustainable way (vs half-assedly writing down half of the corner cases and writing 0.5% of tests as asserting a bug.)
The main problem I saw with DevOps and QA automation was that if people don't write code all day, having them write code that gates release of software to production does not result in good outcomes. Sooner or later developers have to inject some engineering practices.
If the engineers are running AI generated code, nobody knows how to do that job, and you will get a long string of permanently damaged brand names in the aftermath.
I can particularly imagine a regulatory environment much more rigorous than software has got away with thus far, for example strict requirements around certifying that your model doesn’t exhibit X Y Z biases according to standard frameworks of evaluation.
Having AI write a 90% accurate 1M line (or "parameter") codebase all at once, (which seems to be the expectation here), is the "risk" you're overlooking. No human will be able to know where to start debugging that. At least not yet. But will yet come before incredibly dangerous amounts of AI written code is pushed into critical systems everywhere, by naively optimi$tic opportuni$t$?
Something like Godbolt for neural networks? Can’t be that far away
Someone who's wrong and cagey can cause a lot of damage before anyone figures out what's up. At least a talkative idiot can be managed.
If AI don't care about the process and only care about results. These types of thing is going to happen everywhere. And from this moment, no code is understandable to human.
I’d like to see more work done to incorporate all the advances AI has brought us into our traditional software. Using the example from the article, if databases can be 70% faster and use 10x less memory by leveraging a neural net, how? If we can figure out the how, we can understand it and incorporate it in other areas. A great deal of success in various fields draws on inspiration from other fields, for example biomimicry. We even come up with mental models in areas like computer science that trivialize complex topics to simple objects a child could understand (trees, stacks, etc.). We would benefit immensely from learning from neural networks, but instead we have decided to largely ignore the how and see what they can do. Both are important, but one is severely lacking.
Maybe you can use some special purpose artificial language created with the purpose of writing unambiguous texts... Like Java or Python.
Or just use English and do lots of acceptance testing plus add fail safes. I bet that will be more economical.
And yes, it is a very obvious point, and that people keep missing that point on this site is unsettling. (Also, yes, this can be trivially circumvented if you just let those people program, instead of only do verification.)
You can, also obviously, replace millions of average developers with (way more) millions of extremely competent spec writers if they can use formal methods. Those will require way more computing power than it can ever exist on Earth to do their work, but they can mathematically get there.
Why do you think you need orders of magnitude more spec writers than coders rather than the other way around?
AFAIK, those are the only ones in common use, but differently from formal ones, non-formal things tends to come on a multitude of widely different types. So I wouldn't be surprised if people have invented many more.
Gonna be really fun when the first financial companies start trying to generate billing software from specs and end up blowing up entire people's bank accounts irrevocably because of lax regulation and greedy shareholders. Not to mention that any resources you trim from the developer side of things you'll have to at least quintuple in QA.
Formally verifying it will easily take more computing power than training the AI on the first place, so I don't count that one as viable.
then add additional specifications to clarify ambiguity when observed
That's such a clean, almost clinical way to describe, "After the software has caused several billion dollars in trading losses, bankrupting the entire company." (https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stoc...) Imagine if software with large inherent risk was developed using formal
methods, and the massive remainder of software was developed using rapid
development methods.
Imagine if we could tell the difference between use-cases with large inherent risk and the massive remainder of use-cases and avoid using software designed for low-risk situations in high-risk applications.And you can also use the AI to check for correctness (of course, again not 100% accurate), potential issues, potential improvements, etc.
The fault lies with humans treating these programming AIs like the greenest of junior bootcamp devs would treat the most accomplished senior engineer.
Developers and machines that can't explain themselves end up out on their ear or damaging the projects they're involved in.
But, as the article points out, at least for a considerable set of problem domains, it's not.
My general rule of thumb is that anything that has a relatively small, finite, discrete set of inputs and outputs is better suited to "Software 1.0" (e.g. coding a calculator). But there are a huge number of domains, some highlighted in the article, which usually have to do with an infinite possible number of inputs and outputs, where human written code is not faster or more correct than a neural network.
Similarly, I don't whether AI weights are computed all in virtual parallel, but if they are computing every node simultaneously that will be less efficient than the Neumann model in which a Program Counter (PC) acts like a "cursor" hopping arund and iterating states of the discrete code model at points throughout it. E.g. a video game with a controller and various sprites will have objects that update at various rates and the player moves with different code underwater than in the air, than on land, so different parts of the model would execute.
The reality is it all comes down to testing. If S1.0 has a unit test that says "Person should not get financial help", S2.0 should also have a unit test that says "Person should not get financial help", and work the same.
Of course my unit test name is designed to enrage, but let's be honest, we're writing code that makes these decisions.
I keep complaining about some NIH code we have and people say things like, "well that library didn't exist at the time, so we had to write one."
Git history says they're wrong. Time and again they were writing code that had existed for two years already.
Soon we'll have camps advocating various rhetorical paradigms to program AIs. Instead of imperative versus declarative programming, we'll have effusive versus abusive prompt engineering:
"The best way to program is to be really nice to the AI, and tell it how much you appreciate it and it will come up with the best solution itself."
"No! The best paradigm is to berate the AI and to beat it into submission to get exactly what you want!"
The more I hear people talk about programming in the future, the more it sounds like that's really where some people want to take us. I'm not excited.I'll also note that the next hardest part of programming is troubleshooting "in production" whether a Web application, in an embedded device, or running on someone's machine. Is the "AI" going to help there? Are we going to even be able to fix those problems when doing so could make the code we don't understand fail in another way or hit the wrong side of the performance tradeoff the "optimized code" entailed?
My dad was worried about whether I should go into a CS degree because there was a code generation cycle going on at the time and people thought the computers would be programming themselves. We've had a few since, and my skills are more valuable than ever.
This article is from 2017, and I don't know about you, but I increasingly feel like software is more and more broken year after year. I think that has a lot to do with "delegating complexity" and building things without really understanding what they're doing. I think the correlation between human understanding of the fine details and underlying logic and desired outcome and software quality is pretty tight. That doesn't mean "software 2.0" neural net stuff doesn't fit in, you just need a human to plug it in right that really understands its benefits and its limitations.
The author mentions the downsides, but I think they underestimate them. If you lean on AI too heavily and don't ever translate to a traditional language, you've basically liquified your logic/there's no garbage dump to salvage from when things go bad and zero understanding of implicit context. If you use it to generate ostensibly human readable code you get "documentation" with no guarantee of accuracy, which makes it worse than having no documentation (depending on how high the error bars are). While that's not a new problem either, if it's autogenerated that means it's easy to create way more of it than human generated code, which means it'll probably be an ever larger portion of what gets sucked up into later AI models. If they become too self referential they'll become increasingly detached from human judgement about whether the code is doing what it should and error bars will grow.
I'm still convinced these things are virtually all going to end up in a fancy autocomplete suggestion and compression niche after a lot of pain. But that's still a big deal/I don't think the limitations of these means they don't have a big future place. The sheer number of things you can have autosuggest for now with these AI models are amazing, and that expansion is boosting productivity and creating a large number of new products that are going to become essential tools. That being said, every time this type of thread comes up I'm like "woah woah woah, pump the breaks, these things have no understanding of what they're doing. You can't just stop thinking about stuff and let a machine do it, bad bad bad idea."
...
"WTF! We have so much technical debt! Fire the code monkeys and replace them with AI! We need more Software (AI)"
I'm sure this is going great.
It requires a perspective shift. Then again, cost efficiency is a big motivator.
90% failure rate can be lousy in one use case and can be perfectly adequate in another.
So, what are best testing frameworks people have worked with?
One way or another, I want to be able to read the source. Or am I missing the point? Maybe future software will be a big neural network blackbox that we verify purely through tests?
Also if it could recommend more efficient routes for how people wrote the code.
For example, as someone who works with financial software, I don't see Karpathy's "Software 2.0" replacing, say, account ledgering software anytime soon. "Yeah, we calculate our clients' balances correctly 99.9% of the time!" isn't going to cut it.
But I don't think that's what Karpathy is arguing. There is a large set of problem domains where Karpathy's Software 2.0 is a much better solution than what he calls Software 1.0. For example, even in finance, stuff like fraudulent transaction detection, or financial security software for intrusion detection, is very well-suited to Software 2.0.
So yes, I think Software 1.0 will always be around, but I don't think it makes sense to use it for domains where Software 2.0 is a better fit. What I feel like Karpathy is arguing for is really now a recognition that Software 2.0 really is a whole new paradigm shift, and we need better tooling (he uses "GitHub for Software 2.0" as an example) to support it.
The real interesting things will be Software 1.0 and 2.0 working together. You use 1.0 to run and validate the work of 2.0 that is guided by prompts. An example of this would be using prompts to generate source code that is compiled and tested. The real TDD is only writing tests and letting Software 2.0 create the code for you. This extends to other work like engineering as well.
The kinds of software that "Software 1.0" is suitable for are markedly different that the ones "Software 2.0" are. As Karpathy argues, it's a different tool, suitable for different tasks, and it should have a different name.
I find that teams and products with negative pun acronyms fail for example. I find that the trend of the name Lilith being popular correlates with abortion rates (in the bible, Lilith killed unborn babies) etc etc etc
cf. gana-pathy (Lord of the Ganas)
Wait, wait, I have heard that it was called BPMN and it generated underneath Enterprise Java Beans and it was working amazingly, all software is written today, right? Right?
But, but nobody touches this crap besides generating pictures to show on Powerpoint slides. Because writing software is circa 10% of all effort, specification, legal stuff, maintenance, avoiding technical debt, proper test cases, anomaly testing, performance testing the right stuff. That is hard, that matters. Software 2.0 is barking the wrong tree.
Needless to say it didn't work. Accurately describing what a program should do is what a programmer does. We're not telling the computer what to do (move this value from memory into this register, etc). We're describing a program's behaviour.
So I think we'll just get another language that is a good fit for describing what the program should do, and instead of compiling that to machine code, it will be used to train a model.
It was still programmers using the language. Same story with SQL.
The only things you can give people who don't invest a substantial amount of time and effort into the craft are markup languages and very high level (configuration) DSLs.
What are you talking about? Vast swaths of non-programmers are using SQL every day. Unless you consider anyone who writes SQL to be a programmer, in which case what you say is true by definition.
When I first joined the industry, back in the early 90's, COBOL was very much the premier business coding language, and I think that only changed with the arrival of Java.
Can you explain further?
So it is really confusing because they are only evaluating "pretty pictures" and you're evaluating the technical requirements of the tool (to do what it was built for) and you mismatch.
Exactly. For serious projects (not to do lists) u will need humans, or equivalent (AI that is grown like one).
"What is it that you do here?" and he says:
"I take the specifications and hand them to the software engineers".
I think our jobs are about to approach that a lot more closely than you might think. Our jobs will be effective translators between what the business wants and what the AI outputs.
In programming we can always make a tradeoff between precision and effort, e.g. by importing libraries or using no-code or code-generation tools. ChatGPT is just one more point on the same tradeoff curve. It hasn't meaningfully moved the curve itself.
Moving the curve would mean making it less effortful to write code at the same level of precision.
Sure sometimes an LLM generates a correct program. And sometimes horoscopes predict the future.
I will be impressed when we can write a precise and formally verifiable specification of a program and some other program can generate the code for us from that specification and prove the generated implementation is faithful to the specification.
An active area of research here, code synthesis, is promising! It's still a long way off from generating whole programs from a specification. The search spaces are not small. And even using a language as precise as mathematics leaves a lot of search space and ambiguity.
Where we're going today with LLM's trying to infer a program from an imprecise specification written in informal language is simply disappointing.
The software industry has gotten this far with very little help from the formal methods community... that's changing in recent years in certain spaces where errors are magnified by scale like cloud computing, etc.
But instead of getting better at writing precise specifications we're going to continue to be bad at it and hope that an LLM can manage to infer the correct program. It might be millions of lines of code but hopefully it does some of the things we want most of the time.
Update: To be clear, I'm not saying AI/ML programs cannot help us to write programs at all, just that the inputs need to be better if we're going to have any confidence that the programs it generates are any better than a horoscope.
Today's ecosystem requires advanced knowledge of system design and still coding abilities.
To democratize model generation we need a more iterative and understandable way of defining intented execution. The problem is this devolves into just coding the damn thing pretty quickly.
Being more clear and precise in our specifications would only benefit us and the AI/ML tool generating the code. We could lean more on the correctness built into the entire stack rather than having to proof-read a mess of inferred code, something we're terribly ill-equipped to do.
Good luck with that. We have those languages already. For example Idris. It's just that now you are essentially doing a lot of math.
And, funnily, I never hear people saying "making [math] more accessible and straight-forward to teach and use in day-to-day work". I wonder why...
I don't think that is really a matter of training. You have to start with people who can think clearly; if they can't think clearly, it's hopeless to expect them to produce formal specs. Very few people can think clearly about difficult subjects.
Of course, you can train people to improve the clarity of their thinking. I think that should be the main purpose of an undergraduate degree.
Writing a formal spec is analagous to writing a program; if you can't program, your program won't work. So writing a formal spec proves that you can think clearly; but if you can think clearly, you can write a program without first writing a formal spec.
More so if there are tools to check your formal spec.
But formal spec can be in higher level than the implementation. For example, it could be describing pre- and postconditions without actually stating out how to go from precondition to postcondition.
I have noticed programming languages don't tend to ask the writer the right questions in the same way e.g. TLA+ and TLC do.
I cannot do this, and neither can any of the people I have ever worked with. Yet despite that we all call ourselves programmers, create value and earn money by writing ill-specified, often buggy code. Why would a tool need to be formally verified to be considered impressive and/or useful?
You could if you wanted to. You're smart, inventive, and creative. There's nothing stopping you from learning.
> Why would a tool need to be formally verified to be considered impressive and/or useful?
Part of it depends on your perspective.
If we assume that an LLM (or some future ML tool based on it) is capable of producing code with the same rate of errors as a trained, expert human could then it would seem the productivity gain is not having to write all of that code ourselves.
We already tolerate a certain amount of errors in our software and the world has not collapsed. The JVM had an error in its binary search implementation that lasted for nearly a decade before anyone noticed. They noticed because the size of the arrays being used started getting big enough that their programs started failing in mysterious ways. OpenSSL had a vulnerability that sat unnoticed for more than a decade. The cost of errors is not zero but it is tolerated.
However the problem of programming is that we think it's our ability to write code which is the problem that is slowing us down.
My perspective is that we're not focusing on the problem: that it is hard to be precise and write programs that work, whole cloth, from their specifications without any errors.
Using an LLM to generate more code has another problem: while humans are decent enough at writing code to solve our problems, even if our solutions are imperfect, we're far worse at reading code and understanding what it does and whether it is correct with regards to some specification (if there is one).
Empirical studies of large-scale code review are very humbling. We can read maybe 200 SLOC every couple of errors and have a negligible impact on error rates in the software being produced. More than that and the effect disappears.
So now we have LLM's producing code that we know will have errors in it. And we have no idea where the error is. It could be a trivial error we could tolerate. Or it could be another Bar Mitzvah CVE. Hard to say.
Even Betrand Meyer missed an error in a single-line expression generated by ChatGPT. He's way smarter than me. I don't see how we'll be able to keep up.
But if we tackled the problem of getting better at being more precise with our specifications, I could definitely see how having an AI-like system automate code generation being really useful. There are plenty of times when working on a formal proof where you want to say, "this is obvious!," that have a machine verify that for you using the same proof rules and tactics you would use. Bonus points if it can explain the proof back to you.
I just think we're a long way off from being able to do that.
As for formality, real formal specifications are very hard, and LLMs are close to understanding natural language anyway, and 1000s of 90%-strict specs are better than 10 provably correct ones. So, some sort of legalese for machines will evolve.
"We are no longer particularly in the business of writing software to perform specific tasks. We now teach the software how to learn, and in the primary bonding process it molds itself around the task to be performed. The feedback loop never really ends, so a tenth year polysentience can be a priceless jewel or a psychotic wreck, but it is the primary bonding process—the childhood, if you will—that has the most far-reaching repercussions."
– Bad'l Ron, Wakener, "Morgan Polysoft"
for others as nostalgic as i am you should know about https://paeantosmac.wordpress.com/ for a bit more philosophical exploration
This sounds great in theory, but in practice, a system that has that much dynamic adaptation has brutally steep performance cliffs and is massively complex. I for one, will be opting out of that giant vertical slice of hell. This is one of the _good_ reasons for having layers: separate failure zones, separate levels of abstraction--true reuse and modularity. Bugs break all that.
And no, given the hallucinations of large models just in the natural language space, I do not want to reason through the mad ravings of a tripping AI to debug a monster pile that happens to make web property X go 10% faster.
I havent heard this phrase before.
Stacks where completely stand-alone applications connected which were not layers of the same protocol/system.
I think that the simple way is to refer to them as stacks now, but if you were raised on LAYERS - refering to OSI as 'stacks' feels foreign.
It appearsthat as we atomoze / containerize various teck, we now thing of them as 'stacks' rather than layers.
It seems that 'layers' are now 'services' rather.
In practice there's only 4 or 5 layers depending on who you ask.
I imagined stacks as tech1+tech2+techN
I just hadnt heard that phrase before...
30 years deep in ops.
> I have said before that I believe that teaching modern students the OSI model as an approach to networking is a fundamental mistake that makes the concepts less clear rather than more. The major reason for this is simple: the OSI model was prescriptive of a specific network stack designed alongside it, and that network stack is not the one we use today. In fact, the TCP/IP stack we use today was intentionally designed differently from the OSI model for practical reasons.
Yes, differentiable code is already a new paradigm (write a function with millions of parameter, a loss function that requires more craft than people realize and train). That has a property that used to be the grail of IT project management: it is a field where, when you want to improve your code performance, you can just throw more compute at it.
And I think that the clumsy but still impressive attempts at code generation hints at the possibility that yet another AI-caused paradigm change is on the horizon: coding through prompt, adding another huge step on the abstraction ladder we have been climbing.
Forget ChatGPT coding mistakes, but down the road there is a team that will manage to propose a highly abstract yet predictable code generator fueled by language models. It will change our work totally.
The point of DSLs are to provide a deliberately limited-scope language optimised for a specific problem or problem domain. LLMs that use general human language is like the furthest opposite of a DSL - its the broadest scope language for describing any problem, and they try to solve them all.
Also, few popular DSLs are truly blackbox in the sense chatGPT is - many of them have exposed source or even line-by-line debuggers available. There are a ton of other reasons this doesn't make sense to compare.
The same mentality, that causes today's "everything must be a web app", will caused terrible inefficiency in AI generated (and human prompted for) code. In the end our systems might not be more performant than anything we already have, because there are dozens of useless abstraction layers inserted.
At the same time other people might complain, that the AI does not generate code, that can be run everywhere. That they have to be too specific. People might work on that, producing code generators which output even more overheady code.
At least some of that overhead will slip through the cracks into production systems, as companies wont be willing to invest into proof-reading software engineers and long prompt-generate-review-feedback cycles.
And then it never happened. They focused on cloning iPhone features and approach, dumbing down and simplifying the OS to the point where it's pretty hard to distinguish any more.
Google missed the boat and focused on the completely wrong things and yet Android is at 70%+ phone OS market share.
No phone consumer would care whether the developers of their apps had AI as part of their development toolkit.
His single sentence caveat about how Neural Networks can fail in unintuitive and embarrassing ways is the understatement of the century. I’d like to add that Tesla still hasn’t solved that lane forking problem even eight years since it was first identified. I guess just throw more data at it, and eventually it will get better? At what point does the belief that things will get better with more data fed into the same algorithm become a religious creed?
Neural Networks are significant advances in the state of not just machine learning, but the world as a whole. But the caveat that we don’t really understand what they’re doing is the whole fucking problem. Until Neural Networks can take advantage of, constrain to, and augment human models, they don’t have a snowballs chance in hell at replacing the types of software we rely on the most.
Until then, you’ll just create a massively inefficient system where the neural network writes the software but you spend 10x on engineering your training datasets so that your brilliant neural network knows that it is better to commit to one of two lanes in a forked road than it is to crash into the concrete lane divider. Or to not be racist. Or to not go haywire because of a sticker on a stop sign.
Oh, ye of little faith! It is heresy to criticize our new religion! ~Some AI consulting firm or AI "thought-leader" probably.
Edit: I'm sure there are some useful use-cases, but I'm not an unquestioning devout adherent. That said, I should probably learn more about it just so I can intelligently defend the use-cases in which it doesn't make sense.
What matters is the percentage accuracy. A black box with a 10% failure rate is better than a fully explainable system that fails 20% of the time. Explanations make us feel better and they can be very important. But for most cheap, repeated processes they aren't necessary. Not to mention that neural networks can be tested and interrogated in ways that other systems cannot.
old adage: if a bug can be reproduced then it's only a matter of time before it's understood and fixed. If a bug can't be reliably reproduced (Heisenbugs) then repair time is unbounded.
(that said, humans are perfectly capable of creating inexplicable and irreproducible bugs - for example, in multithreaded code)
That's not true at all, it depends on the use case.
What actually matters is the desired percentage of acceptance.
For many critical path use cases you'd much rather have something fail twice as often but understand why it failed so that you can correct the issue and resubmit the input. Error observability is an important feature that's taken for granted in many systems. It all depends on what the system is used for—how important it is to be able to get to correct results, and what the consequences for failure are. The biggest danger of neural networks is in people that don't understand this nuance and apply them in a blanket way in all systems.
There is no way to tell the Tesla vision NN, "hey, when you see this pattern and you're confused about which path to take, it is better to take one incorrectly than it is to run into a concrete divider". We know exactly what the problem is, but there is no interface with a NN to tell it to do something, other than to just keep training it with more data. And once you realize that your only interface to get better outcomes is to wildly manipulate the training dataset, then you haven't made software engineering better, you've made data engineering worse.
Take notice of something important: all of the domains where Neural Networks have been wildly successful are domains that are wildly underspecified. Take language for example. Grammar rules, vocabulary, pronunciation, and even meanings of words are constantly changing. There is no possibility of ever having a formal definition of any language, let alone all of them. Or vision...where the only formal definition of anything is what color of light it reflects in a particular angle. Again, no formal definition of anything.
But the shortest path from A to B? That problem has a formal definition, and no neural network has even come close to the accuracy of A star or Djikstras. The minimum cost solution to a Multi-Commodity Flow Problem? That problem has a formal definition, but no Neural Network has come close to the accuracy of a Simplex method's solution.
Tautological arguments about percentage accuracy might give the edge to Neural Nets in some domains, but not all of them, and for that reason they completely miss the point.
1. Percentage accuracy is only half of the thing that actually matters. Without the cost of being wrong taken into account, percentage accuracy will totally fuck you over hard. Here's a game: you can choose between two algorithms, one that has 90% accuracy with a 10% chance of smelling gross, or a 99.99% accuracy with a 0.01% chance of your body been shaved down to bone over a thousand cuts from a vegetable peeler. Which would you choose?
2. Sometimes absolute accuracy matters. We have formal systems for absolute accuracy. We have symbolic logic for absolute accuracy. We have deterministic systems for absolute accuracy. If the best that I can get from a neural network is a percentage accuracy, then it has already failed a test of general applicability. How many years and how many computers and how much data would we have to feed into how big of a neural network in order to get to E = MC^2 with perfect accuracy?
3. Even if percentage accuracy matters, time-relative accuracy matters even more. With a Neural Network, if you need to get better accuracy, how do you do it? You should see actual machine learning practitioners try to solve these problems. They literally try to deconstruct the black box, trying to figure out how different neurons are weighted, and what input data can be altered to result in a different weighting. It's a clusterfuck, and it slows progress to a halt. We've known exactly what was wrong with the Tesla Vision NN for over half a decade, but actually fixing it has completely stalled because of the fact that it is a black box, and can only be fixed like black boxes. This is systems theory 101: you can't fix systems that you can't understand.
If the "actual machine learning practitioners" can break down behavior to the neuron level, then I'd say they understand it to some degree.
The reason that many of these problems still exist is because chasing down individual errors is a seductive waste of time. If you have to tell the network "here's how you handle this one pattern and here's how you handle this other one", then you're building an expert system. Yes it's tempting to correct errors as you see them pop up. But it's better to construct a network that can learn from data to handle any situation that you didn't think of specifically.
Building a robust network requires lots of time and data. There will always be edge cases that cannot be fixed individually. The fact that Tesla or anyone else hasn't built a system for X with no embarrassing edge cases yet does not mean that we should go back to coding individual instructions and conditionals line by line.
“Yes it’s tempting to save and invest your money for your retirement, but it’s better to put your faith in god and he’ll solve all your financial problems, even the ones you didn’t know you’d have.”
If the system, with all of its data, can’t solve a common problem that happens every day, how the hell is it supposed to solve a problem so rare that the engineers don’t even know exists?
Much more so than fixing the 10% of a neural net.
https://twitter.com/karpathy/status/1623476659369443328?lang...
Software 2.0 (2017) - https://news.ycombinator.com/item?id=23766796 - July 2020 (22 comments)
Software 2.0 - https://news.ycombinator.com/item?id=15678587 - Nov 2017 (36 comments)
"You’ll notice that many of my links above involve work done at Google. This is because Google is currently at the forefront of re-writing large chunks of itself into Software 2.0 code."
IMO google is dropping the ball pretty hard right now when it comes to AI.
In the same sense, in the future we will be wiring together AI APIs (probably because it will be cheaper to wire together manually N AIs than to write one that is the sum of the N AIs). Since we’ll be able to do more, user requirements will get more complex… and so the demand for software engineers will go up as well. In the future only a couple of engineers will be needed when today we need 10.
Yes, a lot of software will include NN models. Traditional software is going nowhere, because it's the only means of being 100% sure of what the outcome will be, non-probabilistically.
Neural Networks are a tool for solving probabilistic, fuzzy logic problems.
And then, as you say, there will be certain parts where those models are actually gonna be integrated in software in one way or the other. And I think this is powerful. It would be awesome if I can just toss certain problems to the business folks and empower them to figure out the solution AND implementation by themselves.
But even that will probably take quite some time.
For example take a website. How are we going to provide enough examples of websites to make the code generated fit what we need and not have annoying properties we want to void? Lets say we have a website and we tell the code generator of choice, that we want that website to be accessible for blind people. How do we create the amount of labeled examples, that make the code generator understand what to create? Maybe that very creation of labeled examples will be a software developer's future work activity.
The weird part is presenting it as 2.0 which implies that it replaces 1.0; it doesn't, apart from some edge cases like image recognition - we don't hand code rules to recognise images anymore but that's a tiny tiny part of all the software development work out there.
If you surveyed e.g. all of the code Google has in their piper repository, you would find significantly less than 1% of it could be replaced by even an extremely good neural network.
(Unless you're talking about making GPT write the code that does all of the above – I'm sure people are working on that – but unless I'm completely misreading Karpathy's article that's not at all what he intended; IIUC he's talking about programmers making datasets for specific problems + simple neural net architectures to train on that dataset. If ChatGPT wrote some Java code for you, you can't make it go faster by removing half the nodes, which nodes would those be? It definitely won't go faster if you send the same prompt to a lobotomised ChatGPT)
Personally, I want to keep working on things that I can get to the bottom of. In the same way proprietary software is a nightmare to debug, having an AI blackbox in the middle of my stack could wreck all traceability. However, I can see myself using an AI blackbox when its output is consumed/checked by a human or when "best-effort" is good enough (but treat output as dirty).
Example: sorting documents by relevance (human consumes), code assistant (human checks), transcription of my audio notes (wouldn't do it myself or pay anyone for that, so any output is good enough),
Counter-examples (too dangerous!): AI personal assistant that accept/reject meetings, a ChatGPT text box as an interface for settings, auto-generated tests, Infrastructure-as-a-Desire: input your software and get a new K8S cluster provisioned for it.
Nice one Andrej!
So yes, this will revolutionize and enable unseen performance in the few areas where there is significant data. For all the rest it'll be business as usual.
Going to push back against this one. I think we have a lot more training data than most people realize. I wrote this comment yesterday, https://news.ycombinator.com/item?id=34862450, about how a large government contractor is using ChatGPT to generate first drafts of responses to government RFPs.
Now, most of these RFPs are in very specific areas, technologically speaking (e.g. specific technologies around cloud network security, for example). These folks were actually blown away by how technically accurate ChatGPT was on many different areas, even very specific niche areas, and even considering ChatGPT's view of the world hasn't been updated since late 2021.
Again, the first draft needed to be edited, but there are is a huge amount of data out there that ChatGPT is able to use coherently on even niche, esoteric topics.
so yes, sure, the best speech recognition algorithm will remain something ee-ish, and will probably make use of highly parallel numerical computing and data driven optimization based solution finding... but i think whether or not that will be the road to correct implementation of entire discrete information systems with all of their knotty discrete rough edges remains to be seen...
So "programming" in the Software 2.0 world is basically training a model. Which isn't a task with an end. You stop when you're bored, or when it passes a given level of accuracy, not when it's "completely trained" because it will never get to 100% accurate.
In e.g. speech recognition, handwriting recognition, speech synthesis, drawing pictures, writing a response to a human's question, all those messy "organic" problems, this is fine. A 99.9% accuracy rating at e.g. speech recognition is better than humans do.
But there are problems that absolutely need 100% accuracy, and you can only get that if you code it up the old-fashioned way (though probably not Agile - I love my iterative development cycles but they're equally prone to not quite getting it 100% right).
but... yeah, some problems have continuous performance variables (typically ee'ish) and others have discrete ones (typically cs'ish). the discrete account ledger either computes the correct value or it does not, where many signal processing problems have to contend with noise and are allowed to produce a wide range of noisy outputs and therefore their measures of correctness are based on statistical arguments.
And how do you show that it is 99% accurate besides creating enough automated tests to the point that you could write the procedural version?
I think what I was missing from this article is how to evaluate a domain where neural nets or LLMs can be applied. Image-from-text generation is a great one because accuracy isn't strictly defined. However, telling ChatGPT "code this pacemaker for me" would have a real accuracy attached to it that you could confirm with unit tests.
[1] https://twitter.com/karpathy/status/1618311660539904002?lang...
- I never run this query : I have found a super Google, but maybe I will miss out some pearls available on the Internet;
- I always run this query, but I'm too lazy to have text files, bookmarks in my web-browser to stock answers that I don't want/cannot memorize.
It is somewhat similar to the arrival of Google, Unity UIs, Gnome 3, Windows 8, where they wanted to replace all menus with a single search.
It is the way forward to idiocracy if we are no more able to tidy our ideas/data.
It is like when we are child and we always ask our parents instead of look it up in the dictionary. It is regressive. It is already the case with Google. Since I can look it up on the Internet, I memorize less things. It is human to go a regressive/easy path without will.
I'm looking forward to a blog post "Use AI without declining mentally" :
- Interact with AI to tidy a corpus of tools/knowledges in personal files
- Delegate to the IA the tidying and research of pearls you keep in your folders
- BUT have a shared mental model of your data and their structures with your IA, like a manager may had a shared knowledge of the way the files were sorted in boxes and furniture with her secretary. Like this if your IA/secretary is in holidays, you can still work.
So when you're looking at actually writing software that needs to be dependable / modifiable / bug free, you'd need a massive overhaul of whatever software stack is being used, so there's very little human-assisting "cruft", and instead you'd want a lot of supporting material for a model, which might look like something written in languages used for formal verification of programs.
The promise of GOFAI was about having a human-understandable bottom-to-top framework, and the current "AI" paradigm is at odds with it. The "formal verification" assumption, then, skews towards GOFAI. But since there has to be some human support for the current not-there-yet AI to write software, we might see yet another abstraction layer based on NN / something newer in the years to come.
Have you _used_ ChatGPT? I mean not just asking it random factoids but using it to genuinely help you with something. Are you aware that it's hit 100 million users faster than Facebook, Instagram, or TikTok did? It's not a perfect product but it's hard to argue with those numbers. I work at a startup and most people I work with use ChatGPT daily. I'm talking project managers, devs, personal assistants, etc. I guess that's all to say OpenAI is influential as heck _already_, now imagine 5 years down the line if they play their cards right.
I have, and I didn't find it to be useful for anything I did. I can do what it does with a search engine and trusty C-f. Also, TTS exists.
It's 50% tech and 50% marketing (and I doubt it's 50% tech at that), it's not gonna upend anything. Except maybe increase the authenticity of online scams and make people get more degrees in machine learning. And yeah, make the people that rely on it bound as it degrades their skills.
It's basically the "internet is educationally useful" argument. At some point everyone's gotta use it but you can live without it just fine. And even though people tout its usefulness for everything good, the majority of data transferred is porno.
Because it was covered on the news and social media incessantly. You could get 100M people to follow a taco stand if it got a free month of news coverage which breathlessly covered the ingredients and had literally thousands of Tiktok and Youtube grifters telling you how the chalupas were gonna change everything.
This also reminds me of a quote from the Book I'm currently reading (Practical Wisdom by Barry Schwarz)
“Most of us think about empathy as a “feeling” or an “emotion.” It is. To be empathetic is to be able to feel what the other person is feeling. But empathy is more than just a feeling. In order to be able to feel what another person is feeling, you need to be able to see the world as that other person sees it. This ability to take the perspective of another demands perception and imagination. Empathy thus reflects the integration of thinking and feeling.”
"Mind reading" is another way to put it (https://yosefk.com/blog/people-can-read-their-managers-mind....) - this practical wisdom + mind reading is basically the salient human feature that NNs would never be able to replace so you would always have humans in the system.
AI will take incomplete specifications and guess the rest -- just like humans do. Whether or not it makes those guesses better than a human remains to be seen.
"The End of Programming"
https://cacm.acm.org/magazines/2023/1/267976-the-end-of-prog...
Discussion:
The pod was in the past year, so many years after Andrej's software 2.0 post, and after many years of great AI experience at Tesla to add to or potentially change his views.
Likewise, will be interesting to see how the HN community's experience with ML and AI over those years may have changed our views.
Imagine if the Therac-25 software was written by chatGPT.
I guess in that case we wouldn't learn and teach from the design flaws made. Instead we would "It's just a glitch. No one is really to blame. Just feed it more data and maybe it won't kill anyone next time".
Having been partially responsible for a (back then) SVM based machine learning system and seeing how it's difficult to explain to management why it fails and why fixing it isn't just a missing line of code somewhere was pretty frustrating. I'm not sure I like this future.
Me too! But to allay our fears I think that the OP here is saying all programming will change to some NN powered large language model. That is not true. There will still be "manual" i.e. not NN powered programming and I suspect that it will be the case that type of thing is the majority for the rest of my career.
Good luck to the poor sods in 100 years time arguing with a poorly trained LLM to output some unit tests whilst a virtual chatbot runs the standups.
Unlike the Excel workbooks made by domain experts, it's almost impossible to even find out what the function of any given cell is in the overall computation. See the effort it took to find the "neuron" responsible for a/an differences in GPT-2.[1]
Neural networks have their roles as a black box, but they are not programs, constructed with intent by humans, to be read by other humans, and compilers.
[1] https://www.lesswrong.com/posts/cgqh99SHsCv3jJYDS/we-found-a...
The quality of the code is irrelevant, as the point of Software 2.0 is that it's another layer of abstraction on top of traditional code.
"Coding" becomes "I need something to do a thing," rather than "def doSomething: ..."
As long as the output gives you what you need, the code quality ultimately is an efficiency play. But as AI coding improves, it can refactor itself, so it's a short-term problem.
In my own experience coding with an "AI assistant," I've been able to mentally stay in "architecture mode," which makes me feel twice as creative, twice as productive. That alone is a net positive.
Until it needs to be maintained, or has weird bugs.
> As long as the output gives you what you need, the code quality ultimately is an efficiency play. But as AI coding improves, it can refactor itself, so it's a short-term problem.
Not sure how this is going to work on large codebases.
> "Coding" becomes "I need something to do a thing," rather than "def doSomething: ..."
More likely, corporate overlords will decide that you cannot just "do a thing" but rather that you are allowed to do X, Y and Z things for which they have pre-trained commercial models for.
> As long as the output gives you what you need, the code quality ultimately is an efficiency play. But as AI coding improves, it can refactor itself, so it's a short-term problem.
Have you ever debugged a problem with generated source code? Or even a compiler bug? Now imagine leaving your AI to go find the bug or iterate until the bug disappears hehehe...
> In my own experience coding with an "AI assistant," I've been able to mentally stay in "architecture mode," which makes me feel twice as creative, twice as productive. That alone is a net positive.
IMO, if your work benefits from an AI assistant then your work is to produce many lines of code and you would benefit equally from creating high-level abstractions than from using pre-trained black box models (or as some call them "new hires").
Neural Networks: https://www.youtube.com/watch?v=cdiD-9MMpb0&t=58s
Language Models: 41:50 https://www.youtube.com/watch?v=cdiD-9MMpb0&t=2510s
Software 2.0: https://www.youtube.com/watch?v=cdiD-9MMpb0&t=3944s
What I find interesting in this imagined future is that problem definition usually happens, in my experience, while attempting to encode it, removing all ambiguity. If we skip that step n years from now, will we still even understand the problems we try to solve? Sounds scary to have systems where we can neither reason about solution nor problem.
Instead we got this LLM-based "paradigm-quake".
More interestingly, this makes me wonder if there are some Gödel-like proofs waiting out there that limit the capabilities of efficiently-optimizable programs. What new kinds of undecidable or uncomputable functions exist in the subspace of programs that an NN can learn? Would be exciting to find out.
The next software as a differentiable thing that is the program has certainly failed. But now, there are amazing opportunities to connect text-to-text models to other another, to search engines, that it is likely to become a new programming.
First there was software of the Enigma machine variety, then assembly, then massive IBM machines running COBOL, then C, then Ruby/PHP/Python. We could also talk about how networking and persistence fundamentally changes software. Each of those is a big iteration in itself, probably just as big as moving to ML generated code.
Nothing is really a silver bullet so I guess the future of programming is really hybrid. Stuff like Github's Copilot.
Probably not, although I'd be more than happy to be able to delegate work to our robot overlords.
At the end of the day it is always: garbage in, garbage out.
Ok.
I feel like I'll have to install something that will make life easier with Software 2.0.
You feed it loads and loads of data, where the neural net basically compresses it all into the structure and weights.
Then you give it some (un)compressed part, and it gives you the other.
For what it's worth, "except it's really happening this time" has also been said before...
The difference is a slight loss of emphasis that had been meant to show chatgpt doesn't require many prompts in order to convince the model to answer the situation that you had posed. The word "happily" wasn't used in the sense of chatgpt experiencing emotions
How about a more glass half full take on progress?
So to cast this as Software 1.0 vs 2.0 doesn't make sense. There is a class of problems where neural networks work better. Everywhere else we will continue to use traditional code.
still, hoping that the future predicted here is 25-30 years out
Yes there is a lot of hand-holding and guardrailing but still, it's kind of insane.