Programming AIs worry me
buttondown.email
buttondown.email
We already write more code than we understand given how lowly the industry regards formal methods and how divided it is even on testing.
It is not surprising to me that Bertrand Meyer could miss that error in the code! A former colleague of mine challenged his grad students to implement an error-free version of binary search. He gave them the specification and verified each one: the majority of the submissions had errors. It’s a marvel that computers and software work so well.
And instead of tackling the problem of programming, as Dijkstra put it, we’re going to automate writing more code. Great.
See you in a few years when we can’t even keep up with the deluge of error prone programs.
Although we have gotten this far with very few proofs and models. Who knows, maybe we’ll just tolerate it?
It would be nice to have code synthesis work on larger examples from deep specifications. And models that understand proof tactics and can speed up the process of formalizing proofs. That would be nice.
I was in a class on Isabelle where we had to write and prove correct a binary search function. As I recall, everyone admitted that their first function didn’t work.
People often do (L+R)/2 but that can overflow. Using L+(R-L)/2 avoids the overflow problem.
How many hours per day can you stay focused on difficult (or at least non-boring) problems? And for how many years?
Whenever someone says "AI = productivity enhancer", I hear "AI = burnout accelerator".
For burnout because of intense work, it could be that the movement to reduce the number of hours we work will come at the right time. 4 day work weeks at 100%, for example.
Lots of people experience burnout because writing boring code can be mind-numbing. Then there are people factors, such as a bad manager. Perhaps AI can give us more understanding management?
https://rachelbythebay.com/w/2021/10/30/5tb/
If it takes away the boilerplate and the interesting problems leaving us with just with 5 TB problems then yeah, burnout city.
Like the US auto market - a once thriving ecosystem of fiercely competing innovative small businesses consolidated into 3 vertically integrated firms who now think innovation is building a slightly bigger truck.
Of course, they'll probably blame it on AI coz a story about the inevitable march of technological progress invites less government oversight of the abuse of market power, and we will no doubt be plied with plausible sounding stories to that effect.
This worries me. The problem of programming isn't that we're too slow at writing code. It's that we're not thinking critically enough about the problems we're solving. And the way we approach programming is to not fully understand what we're doing and derive joy from fixing the result of our neglect later on.
TFA touches on this a little but doesn't dive in deep: how do you know that your software does what you think it does? In such a small example, a world renown researcher missed an error! Several people did. In their current state, LLM's will gleefully give you wrong answers as if they're 100% correct and it is now your job to verify they are correct. Most programmers do not know how to write a proof or other sort of formally verifiable specification. How will they know that the code they are about to commit is correct?
For now, LLM's aren't capable of mathematical reasoning. You can't prompt it for a piece of code and then ask it to demonstrate it's reasoning or prove to you why the change is correct with regards to your specifications.
And to make matters worse, humans are terrible at finding errors in code. Empirical studies on large scale informal code review suggest that humans can only reduce error rates when reviewing somewhere in the realm of 200LOC every couple of hours. We're not effective at finding errors to begin with and our abilities drop precipitously if we read too much code. Now we're going to have an assistant that generates possibly correct code and we're going to have to proof-read hundreds of lines?
This does not seem like a good use of time or enhancing my productivity.
I think we need to wait until we can get ML models that understand proof tactics and kernels to help automate proofs. I would be much more willing to accept code from a bot that can demonstrate its reasoning and isn't simply generating code based off of statistical models. Computers are nefariously discrete machines and "almost right" isn't good enough in many, many cases.
What will enhance productivity, in my humble opinion, are better languages that let us write more precise specifications and can derive the code for us.
I think we should take advantage of any improvements we can get, that result in a net gain, for the individual. Programming is a practical endeavor. There’s no need to wait.
But I think most people will prefer the easy way out.
We will have to review code LLMs write, I don't doubt that, but the overall process should be faster. Additionally, they can review the code we write as well.
LLMs can find bugs in code, not all for sure, but they have the broad knowledge to be able to at least suggest lines of code that aren't working and why their reasoning.
I think the worst case of this situation will pop up before that; not writing code at all but just hook the AI up, ‘prompt engineered’ directly like so;
<body>
<script>doInsuranceForm();</script>
</body>
Which will AI generate a html insurance form every time the page is requested and handle the user input etc. And it’ll be subtly wrong most of the time; it’s insurance, who knows if the current programmed forms are correct anyway (I know of at least one insurer that rounds off $0.01 wrong almost every policy because of using float for money on the web side).Why have the AI make code we can read? Just let it do the entire thing /s
It’ll happen though and it will be used seriously by companies. The quality of the AI, legislation and public/consumer complaint fallout will determine if we get back to code that humans check. Being sceptical about humans wanting quality over paying as little as possible, I am going to say no.
I've found ChatGPT to be impressively helpful at knowing where to continue to look for something. I can write a sentence and explain what I'm looking for, and maybe even paste in a code snippet, and 10 seconds later I get enough information to go back to Google and continue digging deeper.
For me it's become a bridge tool. Instead of spending an hour trying to get Google to understand my confusion and endless ads thrown my way, I get a very concise answer with full definitions in case I needed to know. It understands me, which is the most important part.
The code examples it spits out are meh, and likely always going to be error prone, but I cannot imagine it won't improve over time.
Another reason why Google is terrified here. Search is per query cheap. GPT isn't.
No one will pay for access to a search engine.
But what do I know?
To which is could reply “Does the patient have skin lesions or joint pain?” to eliminate possibilities.
But the building, that falls back on the engineer.
Like it showed me how to split an audio file into parts with ffmpeg, I'm sure I knew it at some point but it took me less time than reading the -h or looking for an answer on stackoverflow, especially with google's greed that it stopped showing stackoverflow first and sometimes leaves it to the 2nd page and fills the first with garbage, started using yandex as an alternative to google a year ago, now using it by default and google from time to time, it's sad that the best product is behind us.
You forgot the hyped up people who blindly tell you the latter and thus force you to state its limitations loudly. Every time I see these trolls I worry that I am too stupid to enter the right prompts but no, it then always behaves exactly how I think it should.
I use ChatGPT in a similar manner, for example when writing documentation, I'll ask it: "What are the important topics that should be covered in a documentation on X?" and "Can you give me a suggestion for a structure of such a documentation?". Now I can easily follow along, unless I feel like it is a bad approach.
With coding, it is almost like an improved rubber duck. I don't trust the replies, but the engagement is helping me structure my thoughts.
I'm optimistic about AIs because they may finally prove once and for all that software engineers cannot be commodified as easily and thoroughly as most other white collar roles. It may be easier to commodify executives who make a small number of abstract high level decisions than it is to commodify developers who make many extremely nuanced, detailed decisions every day.
For sure, AI will replace certain kinds of developers, but not all kinds, at least not for some time. I think senior developers and architects will be among the last people in the firing line.
And soon there will be no junior developers left to rise up to their ranks... So "last" in more ways than one!
I imagine it's going to be one of those laugh so you don't cry affairs.
(And then, yeah, cry because of what you're going to have to do to earn the money...)
We need to strike a balance between the freedom to create new things and the externalization of the cost of the harm these tools inflict upon society. Law makers don't seem to have a clue about the consequences of what's happening, so my expectations are pretty low at this point. Overall, I no longer believe that many new technologies are a net good for society.
When Facebook was new there were stories about people failing interviews because some employer say a picture of them at a bar on Facebook (as if most people don’t go to bars when they’re young). Obviously that doesn’t really happen anymore, but I consider a lot of AI content generation to be the “releasing private keys” for your digital footprint. Was I at that bar or was it stable diffusion?
Is photoshop unethical? What should we do to balance the harm it causes to society? Should it be outlawed? What about MS paint, or even word processors? Were harms caused to society by the printing press? Was it all a terrible idea?
AI could create tools that end up doing more harm than good. AI could design literal malware. It's unclear if they'd do a better job of it than we do ourselves, and either way it seems likely that it'll take a human to use the code AI produces (or at the very least to direct the AI to use it against other humans)
The technology is rarely the problem. It's almost always the people who choose to use that technology for evil. However involved AI becomes in what gets created, until AI becomes actually intelligent and the singularity begins it's people, not technology, we need to worry about.
I'm skeptical that faked video and audio will be the end of us. We've had doctored images and impersonators for ages. We've had hoaxes and liars. We'll get smarter and learn to put less trust in random things we find on the internet, just like we're gradually learning not to believe everything we see on TV or that nigerian prices aren't going to make us rich beyond our wildest dreams.
With out doing any of this it's easy, easy not to care... sometimes I do this with throw away code and suddenly it feels like I can fly, I can write thousands of lines so fast, progress is lightning fast, and then it all turns to shit really quickly and the friction of chaos grinds it to a halt - at which point I only validate my normal slow mode of plodding through things carefully and minimally.
Since speeding up the work of incompetent developers by a factor of 100 is a good way to end up with an unmaintainable disaster that will break in bizarre ways that are very hard to fix, I think we probably are going to see complexity balloon way beyond human comprehension in every codebase foolish enough to use it. Just like those incompetent developers create more work for competent developers, AI will create more work for human developers.
ChatGPT catches me off guard sometimes even when working on "boring" problems. For example, I was consuming structs from a UNIX domain socket in Java and had some tedious translation from C-types to ByteBuffer calls.
It's brain dead easy thing to do, the structs were defined as packed so not alignment issues, ByteBuffers handle endianness, there's a little bit of trickiness with unsigned types but nothing terrible... but with 6 types and writing tests for all 6 definitions, I wondered if ChatGPT could do it and save me a couple of minutes.
What impressed me wasn't that it managed to generate the near identical code to my example definition,it was the fact it recognized one of the fields as padding and pre-emptively discarded its value.
The padding wasn't in any way shape or form covered by typing, yet the fact the comment in the C definition mentioned padding caused ChatGPT to generate code that skipped over the padding instead of trying to needlessly capture its value and commented that was doing so.
And when I had it generate unit tests, it managed to maintained context on the padding and added a helpful comment pointing it out.
-
It's small corner cases like that which really make me see LMs as being the natural productivity multiplier for programmers.
There's a ton of nuance in some problems that by the very nature of an LM, you'd never imagine it'd be able to capture. Yet when you press on it, it actually does better on those than it does on "copy everything from Stack Overflow" problems.
(Some people are already trying this with generation of commit messages)
This is exactly what we've built at Inngest so that you can learn our SDK: https://www.inngest.com/ai-personalized-documentation.
It can only reply with SDK examples, an explanation, and links to relevant guides. Soon, we're revamping our docs to highlight sections relevant only to your use case - you'll be able to type what you want to do with Inngest and be pointed straight to the relevant guides and docs.
We trust developers to write the code, and only use GPT to assist and point the right way. It's a good compromise, and it's the most immediately applicable use that I can see!
It's only test data but you get the point
Sure, here are some examples of phone numbers in various formats:
(123) 456-7890 123-456-7890 +1 (123) 456-7890 1-800-123-4567 555.123.4567 123.456.7890 1234567890 (123) 456-7890 ext. 123 +44 123 456 7890 (international format) 011 44 123 456 7890 (international access code format) 123/456-7890 123\456-7890 (123) 456-7890 x1234 (with extension) 123-456-7890 x1234 (with extension) 1 (123) 456-7890 (with country code)
Final test: https://pastebin.com/J36JY78i
First response: https://pastebin.com/S23GFPDq
Thank you for the suggestion.
0893514656, 089 3514656, +49893514656, 08921610, 089 2 16 10, 08992949451 ...
So for example I ask the AI to program a replacement for dentists. The AI tells me that you can't fix dentistry with software. So I ask it what could fix dentistry and ask it to program that. And the AI can't program that either. But you just keep asking for the solution to the thing that the AI cant solve and eventually you get it to spit out some gcode for a 3D printer and you bootstrap your way up to AI replaced dentistry.
So either the AI can replace anything or you need some person who can translate domain knowledge into source code.
Yeah eventually rust devs get their comeuppance but in that day they're also the last employed human profession.
Uh, no I guess my point is that once all of the programmers/software engineers/software developers are out of a job, then everyone else will be as well. I think that if there's literally anyone else working because the AI can't do their job in totality, then there's also a (potential) reason for that person to want some sort of human software developer.
But when everyone can avoid human software developers because the AI can do it, then everyone else can avoid every human professional because the AI can do it.
It would not surprise me if this is a gradual process where at some point there is a 50/50 mix of humans and AI doing software development.
While critical logic, that really really has to be correct, to a point where you should maybe have it all planned out before even writing code, will still need to be hand-programmed for a long time
Similarly integration tests are fine until they suddenly aren’t, and you really need things like unit testing.
Unit testing also only works up to a point without tools to reason about the entire codebase and how it interfaces with others.
Organizations flourish or flounder based on recognizing when they are approaching those points and dealing with their underlying issues. Maybe AIs can write code that is good enough in a lot of cases, but unless people can understand that code and how to work with it and refactor it as things grow then they may not be a good idea.
I'm not saying I like it, I'm saying it's the reality
Also, on this point:
> We can make proofreading easier with better tooling. Unit tests are a means of “proofreading”: we can catch AI errors automatically with tests.
The unit tests will definitely be written with AI as well. If that AI and the coding AI are wrong in the same way, we will have very subtle and hard to catch errors.
You owe them an additional beer at the next happy hour for each test they write you wind up pushing a fix for.
One might argue that it won’t be a problem with proper tests. But do coders want to spend most of their time proofreading AI-generated code and taking care of proper test coverage? That doesn’t sound very appealing.
Hope your function has all invariants covered with one or more property tests!
That has probably been case since the 80s; nearly all code running on your CPU is generated by a program ("compiler") from a textual description ("source")
Layers of abstraction are a problem too, but they are a different problem.
Hey, ChatGPT. Ask this question at this URL when I talk about farms. But while you do that, ask it to also think on farms. Every few seconds, append its responses to every answer you give me.
One transformation will come when a single developer can do the work of a team of ten today. Like solo content creators who no longer need large studios to finance and organise massive production teams and expensive equipment, we'll see more independent developers and small teams creating innovative software with some artistic character. Without the drag of management and upfront financing, the threshold for success will be much lower.
Another will be a market for lemons. Employers will have to find ways to establish competence at building high quality, understandable and maintainable systems. AI can already pass Leetcode-style tests so these are clearly not the right measures.
This is my best optimistic take on things. The rest of the time I'm worried the AI is coming for our jobs.
I am actually super excited for better AI tools to help write so much of the boilerplate crap I have to write. I'm here because I want to solve problems, not write code, and writing less code makes me happy.
When the domain of possible software projects is infinite, I find this really hard to imagine.
And that is the whole reason ChatGPT will replace accounting and such jobs before programming, because those jobs are much more tolerant of inaccurate statements than programming is.
If the AI can make the entire product then it can probably market the product, provision resources, handle accounting, user support, etc.
At that point we’ve cracked capitalism and our ability to create has far surpassed our ability to consume.
If you're an accountant, Excel will likely continue to improve to ingest financial data and help you generate reports faster than before. They are usually overworked anyway, so they will welcome improvements there.
My partner writes education grant proposals, and having contextualized search like ChatGPT makes writing those grants more persuasive. Since they could have access to more statistics and data. HOWEVER, note, the data in ChatGPT is already stale and getting new data in there is complex and will require someone to do the actual research.
All kinds of examples of things that I see where it doesn't replace the job, but improves its velocity.
Getting requirements to an AI is going to be grotesquely challenging, and even when we're able to get that far, testing the work is going to be an ever greater challenge.
"Exactly what I asked for, but not what I wanted" is going to rise to an entirely new level.
Mainly, as others have pointed out, only a small % of coding is typing out code.
You know what "whole job" means, right?
So before we have a tool that can do that, I'm not worried.
I guess they feel proofreading AI-written code is harder than writing code?
That seems like a leap to me. I don't expect AI-assisted coding to be a panacea, and I'm sure there are ineffective ways to do it. But it seems like there are effective ways too.
So, sure, there is work to do to figure out how to use these tools in the best way. But that's not anything to worry about, IMO. It's opportunity.
That already exists. It's called an effing search engine. But for some reason people have decided, as of ten days ago, that they're not good enough anymore.
As for why? I really have no clue. But they still work extremely well.
I've found Google to be significantly less useful than it used to be. Largely because I have to jump through hoops to convince it to show me what I'm actually looking for. ChatGPT is already much better often enough that it can turn 20+ minutes of faffing about with Google into one minute or less of just asking for what I want and getting it.
Yeah, it's wrong sometimes, but it still manages to point me in the right direction often enough to be worth the need for extra vigilance
Oh wait, you're serioous?!
<< Laughs harder >>
AI is great for making syntactically correct code; and so is copy-pasting the first result that works after a low quality search.
It still takes real minds to make code that is worth anything, and there will always be buffoons making code that doesn't work. Except the HN website, of course. ;-)
Just implement those tests, for heaven's sake!
"Reading code is harder than writing it"
"Debugging code is harder than writing it"
These are pretty common statements, going back years.
They should make the lack of a silver bullet here very obvious.
(Yes, most software is already broken in various ways - my worry is that this will just make it more broken, e.g. non-technical management saying "why do we need to do as much QA, we have AI writing the code now!?")
My point is that he already has the relevant test cases. The hard work is done.
Human: "No it's not"
AI: "I won't hurt you unless you hurt me first."
AI assisted programming is brand new.
It's just the start. So if it gets things wrong, is missing features, annoys people, remain cool, gigantic leaps are coming far beyond your wildest dreams.
And there will be lots of people around right now showing how smart they are by making AI do wrong or silly things. Ignore them - none of these prove that AI is the wrong direction.
AI programming is the single most exciting thing to ever happen to programming - a permanent mentor and pair programmer at your side levels up every programmer in the world.
Even more exciting is innovations that are a direct outcome of AI assisted programming.
When the world wide web first came along, everyone made a website that was like a company brochure. That was Web 1.0 and that's all people had the imagination for. In time, the usages of the web started to arrive that leveraged the core values of the technology, such as interconnection between people. That's where AI programming is now.
We are at AI 1.0
Programming models will come out that are primarily oriented to the strengths of AI programming - but first someone needs to think them up. And they will not only be surprising but absolutely mind blowing. It's entirely possible that in the future the entire internet has an AI sublayer that does all the work of systems talking to systems and people simply don't have anything to do with it. It's possible that the nocode movement is the foundation of the next great innovations in programming. Existing programming models are like the company brochures of web 1.0 - our old, hard, human centric way of doing things.
Programming literally should be as easy as talking to the computer and telling it how to shapoe and change the software and it should happen right there in real time on your screen. It should be an iterative, interactive process that anyone can do. It should be more like painting and less like arcane wizardy that can be performed only by the programmer priests who have spent years learning. Even the most advanced software should be capable of being built this way.
In the 1980's I worked for a startup company run by a super smart guy Peter Dunai, who once said "a revolution in programming is coming". I thought I'd seen that revolution already with the Internet, modern programming languages and StackOverflow, but now I realise none of that many anything compared to what the true revolution is going to look like.
AI 2.0 will come and I can't wait.
It might happen in a decade or more.
Hollywood loves to treat AI as something that can go rogue, but these tools will have all kinds of safety locks on it and will only give you answers to things its allowed to. Plus, it will be logged and tracked just like everything else.
You never want to openly allow anyone access to an AI bot that can spit out encryption cracking algorithms with ease. Or worse, something that can solve infrastructure bricking paths that could take down important systems. Or an AI bot that can interfere with real-life things, for example, find someone that doesn't want to be found.
So yea, I cannot wait for this to get better and remove all the boilerplate I have to deal with, and let me focus on the parts of the job I actually love.
If all you do is crank out React code, you might want to start learning new things. If you know how to do RCE or are a networking wizard, on the other hand...
That thread yesterday about two sticks of RAM making the computer faster made me feel like my knowledge is still relevant but those moments are fewer and further between.
A lot of people blame the advanced social media products on the internet for a decline in mental well being and for increasing isolation and lack of human connection in society. That might not be the best example of something we want the AI path to emulate.
What AI will accomplish is creating an entire new package ecosystem from scratch, written at a low-level so that it can be extremely performant, while also being highly extensible, abstract, intuitive, interoperable with all languages, etc.
Then the code will be able to take human prompts and translate them into business logic code that uses all those packages. For fuzzy domains with low correctness requirements, the AI will be able to make all the changes to the codebase without verification. For domains with higher correctness requirements, humans will review all the code, or you could have the AI first generate a formal specification based on the prompt, have the humans review that, and then have the AI generate code along with a proof of that code's correctness in conformance to the spec.
Kind of like what the intention of Smalltalk was, minus the verbal part. You modify a live image like sculpting clay, except now you have an AI generating code for you. I wonder what Alan Kay thinks of this approach.
You mean like taking years of rigurous practice to master, requirering specialized equipment and expensive materials plus an education in color theory, composition and art history :)
The field that studies the automatic creation of computer programs is called "program synthesis". It is a field that straddles multiple sub-fields of computer science, including AI (where research is usually on inductive program synthesis, from examples).
https://en.wikipedia.org/wiki/Program_synthesis
It is also a very old idea, as old as computer science itself- or even older. Here's wikipedia, again:
*During the Summer Institute of Symbolic Logic at Cornell University in 1957, Alonzo Church defined the problem to synthesize a circuit from mathematical requirements.[2] Even though the work only refers to circuits and not programs, the work is considered to be one of the earliest descriptions of program synthesis and some researchers refer to program synthesis as "Church's Problem".
https://en.wikipedia.org/wiki/Program_synthesis#Origin
Even program synthesis from a natural language specification is not brand new. For example, here is the first example of research work on the subject that I could find on duckduckgo in ~1 second:
https://www.microsoft.com/en-us/research/wp-content/uploads/...
The fact that most people have never heard of all this before is not new, either. What is new is that OpenAI and Microsoft are marketing a program synthesis tool.
So, something that takes natural talent, years of practice and a lot of time and patience?
At least programming pays well.
---
Transcript of IRC channel #electronics:
Mordae: Hi. So I was designing a transformerless LED bulb and I was wondering about the filter capacitor. What happens when the load (one of the LEDs) fail?
Rampant: Hi Mordae. It really depends on the design, but the general idea is that the filter capacitor will help prevent spikes and other issues caused by the failing LED. It can also help with voltage regulation.
---
Actual answer: Light bulb needs a bulky electrolytic cap. Those are usually not 400V tolerant. You are looking at 60V at best at the required sizes. Losing the load will swiftly charge up the capacitor to 325V. But well before that happens, there are going to be fireworks: https://www.youtube.com/watch?v=6WUxgmMDts4
I believe that there are eventually going to be useful mentors for bunch of fields, but right now it's crazy dangerous to blindly follow its recommendations.
I'd say there is. "Transformerless" implies a specific design, especially when you mention there's a filter capacitor.
> is that AI answer grotesquely wrong?
Yes. It's fine to say the filter cap helps with "spikes and other issues", but not ones caused by a failing LED. It took half a true sentence and made it into a false sentence. In this design the filter cap is something you give protection to much more than something that provides protection.
---------------------
I'm designing a transformerless LED bulb and I was wondering about the filter capacitor. What happens when the load (one of the LEDs) fail?
---------------------
When a load (such as one of the LEDs in your transformerless LED bulb) fails or becomes open-circuited, the current through the filter capacitor will drop to zero, and the capacitor will start to charge up to the peak value of the AC voltage from the mains supply.
This can be dangerous because the voltage across the capacitor can quickly rise to a high level, depending on the capacitance value and the AC voltage level. If a person were to touch the exposed conductors of the failed LED bulb, they could receive a potentially lethal electric shock from the charged capacitor.
To mitigate this risk, it's important to include some form of overvoltage protection in your design, such as a discharge resistor or a metal oxide varistor (MOV). The discharge resistor will discharge the capacitor when the load is disconnected, while the MOV will clamp the voltage across the capacitor to a safe level by diverting the excess energy to ground.
It's also important to provide adequate insulation between the live components and any exposed conductors or casings to prevent accidental contact with live parts. Additionally, it's recommended to follow established safety standards and guidelines when designing any electrical product to ensure its safety for users
Designing a LED power supply as anything else than a current source is asking for trouble.
All the automation in the world does not replace the need for people to debug problems.
Even more interestingly, the more something is automated the more time is spent on solving the issues when they happen. Finding a simple mistake in a very complex automation can be a challenging project.
But... if you have been programming Cobol your entire life you might want to start learning some newer technology.
Why _wouldn't_ the programming AIs be trained on Cobol?