Software Complexity Is Why AI Won't Replace Software Engineers
softwarecomplexity.com
softwarecomplexity.com
But for generating templates or snippets or helping me with documentation for proof of concepts...or being an amazing search friend with great abilities to format data for me, absolutely an amazing friend to have onboard.
Shouldn’t that stuff be in a library or API call with examples a quick google search away?
It came up with the following along with a verbal explanation of how it works:
def quick_sort(arr):
if len(arr) <= 1:
return arr
pivot = arr[len(arr) // 2]
left = [x for x in arr if x < pivot]
middle = [x for x in arr if x == pivot]
right = [x for x in arr if x > pivot]
return quick_sort(left) + middle + quick_sort(right)
# Example usage:
arr = [3, 6, 8, 10, 1, 2, 1]
print("Unsorted array:", arr)
sorted_arr = quick_sort(arr)
print("Sorted array:", sorted_arr)
I then asked it to make the code more efficient. To this, it offered "a more efficient in-place implementation [...] using the Lomuto partition scheme". The code was reasonable but had an off-by-one error. It took me about thirty seconds to find and fix the error.Since ChatGPT said that "the Lomuto partition scheme is generally less efficient than the Hoare partition scheme, but it is easier to implement and understand", I asked it to show the code for the Hoare scheme.
The code was correct and was accompanied by an explanation that "this version of quick sort is more efficient because it performs fewer swaps on average compared to the Lomuto partition scheme."
I asked it about the average number of swaps for both schemes. It gave some answers. However, since I didn't know the correct answers off the top of my head and couldn't be bothered to work them out, I couldn't just take ChatGPT's at face value (but I suspect that one was the correct while the other wasn't).
Finally I asked "What's the worst-case input for these three implementations?"
Here things went completely off the rails. It repeated the same answer for all there snippets. The answer happened to be correct for the last two schemes but for the first it read:
<quote> Basic Quick Sort (with middle element as pivot): The worst-case input for this implementation is a sorted or reverse-sorted array. Since the pivot is chosen as the middle element, the partitioning will be unbalanced, and the algorithm will have a time complexity of O(n^2). </quote>
This is clearly nonsense.
This stuff will be amazing in coming years but I’m shocked people say they use it all day
1. I get stumped and ask it “can I do thing X with technology Y?”
2. It always says “yes, here is how”
3. A majority of the time, it’s lying and you can’t actually do whatever I asked it.
It’s incredibly frustrating - it’s like it doesn’t know how to say “no”.
Granted, I am usually asking it pretty niche stuff about iOS development, but so far it’s proving to be a waste of time. I guess it does feel good to feel like you have a lead on some problem you’ve been beating your head against for hours, but then the harsh reality of actual reality comes quick.
chatGPT is basically a junior dev. Does grunt work, gets it wrong from time to time, and doesnt ever say no even when its appropriate
Over the last few days I’ve been cobbling together a little service which asks GPT to identify issues with code, then if it can generate solutions, the service attempts to apply the solution and run the code.
I’ve been learning to dial in prompts and finding this truly is a major factor of the job that needs to be just right. I went from unreliable response formats, hallucinated solutions, naive solutions, and so on all the way to fairly conservative, pragmatic, consistent responses.
My success rate isn’t as high as I’d like, but I find it absolutely remarkable.
My key takeaway so far has been that a lot of perceived limitations are in fact a lack of exposure to what effective prompting makes possible. Another feature that’s powerful yet overlooked is what recursive prompting can accomplish. This is what makes me think complexity might be overcome in the not-so-distant future.
For example, if I apply a solution and it breaks something else, I can then reform a prompt based on the previous prompt along with a new one informing it of what went wrong. These little recursive exchanges have actually yielded useful results in multiple tests. That’s a huge deal in my opinion; this is totally automated by a smooth brain like me.
As token limits increase and the models become more capable, I suspect this tool will actually improve without me changing much. I suspect it could still work quite a bit better and more efficiently if I could more succinctly provide context about the code in my prompts, too. That would overcome token limitations for the time being and likely improve accuracy quite a bit.
So, if things like this are possible already… I have a feeling we’re going to overcome the complexity issue sooner than it currently seems.
> effective prompting... recursive prompting
It feels like you are massaging a giant library of macros :)
The way you describe interacting with GPT by iterative attempts, constructing working primitives is pretty much programming, right?
AI might get more folks into computer science, or even programming; more hackers, more inquisitive minds, and maybe--just maybe--this "assault on complexity" earns some real victories.
I could write the code myself, but the thing is, if this program works then in theory it could address trivial errors at any time of day, at very high speed. The better it gets at doing this, the more useful it could be.
Most automations start out very inefficient, and many stay that way. I mostly do it to learn but I also see a real opportunity to remove a certain type of task from my plate if this service works well.
Of course, other companies (like GitHub) have likely already prototyped this and whoever runs with it will do a far better job than I'm doing. That's totally fine though, I'm just doing this to learn.
At first I was thinking I would have this inspect errors from something like Sentry, try to solve the issue, then spin up a PR or something if it succeeded (I have a version which also writes tests to verify its changes succeeded, but it's flaky when the intent in the code isn't crystal clear).
I realized though that this would probably be better as an extension in an IDE which can intercept error stacks in a terminal or over a port. That way there's full control over apply patches and no need to arbitrarily run code, which requires a whole virtual machine approach. Anyway, fun to learn, might be useful somehow, but totally cool if ultimately I do just write the code myself.
Today there are musicians who play instruments or compose music in "traditional" ways, and there are other composers who have maybe never touched a piano, but are very adept at assembling high-level building blocks.
Both are musicians, though their skill sets might overlap in only the most basic of ways. Both can write/deliver music. But the one whose expertise is centered around assembling high-level blocks can only make certain kinds of edits, changes, and fixes.
The challenge in education and implementation becomes: as these creation toolkits become more integrated into the way we learn how to make music (or code), do we remove opportunity for folks to develop skills at a deeper level? Do we end up with a big but shallow pool of talent, as opposed to a smaller but deeper pool? (I fully recognize here that this is a question that might apply to any "shortcuts" introduced into any skilled field of work)
End users can now type whatever they can imagine into a computer and it will program itself.
This miracle technology is called a "compiler".
If this upstart AI is going to compete, people had better start testing it on the Linux kernel source rather than the SAT or bar exam.
Rephrased as:
Requirements in your head are easy (or at least seem so); requirements expressed in AI-grokkable form are less easy.
And therefore, according to major-citation-needed pie chart, editing code is only 5% of the job, thus AI will always and forever only be able to save us 5% or less of our time.
Seems pretty wooly.
One reason I think ChatGPT is so impressive is because it's remixing and suggesting code, from already well structured code bases, functions etc, calling on well designed libraries etc. My gut feeling is that MS have allowed all of Github to be "sampled", which is why the code ChatGPT is serving up seems to "familiar".
Unless these models start to actually improve the code, ie suggest better, less complicated paths and architectures, there is a chance software will go backwards to the point where just getting LLMs to pattern match a change into a 500,000 long line codebase of auto-generated, randomized solutions stops working, then of course, we'll need a "bigger" model to grok it all, explain what the code is doing etc.
We could call it "drift" or similar, but I think the fact we write code that is easy to understand for others, including LLMs is actually a feature and not a bug. Anything "cognitive" would have to agree at some stage, unless we have unlimited cognition, for every code change available, which is honestly a fantasy at this stage.
I might be wrong, and I'm happy to be wrong, but this is probably the best time in history to be using auto-completion for coding. There's plenty of well thought out, nice training data available, I don't think that will hold water if we just get copy-pasting, neglect using nice libraries etc.
One interesting example I saw recently was someone showed me how they were using ChatGPT-4 to check their code for vulnerabilities, it actually did quite well, albeit it was very basic / best practices stuff everyone should know that needed fixing. Then i watched someone else in a different context, use it to generate code, that was vulnerable and susceptible to the same problems.
We could automate the vulnerability checking (many people already do), but then it becomes an interesting back and fourth between systems writing and fixing code, at some stage it probably becomes inefficient to do this in it's own right. Brute forcing.
Or is real-world code an impossibly complex algorithm soup ?
Some other people see code as a way to communicate about systems specifications. This is the foundation of collaboration on complex software products. In this arena, lowering entry barriers and the cost of writing code will reduce quality and introduce complex, hard to debug failure modes.
An automated AI system should be able to ask a human for help whenever the confidence score is below a certain threshold or even spit out a backlog of all the tasks it can't confidently handle itself.
So laugh all you like.
ChatGPT gets an import statement.
Prompt Engineers will have to maintain that. It will be written in a nightmare of English, without comments, and heaven help what it means for the tool to "break production."
The reference manual contains a chapter on Forbidden Phrases, but it's marked up on paper, scrawled and crossed out by various predecessors.
Initialization is invoking special phrases, runtime feature flags, to disable whatever behaviors were previously tagged.
Yep: we humans will need to tag behaviors of the AI to describe what is happening, without the benefit of breakpoints.
Where are the breakpoints, anyway? What's our `gdb` of ChatGPT, I wonder.