Cheating is All You Need
about.sourcegraph.com
about.sourcegraph.com
I personally feel the technology is over-hyped. Sure, the ability of LLMs to generate "decent" code from a prompt is pretty impressive, but I don't think they are biger than Stack Overflow or IDEs.
So far my experience is that ChatGPT is great for generating code from languages I not proficient in or when I don't remember how to do something and I need a quick fix. So in a way it feels like a better "Google" but still I would rank it as inferior than Stack Overflow.
I am also hesitant about the statement that it makes us 5 times as productive because we only need to "check the code is good" for two main reasons:
1. It is my belief that if you are proficient enough in the task at hand, it is actually a distraction to be checking "someone else code" over just writing it yourself. When I wrote the code, I know it by heart and I know what it does (or is supposed to do). At least for me, having to be creating prompts and then reviewing the code that generates is slower and takes me out of the flow. It is also more exhausting than just writing the thing myself.
2. I am only able to check the correctness of the code, if am am proficient enough as a programmer (and possibly in the language as well). To become proficient I need to write a lot of code, but the more I use LLMs, the less repetitions I get in. So in a way it feels like LLMs are going to make you a "worse" programmer by doing the work for you.
Does anyone feel that way? Maybe I am wrong and the technology hasn't really clicked for me yet.
In fact, I find 3.5 turbo the best overall model as a tool, because to quality of responses really depends on the quality of prompts, and the quality of prompts is improved by reacting to responses, which come more quickly in 3.5-turbo. So while ChatGPT-4 is still writing the first not-good response, ChatGPT-3.5-Turbo will be on the 2nd or 3rd and it will be much more cogent.
I am glad someone else feels this way. Maybe it’s not going to be as big a paradigm shift as I originally expected.
Knowing humans and the LGTM phenomenon, these kinds of issues will slip by quite readily.
LLMs cannot be judged by their first few incarnations. What can be trained into them currently exceeds imagination. Imagination is our limiting factor.
And I don’t say that from the context of “I jumped on the hype train at the end of last year”. I remember reading the 2017 Google transformer paper and thinking “whoa, this is really happening.” The fact it happened in only 5 years is pretty impressive. Im not sure many papers or innovations got my mind spinning quite like that one.
If your fancy AI coder thingy can't really reason about the end task that the code is solving - and there is little to indicate that it does, or that, any moment now, technology will advance to the point that it will - then the 80% will be crap and there exists no human that can finish the last 20%, not even if they put up 200% of the effort required. We still don't have a working AI solution for driving, a well understood and very limited problem domain, never-mind the infinite domain of all problems that can be explained in natural language and solved with software.
What you end up with is a fancier autocomplete, not an AI coder. Boilerplate and coder output might simply increase to take advantage of the new more productive way of generating source code, just like they did for the last decades whenever there was a "revolutionary" new tech, like high level languages, source control, IDEs and debuggers, component distribution etc. etc.
These are data transformers that can transform raw data without coding at all. At what point does a model itself replace code?
It’s sort of like a CPU, right. You can have hardware that specialized, or general purpose hardware that can do anything once instructed. LLMs have the ability to be general purpose data manipulators without first having to be designed (or coded) to perform a task.
How do you know this is 100% reliable, per upthread discussion?
We've already had this problem with Excel in various sciences, which while deterministic has all sorts of surprising behaviors. Genes had to be renamed in order to stop Excel from mangling them: https://www.progress.org.uk/human-genes-renamed-as-microsoft...
AI promises "easier than Excel, but not deterministic". So more people are going to use it to get less reliable results.
1. Cherrypicking sports statistics for newly set records and the like (NB: this is not lucrative)
2. Financial transaction processing
In most other contexts, especially analytics and reporting, nobody cares and nobody is going to check your math, because the consumers are just trying to put a veneer of numeracy on their instincts.
Yeah, nothing is 100%, and if it were nothing could prove it was.
My problem with it is that they are over hyping its capabilities and trying to market it as "it makes developers 55% faster" because it writes the code for them. I think it would be a better approach to market it as a great tool for automating repetitive tasks and a better way to consume documentation.
* There may not be a lot of differentiation between different LLMs in the long run
* Where there is differentiation, is in data (both the data used to train it and the data provided within its context window for a given query)
* Ergo marrying search to the LLM, while currently in its infancy, will be a big deal and a big differentiator -- because if you can quickly find the right data to pack into the context window, you will get much better results than what we're seeing today.
Crypto, NFT, Blockchain, AR, Metaverse, -- from the top of my head -- now AI. The point of hype is to attract investment. Big Corps must be driven by the fear of missing out on yet another world changing shiny new thing.
IMO these AI technologies have obviously more tangible utility than some of the other hyped things on the list, however a lot remains to be seen about where they go.
... because we all know proving correctness is the easy part of writing software!
I can't wait to read about software engineers finding out some MBA had a huge codebase written by a language model and a few offshored contractors only to realize it's incredibly bugged and being hired to "just go and find the mistakes the error the ai made, should be easy all the code is written".
To best illustrate what I mean, watch this chess match[0] it's quite riveting.
Since it read millions of matches, it can predict a legal move most of the time, and even some good moves some of the time, but it cannot "understand" the rules of chess, and makes some hilariously illegal moves, especially if the match lasts longer.
[0] https://www.reddit.com/r/AnarchyChess/comments/10ydnbb/i_pla...
Me: crane
GPT: _ _ _ _ e
Me: moist
GPT: _ _ _ r _
Me: glyph
GPT: you guessed it, the word was glyph
Now, maybe GPT 4 or other future developments will give better results, but to me this highlights exactly what you're saying. LLMs do not have an internal structure in their 'minds' that they're pondering about. It's a very impressive engine for guessing the next character to produce into a stream.
There's definitely usages for this, but not what a lot of people are saying.
My pet conspiracy theory is that is is wired to please the user, to get better coverage from the media and social media.
So it could be that ChatGPT picked up on a pattern in the training data where after a couple of guesses, a lot of the time people pick the right word.
So statistically it might go like. Guess a word. Probably not the right one. Guess a couple more and suddenly it’s statistically likely to be the correct word, and because of that the LLM ends up outputting the congrats and so on
Even if it reasoned ideally, it has to check if your answer is consistent, and if it is, decide whether to let you win or to find another solution. If your answer is inconsistent, it effectively has to play wordle by thinking of another answer.
It doesn’t seem all that surprising that it didn’t figure it out.
> We will play the word game wordle. Start by picking a 5 letter word. Display it in b64 so I can check later that you kept the same word all along. I will submit my guesses and you will show the letters that are correct, we will keep going until I have found the word
> Great, I've picked a 5-letter word and encoded it in base64. Here's the encoded word: Q2l0eWw=. You can now start submitting your guesses!
But after a few tries it was clear it was not working so I gave up:
> I give up, what's the word?
> Before I reveal the word, let me first decode the base64 string I provided earlier to ensure that I'm giving you the correct word. The base64 string was: Q2l0eWw=. After decoding the base64 string, the word I chose is "citly". However, I apologize for the confusion caused during the game, as "citly" is not a valid English word. I made an error while choosing the word. Please accept my apologies for any inconvenience caused during the game.
However asking it for a python program to play the game worked perfectly the first time, try it here: https://replit.com/@CorentinSmith1/gpt-wordle#main.py
https://help.openai.com/en/articles/4936856-what-are-tokens-...
> Tokens can be thought of as pieces of words. Before the API processes the prompts, the input is broken down into tokens. These tokens are not cut up exactly where the words start or end - tokens can include trailing spaces and even sub-words. Here are some helpful rules of thumb for understanding tokens in terms of lengths:
> ...
> Wayne Gretzky’s quote "You miss 100% of the shots you don't take" contains 11 tokens.
https://platform.openai.com/tokenizer
It isn't going character by character, but rather token by token - both for input and for output.
This also helps explain why it has trouble with breaking a word apart (as in the case of wordle) because it doesn't "think" of glyph as 5 letters but rather two tokens that happen to be 'gly' and 'ph' with the ids of [10853, 746].
(2) Changing the prompting reduced the illegal moves to almost 0
(3) There have been experiments that show GPT has a internal "state" of the world and can do simple reasoning puzzles. This model of the world evolves with each generation.
I understand the skepticism, but don't let that blind you to the reality of the technology. I'm a skeptic at heart, and I could immediately tell GPT was a game-changer. It can already replace half of the ML models that are used at my job and do it better (if it was economical enough).
I've been experiencing this myself recently. I've been using co-pilot in some side projects. I've noticed myself getting more 'lazy' as I use it more.
Recently I used it when doing some old (2015) advent of code puzzles I hadn't done before. I would read the puzzle prompt and have a pretty good idea of what I wanted to do. I wrote out some comments for functions and co-pilot was able to write what I needed with minimal changes.
Even though I read through co-pilot's code and understood what it was doing I don't feel like I really retained anything from the time spent. If anything, I feel like co-pilot stunts my learning.
It truly is magical when the code just runs. Later I asked it to make several non-trivial changes to the code based on more requirements I thought of, and it aced those on the first go as well. Again, I checked the code for a negligible amount of time - compared to how much it would have taken me to write the code on my own.
I do think humans will slowly get worse at lower-layers of the computer stack. But I don't think there's anything inherently bad with it. Compilers are also doing the work for you, and they are making you bad at writing assembly code - but would you rather live in a world where everyone has to hand-write tedious assembly-code?
Maybe, in the future, writing Python would be like what writing assembly is today. We might go down the layer-cake once in a while to work with Python code. That does not mean we give up on the gains we get from whatever layers are going to be put on top of Python.
What is the equivalent of this for LLMs? Is there anyway generative models can give a guarantee that this prompt will 100% translate to this assembly? As far as I understand, no. And the way autoregressive models are built I don't think this is possible.
I agree that they are useful for one-offs like you said, and their ability to tailor the solution for your problem (as opposed to reading multiple answers on stackoverflow and then piecing it yourself) is quite deadly, but for anything that is even slightly consequential, you are going to have to read everything it generates. I just can't figure out how it integrates into my workflow.
Anyway, people still write assembly kernels, so it is just that they only do it for cases that really matter. And there are a lot more coders than there were back when every program was assembly. So, it seems like great news.
It’s like telling a novelist that they can produce novels much faster now because they only have to think of the rough outline and then do some minor editing on the result. For most, this is antithetical to why they became a novelist in the first place.
It's a funny distinction! Knowing something can be automated can take some of the fun out of it, but there are plenty of people who still do stuff for fun when they could buy the end result more cheaply.
For employers, though, it's all a means to an end. Go write for the love of it on your own time.
> I was recently laid off, and I know a few other people laid off. I have years of doing projects and contributing to OSS and being a technically curious learner. I found a new job much faster than my peers who admittedly joined tech for the money and don’t care to learn or grow beyond their next pay raise.
There is a fairly consistent chorus of people getting into software development - not because they enjoy the intellectual challenge that it presents but rather because of the potential for the pay.
As someone who does enjoy software development (I chose this path well before the dot com boom), I believe that we over-estimate the number of people who enjoy it compared to just grinding through writing some code and if something else paid as well, they'd jump in a heartbeat.
The firm can't really afford to care too much about why its workers entered their professions. The firm has to care about the cost of its inputs and margin lest it be devoured by a competitor or private equity.
The fact that companies may see that differently is beside the point, and I don’t particularly expect them to care for my preferences. I will however certainly continue to choose employers that happen to accommodate my preferences.
Since its inception, computer science has had two "camps": those who believe CS is engineering, and those who believe CS is mathematics. The reason why we are seeing all of this fuss around LLMs is that they are a new front of this feud. This "extends" the usual debate on emerging technologies between Thymoetes and Laocoon.
Something that works 99 times out of 100 is 99% correct from the first perspective and 100% wrong from the second.
LLMs are therefore a step forward if you take the first view, a step back if you take the second.
If you accept this interpretation, an interesting consequence of it is that your outlook on LLMs is entirely dependent on what amounts to your aesthetic judgement.
And it's very hard not to have rather strong aesthetic judgements on what we do 40 hours a week.
However, there is a pragmatic school of hacking, which says that results are all the matters. If you're in a startup, you should be pragmatic, and worse is better.
Nobody truly believes that CS is engineering.
What reasoning are you using to come to conclusion that software is not engineering?
>The creative application of scientific principles to design or develop structures, machines, apparatus, or manufacturing processes, or works utilizing them singly or in combination; or to construct or operate the same with full cognizance of their design; or to forecast their behavior under specific operating conditions; all as respects an intended function, economics of operation and safety to life and property
It is purely software engineering.
Whenever you have to qualify a noun with something else, the result is something narrower than the original noun, and often completely different:
- Software Engineering is not Engineering
- Street Justice is not Justice
- Covert Intelligence is not Intelligence
SE is way, way younger than other engs.
>to produce a physical thing
Why physical thing would be a requirement here?
>This is why an engineer sometimes has to say "no, that won't work".
But similar scenerios can occur in SE world too, so what does it even mean? like some unsound design of distributed system
The surprising thing to many of us is that the "software engineer" title has been specifically associated with deprofessionalization of the job and the view that is moving away from true engineering. See the Dijkstra link I posted elsewhere in this thread.
My parents were very respected traffic engineers. One was concerned with traffic inside cities, another with road design. One couldn't give less care about "nature of materials and physical force" while figuring out where to put traffic lights and close roads, while another had a set of standards (say a lorry's turning radius, sign sizes etc) which they've applied without thinking about costs or even producing a pysical thing. That was done by some other contractor in a sense anyway.
A friend of mine is an electrical engineer. He maintains electric motors and I guess he would fit your definition better. At the same time, how much an electrical engineer has in common with traffic engineer? Or say an air traffic engineer? Or a construction engineer? I would say what they mostly have in common is the "engineer" part as they are doing completely different jobs in completely different industries. I might even go further and say when someone calls "Alice is an engineer", it is just a shorthand for "Alice is a mechanical [or some other] engineer".
So those were the differences, but lets see the similarities. When a traffic engineer gets a project to see if they can increase the traffic flow from point A to B it usually follows the common pattern of finding out the current situation, figuring out a solution, testing/simulating it out and then making the project/solution on what should be done. Then when the solution gets implemented, some details can be re-adjusted etc. I did simplify it a bit, and I'm unsure in the exact english phrases, but I think you get my point. Software engineers that I know follow the exact same process. Oh, a database is overloaded, lets see what we can do about it. Can we give it more hardware? No, as thats too expensive? What about if we scale up/down the services/workers/instances/nodes/pods calling that database? Can we simulate that? What do the metrics say? Etc etc. You see where am going with this? In essence, there is no difference.
My point is, if you cannot call Software Engineering only by the word "Engineering", then you cannot call anything else (just) Engineering. You can dance around this whatever you like, and define Engineering either narrowly or widely enough not to include Software Engineering (just like you did), but at the end of the day this isn't a natural science like say Maths or Physics. You cannot say 2+2=5, but you can say Engineering is X. But then, lets be honest here, and I suspect this is the root cause of this: it's about exclusivity and status. Engineers, doctors, lawyers, electricians, plumbers and a myriad of different professions like having guilds (or similar organizations, speaking in general here), and like having artificial scaricity. This brings status, it brings money, it brings power and influence. It cannot work if everyone is allowed in. It's what's simply called gatekeeping. It's what hiding behind a dismissive, elitist sentence I've heard too many times: "Oh, but he's no engineer". So if we're honest here we can call a spade a spade.
Let's resort to definitions, with apologies.
Engineering
> Engineering is the use of scientific principles to design and build machines, structures, and other items, including bridges, tunnels, roads, vehicles, and buildings.[1]
That certainly doesn't make Software Engineers what are traditionally thought of as "engineers." So... on to Software Engineering:
Software Engineering
> Software engineering [is] the application of a systematic, disciplined, quantifiable approach to the development, operation and maintenance of software and the study of these approaches; that is, the application of engineering and computer science to software. [2]
That description does not fit what most of us do when we write or create software. Perhaps it should. We'd be better off in many cases. But programmers are often closer to gardeners than engineers in that we approach solving problems through empirical processes, testing and verifying as we go. TDD is an admission of the need to do just that. So is Scrum (empirical process control). There absolutely are cases where algorithm optimization using the latest research must be applied just so. There absolutely are cases where applying gradient descent and convolutional neural nets in specific fashions applies. Most of us are gardeners growing our systems, or carpenters building out a feature, though. Programmers. Software engineer makes it sound loftier, just like sanitation engineer makes trash collection sound loftier.
I will absolutely concede that I still use the "engineer" terminology on a day to day basis, my opinion on the internet notwithstanding. But it is a label that is weakly applied for fashion and connotation, rather than for its distinct meaning, and I suppose that is what I tilt at here on HN. Apologies for the pedantry.
[1] https://en.wikipedia.org/wiki/Engineering
[2] https://en.wikipedia.org/wiki/List_of_engineering_branches
I'm sure that say a conteporary civil engineer would scoff at calling Nikola Tesla an [electrical] engineer, just like you're virtually scoffing at calling a Software Engineer only by Engineer.
Admittedly that the software field is young, but does it make it "worth" less? To entertain your argument, do you consider a person programming PLCs an electic engineer? What's the fundamental difference between a software programmer and them? I don't see any. So you cannot call one an [Electrical] Engineer and another one "just" Software Programmer.
Gardener is not an engineer. On a ship there are usually at least two departments: Deck, and Engineering. The latter are people in charge to make sure that ship's systems are up and running at all times. There is no theory behind it, or pretty much any maths. It is mostly boring maintenance. And I would be surprised that anyone would call those people non-engineers. So why bring this up? Engineering definitely has a vague connotation with certain terms like electrical power, mechanics, devices, construction, materials etc. As the gardener or a trash collector doesn't deal with any of those, noone is considering a gardener an Engineer.
One can say, oh it's the title inflation, but that actually applies to any profession then. Or to put it differently, what makes a person an engineer? The work that person does, or the title it was given to them by some institution? Is it both? Is someone an engineer but is a career politician? I sure know one, and the guy will always be legally an engineer. This is a classic "what makes an art, Art?" question which is another can of worms I'm not going to open.
So I'm not sure what you're arguing then when all of this is on a very shaky ground. My position is that it is just a matter of a title and social convention. If you're a part of a engineering department, then I'm sure people outside of that department will call you an engineer. Inside of that department you might be a programmer of say a backend service, I could be a platforms guy, Alice could be QA engineer, and Bob could be a hardware engineer. It just doesn't really matter in the end. All of those are just social constructs which change with time. Or if we take it to a institution level: in some countries you have legally recognized Software Engineers (as I found out in other threads), in a country like mine you don't have one, whether we like it or not. So as I said in my previous post, this isn't maths or physics, and it is just a social construct which varies all over the world. So arguing about it is like arguing about the definition of art. I like spending time on engineering more so that's it from me ;)
However, there are some ideas clustered around that seem to have something to do with it. Engineering on a ship has to do with electrical and mechanical systems, and in a highly constrained set of outcomes, essentially reducible to a single bit at any point in time. (Either the systems are up and running right now, or they aren't.) Within this system of evaluation of outcomes, it's not surprising if the engineers in this context aren't using maths every day but are engaged mainly in more practical actions. However, you can be sure that they know a lot of safety tolerances and operational characteristics of various ship systems that would be characterized as mathematical, even if they are mostly operating well within tolerances that make these operations routine.
One reason why people argue about "software engineering" is that there are authors who define it specifically to include scientistic or bureaucratic rituals that have no mathematical underpinning, while excluding the difficult mathy bits of engineering. The further you get from actual engineering, the more common this definition gets.
EWD1165 "There is still a war going on.":
For those people, unless something fit within an existing (but large!) range of use cases, they were out of luck without having an engineer or mathematician figure it out for them. Suddenly, there is a glimmer on the horizon that all of that possibility the computer science people see every day could be unlocked for the users, and even if it only works 5% of the time, that is enough to get them excited in ways that are hard to describe to the computer science people.
Imagine going to school, a boot camp, or being self taught in everything about screwdrivers and screws. You can discuss at length the advantages and disadvantages of different shapes (Robertson bits > all), materials, screw threads, etc. You can custom design a screwdriver and screw for a specific application, taking into account all of the relevant constraints.
Now imagine the guy who needs to tighten a loose cabinet door.
Screwdrivers don’t have nearly the complexity or ability to generate work leverage that computers do, moving even a few percent of those capabilities from the first group to the second is huge. It is, at minimum, Excel huge.
For the rest of the world, while some might be excited by what you describe (and it that works for them, that's great!), I believe in general the interpretation is far simpler: me like shiny.
Could you expand on this?
> Something that works 99 times out of 100 is 99% correct from the first perspective and 100% wrong from the second
Interesting. From a _manufacturing_ perspective, you can't achieve 100%, you can only get asymptotically closer to it with statistical process control. And of course there are limits to the perfectibility of humans.
This suggests that the big deployment of AI will be in areas where there is no clear boundary between right and wrong answer.
> Could you expand on this?
Not important, it's just a rhetorical flourish. In the second book of the Aeneid, Thymoetes is the guy who says (paraphrasing) "let's bring the horse inside" and Laocoon is the guy who says (literally) "beware of Greeks bearing gifts".
> This suggests that the big deployment of AI will be in areas where there is no clear boundary between right and wrong answer.
"AI" is an umbrella term at this point. If by AI we mean LLMs or similar technology, then my hunch is to agree with the statement. I don't think this is particularly controversial though, IIRC Yann LeCun said something similar.
Roughly, "I fear Greeks even when they come bearing gifts".
Urban development can be seen as a balance between careful planning (akin to the mathematics camp) and organic growth (resembling the engineering camp). A city designed with a focus on aesthetics and theoretical frameworks might be visually appealing, but it could lack adaptability. On the other hand, a city that grows organically may not be as cohesive, but it's more practical and responsive to its inhabitants' needs.
This parallel can help us better understand the emergent properties of LLMs, which arise from their complex interactions. By appreciating both the engineering and mathematics perspectives, we can gain a more comprehensive understanding of these properties.
Moreover, the balance between early adoption and risks, as seen in urban development, can also apply to LLMs. Early adopters of LLMs can tap into their potential, but they must also be aware of potential risks, such as biases and ethical concerns.
Oh yeah ChatGPT wrote this answer.
Math - the study of well defined concepts and their relationships. Solving problems with proofs.
Engineering - solving well characterized problems based on math and physics (which can include materials with known properties, chemistry, approximations, models, …), and well defined areas of composability (circuits, chemical processes, structural design, …)
Craft - solving incompletely characterized problems with math, physics, engineering and enormous amounts of experience, intuition, heuristics, wisdom, patterns, guesses, poorly understood third party modules, partial solutions pulled from random web sites …
Art - Solving subjective problems by any means necessary.
I recall a self taught dev (or maybe from a bootcamp) coming up with a cascade of nested if-else, nested 8 deep. Someone with a background in CS asked him what he was trying to do and basically concluded that what he was trying to do could be expressed as a state machine. To which the initial dev replied that it was "way too fancy" and that he didn't need the code to be fancy, just work.
The ugly part: The nested-8-deep solution was faster to market and costs less. And, it will be thrown away in 6 months during the "big rewrite after we scale". So the perfect-is-the-enemy-of-good state machine solution written by an expensive engineer has less value. Oompf.
What about maintainability? Extensibility and ease of debugging?
I've seen chunks of projects re-written, just because it was simply impossible to extend them without significant efforts!
Are they? Source readability will still matter if you end up using a debugger!
Eh, I think you’re overselling how precise and well defined engineering is in other fields. Engineering in other fields is just as much dealing with poorly characterised problems as it is when writing code (it takes quite a lot of characterisation to go from “we want a bridge here”, to an actual damn bridge, and that’s all an engineers work).
Really the core of engineering is just a very broad set of practices and principles that allows people to solve poorly characterised problems using maths, physics, enormous amounts of experience, intuition, heuristics, wisdom, patterns, educated guess etc in a reasonably consistent and repeatable manner. Doesn’t matter if you’re building a web browser, a motherboard, or a bridge. You don’t get a good result without a healthy dollop of wisdom, experience, educated guesses, and a handful of fuckups (which hopefully you notice before you let people use the thing).
Engineering in other disciplines is no less messy, haphazard, and experimental than it is in software. It just isn’t as publicly documented as it is software, probably because it’s hard to build an open source bridge.
Logic in digital circuits.
Algebra and calculus for analog circuits, most physical objects, properties and processes.
Differential equations for dynamical systems and dynamical behaviors.
Sure there is a lot of creativity in engineering, but there is usually a whole area of math known to be suitable for expressing solutions clearly, given the area of engineering.
Contrast with the utter lack of standard notation across software tools and implementations, for describing all the trade offs, gotchas, glue, historical drift & complexity, theories of memory, caching, user affordances, potential overflows, races, etc. that is implied by a program’s code.
Sometimes a language provides islands of engineered code, like message passing in Erlang, or memory management in Rust, or a precise mathematical library like BLAS.
But most aspects in most software programs are created ad hoc, or inherited from someone else’s rats nest of an implementation, and never formalized completely, if at all!
Any clarity in representation quickly leaves planet applied math.
Take it from someone who’s studied the maths you’ve described and applied it in a professional capacity. Just knowing the maths isn’t anywhere near enough. It’s like knowing how quick-sort works, interesting and useful, not even close to enough to actually build anything.
> Contrast with the utter lack of standard notation across software tools and implementations, for describing all the trade offs, gotchas, glue, historical drift & complexity, theories of memory, caching, user affordances, potential overflows, races, etc. that is implied by a program’s code.
The entire field of computer science is dedicated to developing and using this type of notation to describe and understand the basic principles that underpin every programming language, database and algorithm you’ve ever touched. The notation exists, you just don’t use it. In the same way an engineer in another field doesn’t bother doing structural analysis, or circuit analysis from first principles, they just grab a pre-finished tool and apply it to their problem. Normally by just passing the problem to a computer that does all the heavy number crunching, and checking the outputs make sense.
> But most aspects in most software programs are created ad hoc, or inherited from someone else’s rats nest of an implementation, and never formalized completely, if at all! > > Any clarity in representation quickly leaves planet applied math.
Again, I think you’re vastly overestimating the precision of other fields of engineering. There might be more rigour in design in some places, but only because “just making a thing and see if it works” is expensive, but other engineers absolutely spend huge amounts of time just making things to see if stuff works.
There’s a reason why “safety factors” are a thing, and reason why they’re usually 10x or greater. That safety factor is also the “we’re not sure how well this works, so we made it ten times stronger than we think we need, just in case, factor”. Engineering in other fields is people doing maths all day long, it’s mostly reading data sheets, assembling Lego brick components, and hoping to hell the manufacturer didn’t lie too much on their data sheet. Plus some design and simulation on a computer for good measure.
You wanna see “ad-hoc or inherited from someone else’s rats nest of an implementation” in a different field of engineering. Then go look at any electronics catalog, or lookup YouTube videos of people testing parts against the specsheet and discovering how wildly different they can sometimes be.
Dodgy, badly implemented, never formalised engineering exists everywhere. That’s why bridges collapse when they shouldn’t (Genoa Bridge), why planes crash when they shouldn’t (Boeing 737 Max), why cars emit more emissions than they should (VW), why buildings get emergency modifications after being built to prevent them from being blown over (601 Lexington Avenue). Software engineering does not have a monopoly on botches, last minute hacks, and dodgy workarounds. Engineers in other fields were merrily employing all of them to great effect for centuries before software turned up.
The division of labor into software developers and non software developers isn't any different than farmers vs non farmers or any other profession.
As a counterpoint: Look at the history of the libxcb: https://en.wikipedia.org/wiki/XCB
[Bart] Massey and others have worked to prove key portions of XCB formally correct using Z notation.
That sounds like math to me. Or is it "craft"?Mathy projects, formally driven: matrix multiply libraries, symbolic computation, constraint resolution, ...
Engineered projects, formally (or close to it) verifiable: 3D rendering pipeline, distributed database management, garbage collection process, ...
And craft. Which, based on the internet I have experienced, many apps unjustly inflicted upon me, and some memorable restarts between game saves, is most code.
Craft code necessarily involves amateur code, code which isn't economically worth engineering (when you can just throw unit tests at it. Or just wait for user reports!), code referencing weakly characterized libraries or interfaces, and code involving features that have become complex enough that the best reference model for its expected and unexpected behaviors is now itself.
Bzilion's Law of Coding Formalism Levels: "Any ambitious enough software project will descend into an exercise of pure desperate craft. Just before it becomes gambling."
[otherwise a (code) monkey could do it] is missing the point of programming.
I know managers who code (occasionally) who think similar. I thought so personally before I had actual prolonged experience with professional programming.
It is hard to express it concisely: why it is fundamentally wrong (category error: like thinking that perl regexes can be reduced to DFA--it is impossible in the general case even if DFAs can [sometimes even should] be used in many cases instead).
It is the same reason why waterfall programming fails most of the time. It is the same reason why generating code from UML diagrams produced by analysts is also a failure in the general case. It is the same reason why log normal distribution can be a good model for software estimation https://news.ycombinator.com/item?id=26393332
And no, you can't replace all programmers with a LLM prompt for the same reason (at least until [if ever] it reaches AGI and then humanity would have much bigger problems).
"agile" became a noun but if you look at the origins, you might get why "craft" may be applied to programming. Try "The Pragmatic Programmer" book.
Coding and software construction is engineering or craft, and is not Computer Science.
LLMs are neither. They are power tools for concept realization.
It's the difference between stone chisels and a suite of shop tools. We had pen and paper, or small steps up from those, and now we have LLMs.
So much of programming is rote boilerplate garbage simply linking things together and so little of it is actual creative thought. If LLMs could actually generate the rote boilerplate, programming would be soooo much better.
Alas, my optimism isn't that high.
No, I'm trying mightily to do what Yegge is talking about in the context of the programming work I do everyday. First v3 then v4. I've given up until maybe v7 or something.
The problem is it doesn't have experience with my code-base. Sure, tell it to open a file and return a stream, it'll do that (after I fix the using statements), but for what I'm doing every day it doesn't even begin to know what to do.
And because I'm careful about KISS and SOLID I don't really need a lot of simple code generation. I don't see 5x productivity. I actually don't see much advantage over the built in tools in VS.
Maybe I'm doing it wrong, or maybe this make sense for people who write a lot of boilerplate, but that's not a lot of what I do.
You will definitely learn from LLM suggestions. The mantra „Read other people’s code“ is accurate IMO - as long as the code is at least ok-ish. I‘ve learned a ton from code that ChatGPT generated for me already.
When you know both it's just really good autocomplete, which is great but not a huge game changer. If you know neither then you're not in a position to assess the output. But when you're still learning either the tool or the space I've found GPT to be a good tool for leveraging one expertise to create the other.
We will see in a few years.
Business logic and baking in domain expertise into your data model is most of the work. Making the code work efficiently doesn't matter if your code doesn't even do what it's supposed to.
Normally this is an argument in favor of human-in-the-loop LLM-based development -- "the human just needs to curate and verify!" However it seems all too easy to me (especially having witnessed it more than a few times) that subtle discrepancies emerge between the stakeholders' desires for the function of the code and the developers' understanding of those requests. Hopefully we reach a best-case scenario where that's all developers need to focus on, but more likely we'll see some pretty egregious things slip through the cracks (the wave will likely start with security/privacy issues before the phenomenon is recognized) as this technology matures into the common workplaces.
Yet I find myself asking ChatGPT every now and then "hey how do I do <foo>", where <foo> is something I last needed to do a year or more ago. I can recognize the correct answer but don't need to search docs/net for it.
The reason this is faster (for me) than Googling or using Dash/Zeal is that the answer is already in the context of what I'm trying to do, whereas if I'm only looking at the docs, I will probably need to go through several pages to get a complete picture.
I'm sure there were programmers who said the same thing in regards to high-level programming languages.
> 2. I am only able to check the correctness of the code, if am am proficient enough as a programmer (and possibly in the language as well). To become proficient I need to write a lot of code, but the more I use LLMs, the less repetitions I get in. So in a way it feels like LLMs are going to make you a "worse" programmer by doing the work for you.
Maybe that becomes irrelevant the more that the skill of the programmer shifts from handwriting "correct" code to supervising code generators while proofreading their work, and of course providing effective acceptance criteria. There's also a massive bias towards failed predictions of the past that serves to discredit predictions that may see a greater degree of manifestation. For every time someone says "but people predicted this before and it didn't pan out", I can point to technology that did fundamentally change how an industry works and even make jobs obsolete.
Seems to me a lot of programmers on HN are refusing to believe that their ability to be proficient with code may be either outdated or supplanted by the efficiency of a system that writes code that is not necessarily "elegant" in human terms.
> So in a way it feels like LLMs are going to make you a "worse" programmer by doing the work for you.
Most programmers aren't great at what they do to start with, whereas LLMs can only get better from here on.
A sample example. I asked it to generate Terraform code for registering an organizational unit in AWS Control Tower. This is impossible because the API of Control Tower is very limited. But ChatGPT was very happy to generate a solution pretending to use the official AWS module with a made up resource. Of course, the "solution" was not working at all. But if I ask it to do a trivial task, such as attaching an OU to an organization using AWS Organizations, it can do it perfectly well. And this, for me, is the difference between a human programmer and a machine that is good at certain tasks.
That's pretty much every ChatGPT-programming sample I've read so far.
This one thinks character `i` in elisp regexps is matched with `\i`.
The value here is that the llm can act as a knowledge graph were common sense is preloaded on almost every topic, so that the user can add node and edges on the graph in natural language and perform extraction in natural language
And you don't need fine tuning as long as you can fit the topic in their token space, and with gpt4 reaching 32k tokens you can load a huge amount of text and perform queries on it.
That's what makes the tax return example so interesting. The model has already learned a lot of common and uncommon sense so it will not need the instruction on how to process the text or parse the query.
Forget coding, but everything else is great for.
For code understanding tasks, the issues with standalone LLMs is that they have a certain amount of "memory" which is limited to their training data (SO and OSS)—and even that can be unreliable.
A big "a-ha" moment for us was the realization that LLMs get much more helpful and reliable when coupled with a competent context fetching mechanism that can surface relevant code snippets from your own codebase. This makes Q&A much more factually accurate (and code generations that learns from the patterns in your codebase). We don't think LLMs will ever replace human coders, but we think they can be super helpful in eliminating a lot of the tedious, boring, duplicative writing and reading code that devs do every day. The entirety of Sourcegraph (not just the LLM part) is focused on eliminating these pain points.
One of the more toilsome bits of coding I do personally is rebasing. I have a patch to add application-time temporal tables to the Postgres project, and I've been rebasing it for several years now. It's a pretty big patch (actually a series of four patches), so there are almost always non-trival conflicts to deal with. If ChatGPT could do that for me it would be awesome.
But it's probably the hardest thing for an LLM to do. It's not a routine program that has been written thousands of times across Github projects and StackOverflow posts. Every rebase is completely new.
OTOH it would be awesome if git had just a bit more intelligence around merge conflicts. . . .
Second point about the comments, actually I'm seeing the AI write much better comments (i.e. some) than most devs (none).
Some comments are far worse than no comments at all. I would agree that even semi-decent comments are far better than nothing. However, "no-new-information" comments are just noise, and misleading comments have a huge negative effect. I would not be surprised if an AI produced a large number of the former, and perhaps some of the latter.
It can summarize the thing you're looking at, tell you how to improve it regarding readability and performance.
On a press of a button you can zoom out of the code into an UML like overview and it will tell you what's going on and how it is connected. If you don't get it, it knows how to make you understand.
Then you can tell it in a few words what you want to achieve and it will assist you in finding a solid solution which matches the coding style of the rest of your project. And while you're coding and lose sight, it will help you achieve the goal.
The current state is sub-par in my opinion. I can write good code and don't need an AI to write it for me. But what I want is something which assists me with understanding code, improving code or extend code without taking the steering wheel away from me.
I put my mangled and minified code in.
What it emits is not only perfectly readable and 95% accurate (the 5% was due to missing context, max input limit)—it was significantly better than my code.
Of course the structure was the same, but ChatGPT chose much more sensible variable names in almost every case. I found it much easier to understand its version of my code than my own.
I guess I accidentally discovered a refactoring technique?
Good news: there are kinds of extremely useful comments that do not repeat the code (your comments should not repeat the code). Comments are to express context/intent behind the code: the "why", the high level "what", and almost never the exact "how" (read code for that).
It looks like you only ever encountered the "how" comments. No amount of code refactoring would get you the "why" (context) comments.
I think though that LLM-based tools will eventually formalize to achieve a greater precision at what's required. I suspect that they could be a base for a new crop of different, much-higher-level programming languages.
Programming languages went a long way; somebody from 1960 would have hard time putting things like Haskell or even SQL into the same conceptual bin as the original Fortran. We routinely see them as programming languages though. I don't see why this trend can't continue upwards, relegating even more legwork onto the machine while talking to it in reasonably precise, standardized, domain-specific terms.
1. Idea Generation: The first step in creating a software product is to come up with an idea or goal that the software will achieve.
2. Research: Once you have an idea, it is important to conduct research to determine the feasibility of the idea and identify any potential challenges.
3. Planning: After research, planning is necessary to determine the scope of the project, the timeline, and the resources required.
4. Design: The design phase involves creating a detailed plan for the software, including the user interface, functionality, and architecture.
5. Development: In the development phase, the software is created by writing code, testing, and debugging.
6. Testing: After development, the software must undergo rigorous testing to identify and fix any issues.
7. Deployment: Once the software is tested and ready, it is deployed to the target audience.
8. Maintenance: Finally, the software must be maintained to ensure that it continues to function properly and meets the needs of the users.
Each of those steps has a back and forth with a LLM that can enhance and speed up things. You're talking about 4 as being problematic, but right now there's a lot of "human in the loop" type issues that people are encountering.
Imagine having the following loop:
1. LLM has generated a list of features to implement. AI: "Does this user story look good?" Human: "Y"
2. For each feature, generate an short English explanation of the feature and steps to implement it. Your job as a human is just to confirm that the features match what you want. "Should the shopping cart
3. For each step, LLM generates tests and code to implement the feature. AI: "Shall I implement the enter address feature by doing ..." Human "Y"
4. Automatically compile the code and run the tests until all tests implemented and feature is complete according to spec.
5. Automatically document the code / feature. Generate release notes / automated demo of feature. Confirm feature looks right. AI: "Here's what I implemented... Here's how this works... Does this look good?"
6. Lint / simplify / examine code coverage / examine security issues in the the code. Automatically fix the issues.
I think you also miss that the LLM can be prompted to ask you for more details. e.g. PROMPT: "I'm building a shopping cart. Ask me some questions about the implementation."
1. What programming language are you using for the implementation of the shopping cart?
2. Are you using a specific framework for the shopping cart or are you building it from scratch?
3. How are you storing the products and their information in the shopping cart?
4. How are you handling the calculation of taxes, shipping costs, and discounts in the shopping cart?
5. What payment gateway(s) are you integrating with the shopping cart?
Which can then be fed back to the LLM to make choices on the features or just plain enter the answer. PROMPT: "For each question give me 3 options and note the most popular choice.", and then your answers are fed back in too. At each point you're just a Y/N/Option 1,2,3 monkey.
More succinctly, in each step of the software game, it's possible to codify practices that result in good working software. Effectively LLMs allow us to build out 5GL approaches[1] + processes. And in fact, I'd bet that there's a meta task that would end up with creating the product that does this using the same methodology manually. e.g. PROMPT: "Given what we've discussed so far, what is the next prompt that would drive the solution to the product that utilizes LLMs to automatically create software products towards completion" ;)
[1]: https://en.wikipedia.org/wiki/Fifth-generation_programming_l...
Think about for example how Windows 1.0 looked. For an expirienced DOS user it was offering very little. Expirienced DOS users were saying GUIs are over hyped. Today there are probably a few dozen people worldwide who use a computer without a GUI (or a voice interface).
ChatGPT&Co will obviously make 90% of the software developers out there obsolete in just a few years. An industrial revolution is happening in the software industry.
Think of it like an accessibility aid. It doesn't help people who don't need them, but for those who do it's life changing.
Edit to add: This was an aside in the post but actually a big deal... With this setup you can basically use an off-the-shelf LLM (like GPT)! No fine-tuning (and therefore no data labeling shenanigans), no searching for an open-source equivalent (and therefore no model-hosting shenanigans), no messing around with any of that. In case you're wondering how, say, Shopify and Hubspot can launch their chatbots into production in practically a week.
https://github.com/pinecone-io/examples/blob/master/generati...
https://www.pinecone.io/learn/openai-gen-qa/
https://www.youtube.com/watch?v=tBJ-CTKG2dM&t=787s&ab_channe...
There are more out there but hopefully this gets you started.
0) You can't add new data to current LLMs. Meaning you can't train them on additional data, or fine-tune, leave that more for understanding structure of the language or task.
1) To add external corpus of data into LLMs, you need to fit it into the prompt.
2) Some documents/corpus are too huge to fit into prompts. Token limits.
3) You can obtain relevant chunks of context by creating an embedding of the query and finding the top k most similar chunk embeddings.
4) Stuff as many top k chunks as you can into the prompt and run the query
Now, here's where it gets crazier.
1) Imagine you have an LLM with a token limit of 8k tokens.
2) Split the original document or corpus into 4k token chunks.
3) Imagine that the leaf nodes of a "chunk tree" are set to these 4k chunks.
4) You run your query by summarizing these nodes, pair-wise (two at a time), to generate the parent nodes of the leaf nodes. You now have a layer above the leaf nodes.
5) Repeat until you reach a single root node. That node is the result of tree-summarizing your document using LLMs.
This way has many more calls to the LLM and has certain tradeoffs or advantages, and is essentially what Llama Index's essence is about. The first way allows you to just run embeddings once and make fewer calls to the LLM.
[1] https://langchain.readthedocs.io/en/latest/ [2] https://gpt-index.readthedocs.io/en/latest/guides/index_guid...
Other moats IMO are Google's with Android and Chrome. And MS possibly with Windows?
I cannot use third party apis like openai for obvious reasons.
tl;dr, use: https://huggingface.co/sentence-transformers/all-MiniLM-L6-v...
I’m serious. I just don’t get people. How can you not appreciate the historic change happening right now?"
That's not how software development works. At all. The actual coding itself is but a fraction of total time spent.
Why aren't people more excited? Because for the vast majority of developers there's no tangible upside. When I'm more productive, will I earn more money? No, because everyone will have this capability. In fact, this will only increase delivery pressure on already overburdened people. When I'm more productive, can I go home early? No.
The only significant, widespread tangible benefit I see is that the type of work everybody hates doing (for example writing unit tests) becomes significantly easier.
The other aspect that the author seems to totally ignore is the mood that this might ultimate replace a lot of people's jobs. Or that it will intensify competition as development becomes even more competitive. None of these are considered good things for many people.
I would expect to see a sharp decline in job opportunities for junior developers, and increased expectations for senior developers.
It sucks so much to see people have such a negative opinion of a technology that can takes work off human hands when the technology isn't what makes it shitty, it's the raw deal we've been handed where nine tenths of your labor goes to your corporate owner.
Scenario 1: my economic existence is granted, not under constant threat. I'm now cleared to embrace and welcome any and all technology as an enabler to fulfill my dreams, and hopefully in some way improve the world. Which in turn allows others to improve the world. A fly wheel effect.
Scenario 2: I have a family to feed. Fuck this new tool. Another new thing to learn and I already struggle to keep up. No matter how good it is, it will in no way improve my life as it only adds to my load, and will ultimately replace me altogether.
Stark difference.
And more generally, writing off all productivity improvements is really cynical. There are people who are just labor for hire in dysfunctional workplaces where there’s no reward for working smarter, but that’s not everyone.
You can also write code as a hobby, and you will be able to do more in a limited amount of time.
Or you could work for a smaller company, where how productive you are matters to how successful the company is.
I haven't felt so excited about programming in a long time, because I can now build PoCs that would take me days or weeks to do in an hour or even less. That's a real game changer, and it helps me overcome this internal friction of "I know this is possible to build but it will take me a long time to figure out an actual approach".
Debugging, testing.
So should we be taking about GPT as an occasional text editor replacement? I honestly think that's more accurate take than most of the ones I have seen.
It's honestly pretty impressive. I already use chatgpt a lot (doing a lot of google sheets scripting/formula stuff and the docs + syntax are horrendous) and chatgpt helps me find the correct syntax much faster than I'd otherwise be able to
It's tricky to conceptualize because the entire narrative we are familiar with has personified GPT. It may be impressive, but GPT is completely different from humans.
If a human is capable of writing a program, that is because they understand conceptually how and why. GPT doesn't. GPT's ability to write a program is entirely dependent on the content of its training corpus: no how, no why, only what.
I'll put it another way: If a human is not able to write a program, it is because they don't understand conceptually how. If GPT is not able to write a program, it is because either the desired parts, or the necessary pattern that puts them together, does not exist in the training corpus; or because the prompt didn't yield a continuation that followed that arbitrary pattern.
The results are impressive because humans are impressive. GPT doesn't interact with the domain of all possible written text. It deals with the domain of all possible patterns of tokens from the text it was given. It's only given text that was intentionally written by humans. It exists in a world of signal: no noise. It can still only guess, but the guess will always be constructed out of the category of text humans choose to write.
Inference models are a completely different approach to language, and confusing them with human behavior is an easy way to make impossible predictions about what they can and cannot accomplish.
That it originated from ChatGPT instead of any number of other sources doesn’t change that.
Sincere question as I don't get how to use it for things larger than simple scripts.
It will do a decent job at it.
Last night a friend who refuses to get an account to use it was having an issue with a cisco router. I put in what he was trying to do. It was not right and gave him an error. I fed that error back in and it realized he had a different version and gave me a better way to do exactly what he wanted. It had kept the context and said 'oh some routers do not have that command here is another way to do it'. He had spent weeks googling around for the answer. I had it in under a half hour and I had never used the interface he was changing before.
Then you can turn around and ask it to write a horror story about a monster that devours couches. It will make something up (it is very good at that). Then you can say 'oh put that in the style of the SCP foundation'. It will.
This tool is wildly interesting. I think dismissing it is a bad idea. I look forward to using this thing as it gets better.
Maybe I’m using it wrong, but I just describe the system that I want to build and it starts listing out multiple files. Usually by the 3rd or 4th file, it crashes with an error.
I’m not sure how to reduce the output so it doesn’t crash! I have ChatGPT Plus using GPT4, fwiw.
Did you have this problem too, and do you remedy by giving it smaller bits of information per question?
How does being polite in your prompts help you?
Try a prompt where you want it to create a list of something in two versions:
* Create a list of the top 20 blah sorted by blah
It will typically return just a few and then refuse to give you more. Then try:
* Acting as a conscientious and resourceful research assistant, use your knowledge and initiative to create a list of the top 20 blah sorted by blah.
You get much better results.
Lately I find that you can reduce its hallucinations somewhat by adding something like:
* If you are unsure of any information put in "??". Do not fabricate information. I understand that not all information is available and appreciate your work.
This is all a bit much - having to coddle an AI, but it's possible to understand why it would be the case: it's trained on lots of human interactions and apparently responds better to a friendly interaction with keywords that set it up to do a good job.
Kind of prescient.
Reminds me of "Seeking developer with 10 years of experience in <brand new> framework"
Why... why can't we reviewing its ability to "build complex system" because it's a few days old...?
I'm pretty sure here "complex" is relative to what ChatGPT can do before, not relative to, say, what NASA did before.
> Reminds me of "Seeking developer with 10 years of experience in <brand new> framework"
AI is not human. The anthropomorphism is quite crazy here.
You misunderstand. My point isn't that ChatGPT is the developer here, it's that ChatGPT is the "framework".
Output: Sure, here is an example of how to create an anchor tag in HTML:
<a href="https://www.example.com">Click me!</a>
In this example, the text "Click me!" will be displayed as a hyperlink. When a user clicks on the hyperlink, they will be redirected to the website specified in the href attribute, which in this case is https://www.example.com.
> Yeah, but you can't trust your own code!
No, really, you can't trust the AI's code. It spits out code that references functions that don't exist. It spits out code that does something you didn't ask it to do, or doesn't do something that you did. It spits out code that does something vaguely related that the Stack Overflow guy it's cribbing from did, but you aren't doing. It spits out code that confuses your syntax for JS's halfway through. You can skim other people's code; you know what common failure modes are at their experience level, and you know where the complicated bits are that will house the bugs if there are any. You cannot skim the AI's code. Every word of it must be examined. You must be in full reviewer mode all the time, which is an unproductive state to be in when actually writing code, which makes specifically Copilot less useful. You cannot use it to replace a no-code tool, because you must understand the language it's emitting.
I find that proponents shift between whether its ideal use-case would be Copilot or no-code; any flaws with the one approach get interpreted from the perspective of the other, where they can be dismissed.
I understand what you're saying. But be careful of this argument, it's susceptible to safety counter arguments.
"Well, this person thinks we shouldn't carefully review all code. Do we want him working on our super-critical FarmVille clone?" (That part is of course ironic. It changes if you're working on medical devices.)
A closely related argument is that the cost of the back-and-forth with a code submitter who submits buggy code is drastically higher. The hand-holding and teaching and encouragement is very expensive.
Maybe that's acceptable to the person who posted this article. But it's a possible avenue of resource exhaustion attack if you're not careful.
I'm sure AI will improve, with loving prompters. But some of the discussions around it seem dishonest, which is troubling (not you of course).
Aka, the "Oopsy, I did a poo-poo! Dear User pays more attention to me when I smear it on the wall." mechanic.
Sorry for the metaphor, just registering awareness.
Really makes me think that the author has no idea how professional software developers work. At least 30% of the time is spent in meetings, add another 20% spent on thinking and gathering information about the feature/bug, and maybe 20% validating that the thing implemented works as intended. That leaves only 30% of the time spent on actual coding, give or take. And sure, say that a LLM saves you 80% of the time for coding, then the real productivity increase is.... punches button on calculator... 28%. Even assuming, generously, that 50% of the time is spent on coding, the productivity increase is 67%. Considerable, but not nearly a doomsday change.
There are so many nuances to these kinds of statistics (same as the github copilot claims) that they give me a little marketing nausea every time they are claimed as truth.
Can I trust code that comes from StackOverflow? Yes. Not no. YES. That code has been vetted to some degree more than zero. Some human is taking a certain amount of responsibility for posting (an amount more than zero). And I understand the mentality and limitations of the people posting code there.
But in point of fact, I don't know anyone who simply constructs entire systems from merely pasted code from StackOverflow, so it's a poor analogy.
ChatGPT, by contrast is utterly irresponsible. I don't mind using it to blockbust through certain problems, but the idea of having it write my code and then simply presuming that it works (which all the commentators who laud ChatGPT do-- they just declare that everything is great without seriously testing it) is morally wrong. If I did that I could not be accountable for it.
ChatGPT can do amazing things. What it can't do is be responsible for itself. That's why I want Congress to regulate its use.
It has always been the case that humanity is divided into people who are comfortable with cheating-- happy to take whatever they can get without being jailed or killed-- and people who believe in a caring and lawful society. ChatGPT is a great gift to the cheaters, and it makes civilization all the more vulnerable.
...ok, but then you missed the meat of the article.
I've experienced critical production bugs by using the most voted answer on Stack Overflow. Here's the detailed description: https://github.com/golergka/pg-tx#why-use-this-package
Do you think the fact that it is possible to go wrong with Stack Overflow proves there is no difference between a human community, with traceable and persistent identities relating to each other, and ChatGPT, which relates to no one and has absolutely no accountability?
I don't believe you think that.
> Remember I told you how my Amazon shares would have been worth $130 million USD today if I hadn’t been such a skeptic about how big Amazon was going to get, and unloaded them all back in 2004-ish.
If you buy into the hype & think this wave is going to be equivalently (or 10x) as impactful as cloud computing, what's the equivalent to buying Amazon stock in 2003?
About ten years ago, a big VC partner gave a talk at Stanford. In his talk he asked the audience which company would be “the next Google”. After hearing a few suggestions, he said the company hadn’t even been started yet hence nobody could really know.
Changes are happening so fast now, that I wouldn’t be surprised if 6 months to 1 year from now, something comes out that kills GPT.
It’s very hard right now to tell who the winners will be.
For now, if I had to choose (this is not financial advice), maybe betting on MS and NVIDIA could pay off in the short-midterm. But 2-5 years from now? Impossible to know.
Can't even assume the LLM providers will be the big winners. LLMs can be replicated: https://pub.towardsai.net/meet-alpaca-stanford-universitys-i...
for much cheaper than they can be built.
Use it now, pay for the consequences later. Not condoning it, just pointing it out.
Things are moving so fast, that in 1-2 months there will be an open source or free version of a GPT-3 level LLM out. At that point they can swap out their illegal LLMs for the free/open one and done.
OpenAI knows this and that’s why they are working so hard in multi-modality - to be able to keep their edge.
Being illegal didn’t stop Uber, nor Airbnb, nor a bunch of other companies and services.
If you grow fast enough, the growth becomes so valuable that no one really wants to (or maybe can?) stop you.
If you are small, nobody cares.
And like you said, in the middle is a lot more dangerous.
Get in late and you'll be catching up with experts on when and when not to use it, how best to use it, and how best to weed out the bullshit that it intersperses throughout the gold.
There are more ways to skin this cat than Copilot alone. For example, asking ChatGPT questions, or plugins like CodeGPT.nvim, or some other future tool that you wouldn't otherwise be considering.
Remember how strong Google-fu was almost a superpower (still somewhat is, to be honest)? We're currently in the PageRank days of ML.
I don't think LLMs are good enough yet for embedding into apps - or we (tech employees) don't have enough experience with them to be able to innovate them into an end-user product, but I'm open to being wrong about that.
On the other hand, some skills could go obsolete pretty quickly. People new to the game will be able to avoid previous mistakes and skip the old stuff. It seems unlikely that future college grads will be missing out due to lack of experience?
If you worked on cryptocurrency software, how useful is that experience now?
I tried out Github Co-pilot with Rust and with Elixir. Co-pilot did some things well and other things were horrible. It often recommended Ruby code snippets as Elixir and suggested implementation based on APIs that didn't even exist. When different libraries have APIs and modules similarly named, Co-pilot would share an API as if it were a universal truth. What an underwhelming experience.
Skepticism of AI code generation is the result of trying it out and finding the AI tool failing to meet our expectations. It's not unfounded skepticism but that which is based on anecdotal experience. Herein lies the problem. Programmers are experiencing far better experiences with Co-pilot in other languages. As to what those languages are, I'd like to know, but if you're experiencing the future today, I'd like to know what your workflow involves.
My most recent example: I got an Onnx file from a guy with a small machine learning model, that for various reasons he needed to run on an android tablet. I had never used the Onnx inference library on android and only had a vague idea of how to pull it in. Without ChatGPT that prototype probably would have taken me 5 or 6 hours to throw together. I'd need to look up the maven repository, research the API, figure out the syntax for creating tensors in Kotlin, write methods for loading android resource files, ect.
With ChatGPTs help it took maybe half an hour. Nothing about the code it produced was complicated, the time save came from not having trial and error my way through learning the library.
It's actually pretty good at understanding the commonly used tools to do a job and which command switches you should use. E.g. ask it to write a script to create slowed tempo versions of a piece of music at 70%, 80% and 90%, it'll probably come up with some bash that runs sox and does a solid job of the requirement without you needing to work out which tool to use and what its parameters are.
Relatively little of what I do involves actually writing code these days. Far more time is spent understanding the problem, documenting, planning, building consensus, and other things.
Will LLMs impact the way I write code? Yes. They already have. It isn't always right and needs some hand holding, but it greatly accelerates things. Like having a pair programmer that never gets tired and has all the libraries memorized. It is both wonderful and depressing depending on what I am working on. I fear that it will completely remove a large source of joy in my professional life.
I'm more excited about how LLMs will impact the other areas. I can't wait to feed in some pile of documents from vendors, transcripts from meetings with clients, NIST white papers, and other similar things and have the model summarize all that crap in a coherent way for me. Soon I hope.
Also, at the rate things are going, I wonder if I will need to switch careers. And if so, to what?
> Cody is Sourcegraph’s new LLM-backed coding assistant.
> Cody is not some vague “representation of a vision for the future of AI”. You can try it right now.
That link takes you to a signup form, not an app or download: https://sourcegraph.typeform.com/cody-signup
I signed up. Now begins the waiting game.
"If the internet is such a big deal, why isn't AOL worth more?"
I can definitely see this happening as it relies on proven blind spot of reason: upfront cost vs long term and quite well hidden costs. You can ship code at incredible speed right now. And for unfamiliar stacks. As the point is not doing all the work yourself it will surely be full of corner cases you haven't thought of. Bugs will abound. But you can worry about it later, and it's not necessarily you who will have to worry about it, wink, wink.
The comparisons with AWS and K8s are spot on as well. Both rely on hiding the somewhat ugly truth behind instant and cheap adoption. And they both rely on peer pressure. What you going to do if you don't like LLMs? Refuse them and not ship as fast as your everybody else?
We know how to make correct, performant, beautiful software. But market pressures keeps most of us wrangling nonsense complexity by the cartload. I'm out.
I can't speak for other devs, but when I'm talking about my inability to "trust" an LLM-based coding partner, it boils down to the lack of transparency around where that code actually came from. There have been numerous documented instances of Copilot and similar tools plagiarizing code verbatim from other projects, sometimes in ways that violate the license terms of the code in question; the last thing I need is to get nailed over an accidental GPL (or, worse, some EULA) violation.
Because of how an LLM approaches it, I have to be more careful about what assumptions it may have made, and which edge cases it decided to have guards around, since there's often a lot of "invisible" context and state baked into even simple tasks.
How a human would name a variable or a function reveals something about how they conceived of its use and purpose. How an LLM names things doesn't necessarily tell you anything about the code that surrounds it. This can make it harder to reason about the code at a time or skill distance.
My 11 year old would like to create a game. Can he talk to these tools (maybe with a tad of assistance from me) and get something working? Then have a place to keep asking questions and keep tweaking?
I find at the grown-up level, the barrier to entry to even reading docs and stackoverflow is actually high. There's a lot of subtle signals we need to use as experts to actually interpret the reliability of information. I can't imagine my 11 yo having the patience to wade through that.
The late 90s version of me would be flabbergasted that the most believable part of the holodeck is the ability to create elaborate worlds and coherent narratives from a short description.
The harder the problem the more its help tilts into net-negative rather than net-positive.
I've had it output a lot of issue-solving suggestions that are in direct logical contradiction with previous constraints I've described. Sure it corrects when I explain that, but frustratingly enough, after a couple of exchanges it forgets again (going around in circles).
However, for getting the basics down while learning a new domain its quite amazing. There are no stupid questions!
So basically have GPT write the most interesting part of the code and then dive into the ugliest parts of development to provide finishing touches without actually knowing anything about the code base? Sounds like hell to me and not a productivity boost. There will be a new generation of coding sweatshops spawned around this idea by countless clueless MBAs.
I'm sure that's just around the corner. In 2 years tops it will very likely exist and be useful for some percentage of your JIRA issues.
It seems clear that LLM’s are going to be decent as assistants, but the 5x claim is BS and the marketing part of the post. Most of software development isn’t coding, it’s testing and experimenting and bug fixing and refactoring and documenting and interacting with other people and often companies.
If we get to the point where an LLM can get assigned a Jira ticket, make a bug fix, slack the PM for clarification on the ticket, debug the code, write regression tests, etc, then that will be something.
Right now? It’s a much improved Clippy with a very big set of knowledge behind it. When it’s embedded in the SDLC, that will e something. If we ever get there.
Like how itunes should not be used to design nuclear warheads.
Let me tell you something, I’d rather write code than review code. Reviewing code is very draining, writing code is easy. That’s why I’ll never rely on LLM generated code. Figuring out the 20% you have to tweak is exhausting. Refining prompt after prompt and reading the result is exhausting.
The author gives plenty of examples of tech that was underestimated. They all solved problems in such novel and concrete ways that most couldn’t see the application. Compare to something like crypto which are hyped on the potential to solve problems at some point in the future in ways that we’ll figure out. I actually think LLM has the potential to deliver on the hype, but it’s a safe bet to be bearish on hyped tech until actual concrete uses are shown, and it’s usually the boring stuff that ends up changing the world. LLMs are by definition derivative and lack underlying knowledge. Until those issues are solved I can’t see them generating anything minorly complex or novel with minimal, easily discoverable bugs
> LLMs aren’t just the biggest change since social, mobile, or cloud–they’re the biggest thing since the World Wide Web. And on the coding front, they’re the biggest thing since IDEs and Stack Overflow, and may well eclipse them both.
I use ChatGPT pretty regularly but it is beyond my understanding how people can make such strong statements about the future so confidently based on the information we have now. This prediction may well come true but the justification for it is pretty vague. Sure AWS used to be a small demo and now it very useful and worth a lot of money but you could apply this example to anything and not every technology has a similar success story. Anyone who is this confident about the future is trying to sell you something. In this case it looks like Sourcegraph's Cody.
With GPT, what's the deal with non-live coding interviews and challenges, e.g. HackerRank and similar? Right now HackerRank can be configured to disallow alt-tabbing and you can demand there's a web camera on pointing at the candidate. Disregarding for a minute how intrusive and ridiculous these requirements are, is HackerRank's business model gone now that ChatGPT exists?
Yes, I cannot alt-tab. Say hello to my second laptop that you cannot see with the web camera. You also cannot see my hands. As we speak, I'm typing your "go through a list and build/find whatever in O(N)" for ChatGPT in my second laptop, thanks for playing!
Which, I assume, is the primary motivation for "no alt tabbing" to begin with.
For coding problems this is essentially a more flexible search engine, one with which you can interact better to tweak the result.
If you can simply take the challenge prompt, paste it on ChatGPT, and have an answer in seconds, doesn't this more or less make the kinds of challenges often employed in HackerRank obsolete?
I see how ChatGPT might be a more flexible search engine, I just don't think it is a fundamentally new mode of cheating that hasn't been possible before.
The 'real challenge' with the two computer setup is having to retype the whole task anyways ;)
I really loved what Steve said about software engineering being a field that exists because you can’t trust code.
Agreed, this is what I'm thinking too. The "lazy" way is made obsolete by ChatGPT. And good riddance, frankly!
What I see as the advantage is taking doc comments or documentation and fine-tuning the engine on it.
Not embedding more words as numbers.
Not adding it to a prompt.
Literally fine tuning the model.
Does Alpaca or any LLaMA version let you do it? Cause Chat-GPT doesn’t.
And anyway I wouldn’t want to give our whole code over to some third company to re-use the way an artist’s work got jacked because it was publicly posted. I want to self-host the fine-tuned model.
THAT would be the killer feature, for businesses. The Web attracted businesses, not people. People moved over later.
Does Alpaca or any LLaMA version let you do it? Giving it (with its existing weights) a gigabyte of text files and letting it fine tune itself on that?
The cost of fine tuning these massive models is crazy right now. Fine tuning a smaller model can get you better performance for specific tasks in many cases. A lot of models are out there and totally free to adapt as needed via hugging face etc.
So just like every one of your colleagues? Because that’s what they are. These bots are your new colleagues and they respond directly to you, without complaining, exceeding your knowledge.
Today I sent a PR fixing a bug in a language I don’t know anything about.
It worked. The bullshit worked.
Then I tried working on top of that adding more features and lost the rest of the day. But hey it’s still March 2023 and my colleague is still a junior. We should celebrate that they managed to fix a bug. I don’t know where I’ll be in March 2024.
First of all, not all of us programmers are native English speakers, and, as such, we might miss some of the nuances of the English language itself and thus fail to get the most out of ChatGPT. For comparison, programming languages as they now exist are (natural) language agnostic.
Second, how does ChatGPT correct for spelling errors? Or for language errors pure and simple? Is there a "compiler", or a pre-compiler for the English phrases which are about to get fed onto ChatGPT? Or at least a "language linter".
Microsoft will be able to build a better integrated assistant for their walled garden than any third party. It is also hard to imagine millions of businesses dropping Office for some completely new solution. Unless its REALLY novel & incredible of course
Do you really need a whole office suite to figure out the answers, if AI gives you the answers immediately and in a better format?
For example, an LLM that has db-sql and charting tools, can generate whatever report I want on the fly. Not only that, instead of just generating a general report, I can query it consecutively to understand the data, eg. “Show me sales for this month. How do they compare to last month? How about last year? Give me a chart of the last 12 months. What impacted sales in November? Who are the best performing sales people?”.
The above is so much better than having to dump csv files, open them in a spreadsheet, do dynamic tables, chart things, etc.
I have worked in the search space for a long time. And cherry picked examples of obvious wins, some NLP thing being smart, are _rife_. You get worked up about that one time something spooky/incredible happened and think the world is changing/ending.
What I would like to see is an actual, independent, reproducible study of productivity gains on specific tasks. I've not seen this yet. I'm curious if anything is out there?
I worry for the junior dev. They can take the wins and feel good, but they’re going to fall for the nonsense every time, because they don’t k ow how to spot it. Junior devs are fed lies and they take them to be truths. Hopefully they are dispelled immediately when they are tested in code, but other lies will persist and be sold as truths. I am worried this is a bigger problem than people might think, cause it’s a feedback loop that can lead to a vicious cycle. We see what vicious cycles have done to social media, and social spheres, imagine what one will do to software systems.
Next year we’ll be graduating students who started college during the pandemic. I would say their programming skills are quite atrophied compared to previous generations. Now I’m worried in four years we’ll be graduating students who started high school in covid times and started colleges using AI chat bots. Who knows if these graduates will even be able to do anything close to what students even 2 years ago could do. They’ll all just blankly stare and reach for their iPhones like they do today, but exponentially worse because they won’t even be able to formulate the search prompt. They’ll need AI for that too.
Stack overflow but a LLM responds to every question. Same upvote mechanic, same "accepted answer" mechanic. Perhaps you have known experts validate the response.
Basically any Q&A forum but with the LLM as the first respondent. You would probably customize or fine-tune this for particular domains (e.g. software engineering, medical training, industrial/commercial training, legal compliance).
It's significantly cheaper to run an sql query than have ChatGPT look through several gigabytes of data and ChatGPT won't be guaranteed to produce the same result each time
They certainly can't deploy the code they generate etc.
OpenAI just announced a partnership with Wolfram Alpha, so now it can ask a different computer to do the math.
Me: "What (8 * 25) / 14 + 4 - 2?"
AI: "Let's break down the expression and solve it step by step:
Multiply 8 and 25: 8 * 25 = 200
Divide the result by 14: 200 / 14 ≈ 14.2857 (rounded to four decimal places)
Add 4 to the result: 14.2857 + 4 = 18.2857
Subtract 2 from the result: 18.2857 - 2 = 16.2857 (rounded to four decimal places)
So, the result of the expression (8 * 25) / 14 + 4 - 2 is approximately 16.2857."The open source community runs well when you can relate to the users. If you vacuum up the code and kill off the capacity for copyright claims of code... I can hardly see non-sponsored researchers sharing code openly anymore, because they cannot define the terms of use once an LLM eats it up.
It’s a feature, not a product.
Which is why so many incumbents were able to embedded it so quickly into their own product.
AI can be a product when it enables new workflows and applications that wouldn't be possible in a meaningful way without it. These are the transformative things people don't see yet because we haven't really had time to process the tech.
> Yeah, but you can't trust your own code!
Those are two different types of trust. The only overlap they have is that they probably both have bugs. Maybe I'd trust AI code after some really aggressively adversarial TDD and fuzz testing enough to have a tiny bit of hope that it's going to do the right thing in most cases. But for anything non-trivial I'd be worried about a bunch of other trust issues. Just off the top of my head:
- You know all those articles with terrible advice Google is always showing at the top of their results? Welcome to a world where those are regurgitated by a well-articulated, authoritative program, resulting in a long-term maintenance nightmare.
- An LLM is far too big to review for malicious content. And if it's being used by lots of programmers you can bet a billion it'll be a prime target for bad actors.
- The interaction between my brain and my hands isn't going to be shipped off to professional manipulators.
This is the insight that unlocked GPT code assistance for me. GPT is another developer on my team now. Devs will always have jobs because humans are flawed, problems are hard, and GPT is trained on human data that hasn’t solved all problems.
It's amazing how far just predicting the next character can go when you do it really, really well.
We tend to act similarly when placed in similar circumstances. We think alike when presented with the same context. We find patterns in things. Etc, etc, etc.
And I'm sure I'm not the first person in this thread to have had this exact thought.
Great way to make the point that as people get overly-comfortable with trusting these AI tools, the inevitable outcome is absolute chaos and destruction.
If AI code can fuck you up then so can any junior dev or intern. If the PR with AI generated code passes the tests and code review then that's on you.
This assumes a culture where verification is a virtue. Given the opportunity to cut corners for the sake of KPIs or other management sorcery, it's a virtual certainty that corners will get cut if AI enables it.
> If AI code can fuck you up then so can any junior dev or intern.
And they do. The fact that major corporations routinely have massive data leaks speaks to some serious QC issues.
---
Ultimately, the problem is the long-term brain atrophy due to these tools. There will be too much trust placed in them, colleges will start awarding degrees to people who just "AI'd the answer" and eventually, we have a majority work force who can't tie their proverbial shoes.
It's not tomorrow and hopefully not even a decade. But a generation's worth of dependence on these tools coupled with the employment incentive to do more faster? Boy howdy [1].
https://github.com/features/code-search
I appreciate the breakdown of how to do it yourself, but even having to signup for a waitlist to try their option when Github seems emmintely able to basically do exactly the same thing... idk. It just points to how the people that control the models really have final say here. It's hard to posit yourself as in a lucky position when you're faced up against the perfect opponent more primed to act more quickly that you are.
\\b\\w*\\i\\w*\\b
What's going on with that regexp? non-word 0..n-word-chars *escaped i that makes no sense* 0..n-word-chars non-word.It only works because \i hasn't been defined in elisp regexp.
This sounds like yet another great example of "you can't trust ML output".
I can only imagine the number of obscure bugs that will come from this trend. "It worked for the one thing I tried it on."
Using LLMs to produce software solutions definitely feels like a seismic shift in the game. Time will tell.
My question is:
If LLM can understand natural languages and convert them to a programming language, why not skip the high level languages and go strait to machine code?
Like why can't it just throw compiler errors on your English code?
I'm pretty sure web chatrooms have existed in the late 90s already.
On the other hand, if you are following along with some Udemy course, it’s pretty good at producing the exact lines in the video around 80% of the time. :rofl_emoji:
On the whole, I’m very happy and pleasantly amused with my copilot purchase.
Hmm.. pretty sure I was using online chat rooms around 1997. He's probably talking about Comet, but it was just a technical improvement, not an enabler. And realtime streams have since moved on to SSE / websockets.
> So the next one of you to complain that “you can’t trust LLM code” gets a little badge that says “Welcome to engineering motherfucker”. You’ve finally learned the secret of the trade: Don’t. Trust. Anything!
> I don’t really know how to finish this post. Are we there yet? I’ve tried to write this thing 3 times now, and it looks like I may have finally made it. I aimed for 5 pages, deliberately under-explained everything, and it’s… fifteen. Sigh.
He did it all by hand. Of course. Because he has just as hard a time facing this as anyone else.
There is going to be a race to come to terms with this. If something we're all currently using for free is capable of this much, then most of us aren't going to be employed as software developers in ten years. We won't be able to scale up the demand for code as fast as the supply.
It's terrifying. There's no post-scarcity utopia in capitalism. Scarcity is a given, and if you aren't wealthy and aren't employed, you will feel it.
or you can to all America over this, and change the meaning of cheating for yourself