I've only found it useful for explaining basic terminology or concepts for a topic I'm not familiar with. Though this is the most dangerous use case since I'm not familiar with the answer, I can't easily fact check it.
I've only found it useful for explaining basic terminology or concepts for a topic I'm not familiar with. Though this is the most dangerous use case since I'm not familiar with the answer, I can't easily fact check it.
It writes lots of code before I even think about what I was going to do in that file. Half the time it just works outright, and if not, with minor changes.
It is fantastic when I have to build functionality in languages I am not great at.
I regularly copy and paste code from my colleagues PR's and paste it into GPT-4, it explains it perfectly and gives tips on how to improve it (which I add as comments to my PR review)
> It writes lots of code before I even think about what I was going to do in that file.
If you haven't even thought of what you want to do in the file, how are you prompting Copilot to write the code you need? Am I understanding correctly that Copilot not only generates the correct code, but it's also able to deduce what code you meant to ask for without any input from you? I don't see how this is possible except in the uncommon case where you are bringing a file to be more in line with other files in your code base that have a similar structure.
I do think Copilot can be useful, but I'm confused at how different my experience is from what others are getting out of the tool.
You can know what you want to program, but not how you would approach the code. "Hey ChatGPT, how do I draw a curved line plot in JUCE?"
If I just ask it "Can you write me a script that gathers the current stories on the front page of hacker news and sorts them by number of comments, then prints the title and number of comments?" it writes a script in Python using requests and BeautifulSoup4.
Otherwise, even just creating a function, generally they will be filled out automatically.
I suppose I'm quite big on naming (vars and functions)
Maybe if your language or code is more terse, it might struggle more.
I wonder if it's a language difference thing (I use C# mostly), but my experience with Copilot has been nothing short of mind-blowing.
It's not writing every line, and when writing truly new feature code it's less useful. But here are two patterns that I've noticed it is especially and consistently good at as a potential place to look for initial value if you are skeptical:
- If you have any kind of repeated pattern, even very complex ones, like performing a set of operations on a set of objects, or initializing a set of things, or whatever, it will guess everything else after the first line 90% of the time. This stuff is almost always plumbing/wiring/boilerplate-type code, so pure time gained. Think about Excel filling incrementing numbers down a column, but for pattern-matched lines or blocks of code.
- For any reasonably testable class, if I write the name of the unit test, Copilot will write the entire unit test perfectly 90% of the time, down to variable names and //Arrange//Act//Assert comments that I stylistically prefer. Seriously, it's sort of scary how good it is at this.
It suggests solutions to problems, I add the complexity. So far my code has been both clear and performant. This should not be the case at my level.
What does this mean?
Copilot was released in March of this year. ChatGPT 4, which the community considers to be the only version of ChatGPT to be competent enough at coding tasks, was also released in March.
My guess is that you didn't mean that statement to be a fancy way of saying "I've been coding professionally for around 6 months", but I don't know what else you mean by the "age of AI".
As opposed to ChatGPT which spits out dozens of lines of code based on a request and therefore requires more involved editing after the fact to understand what’s happening.
I can understand using chatgpt 4 but copilot??? copilot can barely autocomplete... and half the time it does it incorrectly when variables and other stuff are involved
Which isn't easy, because there's virtually no documentation for it.
Once you get the hang of it it's incredibly useful. I really feel it when I'm working in an editing environment without it now.
I am in the latter group, and I don't find it all that helpful. There simply aren't any tools that can plug into a massive codebase with millions of lines. I never just work on one specific repository either - a change generally involves multiple repositories.
On the other hand, if you're writing a smaller standalone utility, or working on a greenfield application, it can be helpful (with all the useful caveats about hallucations).
I think that a lot of ML folks, especially those in research, fall into the former category, which is why there has been so much hype.
If Sourcegraph's offering (https://about.sourcegraph.com/cody) has limits I'm not aware of, this is just a matter of time. But in my experience, this is a "nice to have" rather than a "must have" when it comes to the benefits of GenAI as it applies to coding.
I have interviewed hundreds of engineers across the entire skill spectrum. I think GPT4 right now is about on-par with a mid level developer.
It makes a lot of mistakes, especially around counting, and usually can notice and fix them if they're pointed out. Humans do this all the time -- how often do you get a compiler error from something silly?
It often misapplies interfaces on the first go-around. Again, this is just like humans. We make mistakes, notice them, fix them. If you simulate this by telling GPT that its code produced an error it will often correct itself.
The absolutely killer feature of GPT4 is that it has these skills in every subject. It's fluent in kernel operations. Databases. Networking. Various UX frameworks. Any language.
It's definitely not perfect. But, if the alternative is hiring a mid-level human engineer, GPT4 is a really compelling alternative.
Hooking it up to automatically run the code in question and examine the output is a trivial undertaking - many folks have already done this.
At writing code, that might be a slight exaggeration. At explaining code it's at least at mid level, possibly better.
I’ve used chatgpt recently to build an entire c++ game with sound, keyboard and mouse control, ai, etc, having never used c++ before. I’ve also never built a game before. I’ve worked with a lot mid level developers and I’ve never seen any that can write me a a c++ file giving me working code for a collision system in 30 seconds. Let me know if you find one though.
I’ve got 20 years of experience in the tech industry and chatgpt is far far better than any mid level engineer I’ve ever met and I include FAANG engineers too.
Chatgpt can answer any leetcode programming question in essentially every programming language in existence in less than 20 seconds.
Chatgpt can analyze classes for errors written in every language in seconds.
Chatgpt helped me write animation, ai, sound handling code in c++ and helped me build a full working game, the total code generation time for chatgpt “thinking” was on the order of one hour.
I don’t know any senior engineer who could help me (a c++ newbie) write a whole working game engine with all the necessary systems in less than one hour of “thinking time”.
I’m starting to think that all the “naysayers” of chatgpt simply used it once or twice, saw one error and dismissed the whole thing. Or they’re simply afraid for their own skills and avoid chatgpt as the alternative scares them too much.
In my experience, if you compare the unedited output of ChatGPT to the unedited output of someone with a couple years of experience in the given domain, ChatGPT will have more subtle bugs than the human provided output.
Hence my between junior & mid assessment. It doesn't mean that it's useless though. Having a tool at that level of experience for every domain in the world that can cook up code near instantaneously is damn useful. And I assume it'll only get better.
I've found GPT4 moderately helpful, like getting help from someone with a wide but shallow experience of lots of things.
> I can't easily fact check it
Why not? How did you find out things about basic terminology or concepts you're not familiar about before GPT? Apply the same methods, although you just have to fact check rather than coming up with what to check, so removes a small initial discovery, at least for me.
Because if it tells me something that could just be found on google, I could have found it there first and not have to do as much fact checking. And if it tells me something I can't easily find, I can't tell if it's got some deep insight on harder to locate info, or if it's just made up.
I’m not well-versed in JavaScript (I’d like to be, but neither me nor my company have the resources right now to commit). When JS work comes up, I’ll often Google/ChatGPT to point me in a general direction before pulling up technical documentation on whatever comes back. I’m pretty good at learning, but sometimes it takes me a while to find out where to start.
However, there’s a few things like SQL and JavaScrip that are likely to stick around mostly unchanged for your entire career. The risks are therefore much lower.
IMO, a reasonable heuristic is spend around 10% of the time you expect to work on technology in the next year actually learning the fundamentals until you feel comfortable. Often what looks like a major time sink goes away when you stop fumbling around.
I'm not interested in starting at the beginning and learning a whole new language or API, I just want to get my task done, so I can ask GTP to get me started, then I can run the code it writes and start building tests and things to verify it works as expected.
With everyone and all the news talking about it constantly inevitably people are going to start rolling their eyes.
This is bias. 48% is half way to 100%. Once it reaches 100% you don't have a job. You realize that right? It's halfway their to taking your job and you're underwhelmed. The bias is on your side.
It's typical. It's like an indie band is only popular when not many people know about it. Once everybody starts talking about it all the time it loses it's popularity. People start thinking it's not cool anymore. That's the bias you are suffering from.
Many programmers also do work that doesn’t just involve stitching APIs together; GPT consistently fails badly when forced to reason, making it useless for many tasks. For example, ask it to implement a common algorithm like SHA - it will do a decent job. Then ask it to do it with an arbitrary limitation which no implementation it was trained on would have (e.g. use no integer type wider than 8 bits). In my experience, it cannot achieve this task with a correct result, even with significant hand-holding.
I meant 38% aka 40% made a mistake on my math.
It's not a big assumption. It's the most reasonable assumption. When you drive 50% of 10 miles the next 50% takes the same amount of time. It's the default assumption.
I understand where you're coming from though. Many products made by entrepreneurs follow that model where the remaining percentage is always harder than the beginning green fielded tech that was built. But can we honestly say that this is what AGI will be like? The LLM was an unexpected jump forward by an unprecedented amount. It could be we fill the gap to 100% with another such jump.
And it's wrong, as anyone with experience in ML will tell you. 60% is easy. 70-85% is not too bad. 95% is hard. 100% is effectively impossible. This is the core problem of ML systems. They're good enough for 90% of cases but that last 10% can be incredibly important, to the point that the systems become unusable.
You seem to think the topic is about the current state of training algorithms and how it's hard to bring the model to 100 percent. That's a different topic with a subtle distinction.
You’re implying 100% reliability is actually attainable. If that were the case, wouldn’t that mean the halting problem would have been solved by AI? I’m not an expert but I’ve heard that’s like one of those fundamental laws of information theory that really can’t be broken.
> It's typical. It's like an indie band is only popular when not many people know about it. Once everybody starts talking about it all the time it loses it's popularity.
I think applying the sociology of hipster band fans to LLMs is a mistake and I’m not sure how the vogueness of the technology correlates with the correctness of the actual models. Sometimes an early technology seems docile or useless at first but eventually reaches ubiquity and in retrospect the utility is obvious. But sometimes (more often than not) standard adoption curves don’t make sense to apply to a new technology because it isn’t useful enough to go through an adoption cycle. I think there’s a temptation to apply the analogy of something like the early internet or smartphone to LLMs but those are networked products. LLMs don’t really improve with the various applications built on top of them if the LLM is itself fundamentally broken or faulty to the point that it is unsafe to use in practice. Furthermore given the massive amount of premature hype due to the AI zeitgeist, you can safely assume enough user hours have been spent messing with LLMs to get a verdict on their utility. Unlike the feeble technology that takes a longtime to reach a scale to know its utility, we don’t have to wait 10 years for 100m people to try LLMs, it happened in a week. My only point being, I think we should be apprehensive about trying to draw analogies of other adoption cycles that structurally are very very different to this one.
Obviously, I, like probably everyone else on this website would love LLMs to be reliable to a high degree. Just this morning I had such a good use for an LLM that I was seriously considering building (and admittedly still am pondering) but the second I started to think through the LLM faultiness, I had to consider the complexity of the safe guards and weigh if it was really better to use GPT or just write a nasty regex script and constrain the problem. I’m leaning heavily towards the latter but I’d much prefer a silver bullet if it really killed vampires. Until then (if that day ever comes), it’s lead bullets for me.
Nah even 100% reliability isn't attainable by a human. 100% obviously doesn't imply solving the halting problem.
100% as in 100% as reliable as a human. Even surpassing a human. Sort of in the same way as a computer now beats humans at chess.
>I think applying the sociology of hipster band fans to LLMs is a mistake
Why would it be a mistake? Human psychology is similar across all spectrums. What happens in one area is likely possible in another area. If it happens among hipster bands it can happen among hipster technology fads.
>I think there’s a temptation to apply the analogy of something like the early internet or smartphone to LLMs but those are networked products. LLMs don’t really improve with the various applications built on top of them if the LLM is itself fundamentally broken or faulty to the point that it is unsafe to use in practice.
This is a valid speculation. But it's speculation. Basically your saying that LLMs are fundamentally broken and stuck at 38% forever because of fundamental and permanent flaws. The jury on that one is still out. And your point is highly, highly speculative.
We see quantitative improvements on LLMs constantly AND this is a nascent technology we don't completely understand yet. The most probably and logical conclusion is to follow the technological trendline. That trendline is pointing up.
To speculate on fundamental flaws of the LLM when people don't even fully understand what's going on with the LLM is illogical because you can't derive conclusions from something you don't understand. We can only generalize the trendline and constant improvements we've seen in AI for the past decade. Again that trendline is pointing to further break throughs in the future.
>Obviously, I, like probably everyone else on this website would love LLMs to be reliable to a high degree.
No this is not obvious to me. I disagree. I think some people are like you but other people, for example Geoffrey Hinton are in the apocalyptic camp. Personally I'm in the middle, I think it could go either way. It will definitely harm a segment of our society by taking over work, but whether the benefits of AGI outweighs the harm remains to be seen.
Your point was on whether everybody wants to see AI take over. To that Geoffrey Hinton doesn't want AI to take over, while you do. I for one am not sure.
As for whether Geoffrey Hinton thinks AI is legit or not is a different story. On that topic he's on my side, but that was besides the point I brought him up to point out that your take on "everybody" wanting AI to develop further is incorrect.
> and the doom and existentialism is a marketing ploy to a world imbued in conspiracy theory and institutional distrust in order to compensate for a wildly over promised and under delivered product.
Other way around. With all the money and business interests going into LLMs business interests are promoting LLMs in a future that you want and that is not apocalyptic. The conspiracy theories aren't a thing. It makes no sense as those theories don't align with where the money is being thrown.
>But hey, when you gotta raise money, you gotta raise money, and as you put it, the "trend line is going up" and that's all that matters.
Bro. I am not saying "trendline" as if it's something I have to keep throwing money at to support.
I am talking about a mathematical projection based on data. The pace of technology in the past when graphed points to an ever increasing line on a line graph. When you take the slope of that line and use it to do a quantitative prediction, that line just points up. That's just a fact of reality. The logical outcome of the data we see.
Look. People in reality are basically never this gigabrained. Maybe consider as your first-line theory, that when people say "AI will destroy the world", that they mean to express that they believe that AI will destroy the world?
But to be clear I’m not saying this is a methodical highly coordinated ad campaign amongst multiple companies. I don’t think human beings are that competent. I think it’s way more grassroots and feeds on cultural tropes that have existed since at least the 1950s but probably the early 1900s. I think Anthropic’s leadership probably genuinely believes their own bullshit but I also think they understand that that same bullshit has raised over a billion in funding.
What this study looked at was feeding StackOverflow questions to LLMs and then looking at the quality of the code. If you think a programmer's job is just turning an english-language description of a function into isolated code that never gets modified, I don't know what to tell you. In my opinion, anybody like that should not have a job today, never mind the future.
A proper professional programming job involved AGI-level understanding of human users and their needs as embedded in a social context. Plus the ability to create novel solutions. Plus the ability to use code as a collaborative medium to make a code base that is sustainable over the long term by their colleagues.
LLMs are not even 1% of the way to replacing professional developers.
To answer your complaint, I am absolutely looking at reality. Please point us to an actual real-world professional programming job, matching the criteria I list above, for which a current LLM can economically replace a human for over a 5-year period.
My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming that there are at least 269k fully automatable programmer jobs [1]. It shouldn't be hard for you to find at least one to start.
[1] Wikipedia says there are an estimated 26.9m professional programmers: https://en.wikipedia.org/wiki/Software_engineering_demograph...
If I built an entire car but I am missing the key. I cannot drive the car but the car is 99% complete. It's just missing the key. So what happens is, it can replace 0% of human locomotion but it's 99% complete? get it?
Just because the tool can't be used doesn't mean it's zero percent of the way to being drive able. Same with your job. The LLM can't replace any job yet, but it doesn't mean it's 1% of the way there.
But you see what I just explained to you is obvious. You already know this just like I and every one else on the face of the earth already knows this fact.
You're taking the discussion into a play on words. What does 1% apply to? Rather then use common sense and derive what I mean you prefer to redirect the conversation into a very specific definition of 1% that serves your own purpose.
We can play this game all day. We can discuss whether your application of 1% is more fitting than my application of 1%. What a waste of everyone's time.
I think it's better if rather then playing these games to "win" discussions, use your common sense to move the discussion past these games. Otherwise we're going to be talking about obvious things all day and getting all worked up about personal definitions.
Clearly the LLM is more than just 1%.
Mine is clearly measurable. You wanted to talk about job replacement; I'm measuring jobs replaced. You believe that we're at 48% complete (or 38% or whatever) but you can't point to even one programming job that ChatGPT can take over. If your fantasy 48% measure means 0% real-world success, I think the 0% is the more useful number to look at.
Humans may not be smart enough to build human-grade minds, in the same way that cats will never learn to program no matter how much they paw keyboards. People mistaking ChatGPT for AGI strikes me as the same kind of error where they mistook Eliza for a person, or 1970s/1980s computers as the things that would develop consciousness and maybe take over. It doesn't tell us much about the power of computers, but rather the inability of some humans to understand how complex and powerful minds are.
unfortunately the hard topics and the topics that matter are the questions that matter most are often qualitative.
Measurable answers are easy. Nobody cares about those things because it's obvious. Has ChatGPT replaced any jobs for programmers right at this moment? Overall no. But why argue a point everybody already knows?
Will chatgpt replace you in the future as the technology develops? Is it a precursor to the machine that will replace you? Qualitative trend-lines point to an unknown artificial entity that does replace all of us. It is useful to speculate and extend the trendline in that direction.
You may wish to shield yourself to reality and only look at things with quantitative numbers and lengths that are measurable by rulers but the shield is an illusion what you are doing is blindness.
Natural selection and natural history is built from qualitative understanding. Without subjective analysis on qualitative data we would not be able to understand natural selection from the macro perspective. You would be ignorant of the concept of evolution and natural selection and not much different than a creationist.
Obviously you aren't like that, you're just using a tactic to win a discussion. Always use hard data works when it works. But ultimately this strategy has failed.
And 38% is not "a third way to 100%". I'm sure you know what diminishing return is.
No?
I assume you actually don't know the concept of diminishing return, and it's totally fine. Every other comment here is explaining it to you already. I won't bother to repeat. Please read the sibling comments carefully.
You brought up this question of diminishing returns out of nowhere and offer no evidence for it. So why would my point not stand?
I read the sibling comments. One person brought it up but offered no proof that this is what's happening. Simply did what you did in a manner that was less rude, just stated the concept of diminishing returns applies to LLMs without offering a shred of evidence that indicates this is what is happening.
There is a difference between introducing the concept of diminishing returns out of nowhere, and providing evidence that diminishing returns is WHAT is actually HAPPENING with LLMS.
We're about a year out from the introduction of chatgpt, that's not enough time to know if all the gains in the past decade of AI has suddenly hit wall of diminishing returns.
It's a great assistant. I use it. It's not taking anybody's job.
It works but only edge cases are a problem. Next step is to fix the edge cases.
That said I'm not asking ChatGPT to write code that I use directly, I'm basically asking clarifying questions about the papers and other implementations I have on hand, and that's similar to how I use it in most cases.
For code generation I rely on co-pilot, which actually has the context of my codebase.
Edit: I will say this because someone else made me think of it: I find this tech all super useful on my like 20k loc and smaller side projects which are all self contained where as when I tried co pilot at work on a fragmented big corp code base I found it comparatively lackluster.
I work in data science and have to write a lot of repetitive stuff for parsing and cleaning data. It’s reduced my toil so dramatically in this respect, I can’t imagine going back to writing all that stuff again. I’m now much more ambitious in what I’ll experiment with as well because I know setting up the first stages of the data pipeline are going to be 10X less work than before.
At times, I need to correct the LLM on some usually obscure detail. It got the memory layout of the video screen on the BBC Micro completely mixed up with the ZX Spectrum, and I had to correct it about five times before it got it, and then it stuck for the rest of the conversation.
I've been on some wild goose chases, where I have fooled myself into thinking a particular approach would work and ChatGPT has been enthusiastically right there cheering me on.
It's like my dog going on an adventure with me, the dog doesn't care about where we're going or what we're doing, it's just excited to be part of the journey.
And in other cases, ChatGPT has been exceptionally useful in pointing out some blindingly obvious mistakes. It is like having a very knowledgeable, but exceptionally junior developer at your elbow.
If you ask the right questions, and don't except one shot questions to provide perfect answers, it works well for the most part.
The generated code will likely have a few small issues, but it still makes me way more productive. It allows me to work ~10 hours a week less and enjoy my life.
The most annoying part is it's impossible for it to say that it doesn't know something. Hard to train for, I know.
I think we'll have AI some day, but today it's just not at the level that all the hype claims it is.
Now the AI bros here are attempting to sell us their new snake-oil in the form of a stochastic parrot which promises to be the solution to everything.
Well, unsurprisingly it is another unexplainable AI black box which still requires the human to check every single output so that it doesn't hallucinate something incorrect, which it does almost all the time with a lack of transparent explainability as to why it hallucinated.
So the fact is, it is already overpromising and under-delivering for serious use-cases. Just like FSD (Fools Self Driving) was when that was over-hyped and with little to no Tesla Robo-taxis on the road.