I would say it improves my productivity by maybe 5%, which is an incredible achievement. I’m already getting to where coding without it feels very tedious.
I would say it improves my productivity by maybe 5%, which is an incredible achievement. I’m already getting to where coding without it feels very tedious.
Recently at work, for example, I've been setting up a bunch of stuff with some new technologies and libraries that I'd never really used before. Without ChatGPT I'd have spent hours if not days poring through tedious documentation and outdated tutorials while trying to hack something together in an agonising process of trial and error. But ChatGPT gave me a fantastic proof-of-concept app that has everything I needed to get started. It's been enormously helpful and I'm convinced it saved me days of work. This technology is miraculous.
As for my job security... well, I think I'm safe for now; ChatGPT sped me up in this instance but the generated app still needs a skilled programmer to edit it, test it and deploy it.
On the other hand I am slightly concerned that ChatGPT will destroy my side income from selling programming courses... so if you're a Rails developer who wants to learn Elixir and Phoenix, please check out my course Phoenix on Rails before we're both replaced by robots: PhoenixOnRails.com
(Sorry for the self promotion but the code ELIXIRFORUM will give a $10 discount.)
Better to ask it for a bunch of small things and piece them together
I’ve found it to be very forgetful and have to work function-by-function, giving it the current code as part of the next prompt. Otherwise it randomly changes class names, invents new bits that weren’t there before or forgets entire chunks of functionality.
It’s a good discipline as I have to work out exactly what I want to achieve first and then build it up piece by piece. A great way to learn a new framework or language.
It also sometimes picks convoluted ways of doing things, so regularly asking whether there’s a simpler way of doing things can be useful.
GPT3.5 is 4k tokens and has a 16k version GP4 is 8k and has a 32k version.
You are correct that this needs to account for both input and output. I suspect that when you feed chat gpt longer it prompts, it may try to use the 16k / 32k models when it makes sense.
For features that probably should exist but don't it does a really good job of sending you on a wild goose chase.
GPT-4 reduces hallucinations by at least an order of magnitude, and hasn't failed me yet.
In that case they become complications.
You really need to be quite competent in the thing you're asking it to do in order to ferret out the hallucinations, which greatly diminishes the potency of GPT in the hands of someone who has no knowledge of the relevant language/runtime/problem domain/etc.
She asked gpt to help get an html version since apparently she got stuck with the wysiwg editor.
However gpt gave back a full html structure, including head and body. Pasting that into listmonk breaks entire webpage. Then she freaked out and told me listmonk sucks :)
But no, you're fundamentally right. It just goes to the question of whether an LLM assistant can in any sense replace or displace human programmers, or save time for human programmers. The answer seems to be somewhat, and in certain cases, but not much else.
If I already know the technology I'm querying GPT about, I'm going to spend at least some time identifying its hallucinations or realising that it introduced some. I might have been better off just doing it myself. If I don't know the technology I'm querying GPT about, I'm going to be impacted by its hallucinations but will also have to spend time figuring out what the hallucinations are and why this unfamiliar code sample doesn't work.
1) It could use the JSONformer idea [0] where we have a model of the language which determines what are the valid next tokens; we only ask it to supply a token when the language model gives us a choice, and when considering possible next tokens, we immediately ignore any which are invalid given the model. This could go beyond mere syntax to actually considering the APIs/etc which exist, so if the LLM has already generated tokens "import java.util.", then it could only generate a completion which was a public class (or subpackage) of "java.util.". Maybe something like language servers could help here.
2) Every output it generates, automatically compile and test it before showing it to the user. If compile/test fails, give it a chance to fix its mistake. If it gets stuck in a loop, or isn't getting anywhere after several attempts, fall back to next most likely output, and repeat. If after a while we still aren't getting anywhere, it can show the user its attempts (in case they give the user any idea).
It should suggest, lint the suggestion in the background, and if it passes offer the suggestion and if not provide the linting issues output to rework the suggestion.
In general, token costs going down will in turn increase the number of multi-pass generation systems over single-pass systems, which is going to improve dramatically.
Combine all that with persistent memory storages that can provide in-context additional guidance around better working with your codebase and you, and it's going to be quite a different experience than it is today.
And at the current rate of advancement, that's maybe going to be how things will look within a year or two.
This makes a big difference, I'm making code writing stuff at the moment.
Injecting results from a language server while it's generating would be huge imo - same as giving humans autocomplete & hints.
Basically, these days before I dig into documentation I ask "How do I do X with Y framework in Language Z" and if it's pre-2021 tech it works amazingly well.
Even in the IDE I'll sometimes just write comments like (arbitrary example out of thin air):
// Q: Should we use a for loop or a while loop here? // A:
It doesn't always have a great answer, but as you say, it almost always helps my own thinking about it, which is often much more valuable.
Its hard to say if it improves my productivity because I just wouldn’t have done those things
But for the overall applications I think its improved a lot because we can implement best practices more consistently and catch regressions due to the aforementioned unit tests and documentation
But even then, it's not 'replacing' you.
It's just going to let you spend less time on BS and more time on the things that are your maximal value contributions to a project.
If I had a dozen junior or mid level devs you could hand work off to, would that save you time? Would you kick back and not review what they were doing, particularly around business critical parts of the software?
The conversation around AI has become obscenely binary, pulling from (now obsolete) SciFi influences to cast it as humans vs machines.
But it's a false dichotomy. Collaborative efforts are almost certainly where this is going, and 100% human or 100% AI will both be significantly inferior to a mix of both.
The question is if generative AI is powerful enough to reduce the number of programmers needed to achieve a task, without creating enough opportunities to replace those programmers.
Before we are all replaced there could be a moment where demand for software engineers is 10x less.
For industrial applications in particular they need to be functional and operable, not shiny.
The market may adjust over the longer term, or it may just continue to be volatile as the rate of change accelerates. In that case, we can't fix the work market, and we instead have to address the need for people to feed themselves another way.
I've started developing in a new language and I can hardly do any work without the LLM assistance, the friction is just too high. Even when auto-competitions are completely wrong they still get the ball rolling, it's so much easier to fix the nicely formatted code than to write from scratch. In my case the improvement is vast, a difference from slacking off and actually being productive.
Theyve been debugged.
Review and testing.
Reviewing is easier when there is less code (i.e. libraries are in use)
These were tricky problems that were small scope - I've picked them so I could easily provide it to GPT for review.
So I doubt larger context window will do much.
Notably, it's not GPT 3.5, it's 3.0, which is pretty stupid as far as the state of the art goes.
The upcoming Copilot X will be based on GPT 4, which has "sparks of AGI".
In my experience there is no comparison. GPT 3 is barely good enough for some trivial tab-complete tasks. GPT 4 can do quite complex tasks like generating documentation, useful tests, finding obscure bugs, etc...
Sadly, you’ve just described the majority of the developers I’ve had to work with recently.
Most have no agency, write boilerplate code with no creativity, need their hand held every step of they way, and won’t do anything they’re not explicitly ordered to do.
You probably work in an SV startup with a highly skilled workforce. Out there in the real world there are armies of low-skill H1Bs and outsourcers that will soon be replaced with automation.
It’s a recurring theme in economics. Outsource to low cost labour, insource with automation, repeat.
We need an AI that iteratively tweaks its own architecture (to recreate and surpass those modules which are necessary for human thought), and maps out hardware enhancements* to accommodate the new architecture.
*I seem to remember Google working on ML software that proposes new chip designs a few years ago