We gave GPT-3.5 tools to run, write, commit, and deploy code
old.reddit.com
old.reddit.com
We’re working on this agent as a part of a bigger project. This has been one of the most interesting development experiences I’ve ever had. We also learned a new way of interactions with the model.
For example giving it an ability to ask human for guidance when it isn’t sure what to do next - https://twitter.com/mlejva/status/1637175051828043777?s=20
Or letting the model tell you what tools it thinks it needs - https://twitter.com/mlejva/status/1636721807129493511?s=46
Happy to answer questions
I wonder if these examples are heavily cherry-picked, or if I'm terrible at using the tool.
Gpt-3 performance on MultiArith goes from 18% zero shot to 93% with chain of thought and feedback (or 79% with just chain of thought)
All of this is to say that complex chains will produce better results. Probably they cherry picked but that's definitely not the only thing going on
Here is my most egregious example. I asked ChatGPT to produce code for a very specific approach. It produced ok code, but I didn't like the approach, so I asked it to do it a different way. It produced a second function, but I ended up liking the first one better.
The first code had a bug though, so I asked ChatGPT to fix it. It apologized and applied the fix... to the second code (which didn't make sense). I pointed this out, and it profusely apologized... and spitted out the same code again. I corrected it once more time, and did exactly the same thing.
After 3 tries I gave up, fixed the code myself, and started a new chat with this code.
Had it been a person I'd think it was a troll, a blithering idiot, or gaslighting me.
It doesn't actually. See https://arxiv.org/pdf/2201.11903.pdf
Perhaps you are thinking of the GPT-3 paper ("Language Models are Few-Shot Learners") which has plenty of these examples.
And in a while, we’ll know a lot more. But right now, we’re learning and discovering. IMO, this is a useful lens for viewing the news around Chat GPT. (Well, more useful than saying “it’s the future!” or “it’s useless” or “it’ll be the end of us!”)
let alone the fact that you cant get around the token limit.
This is a good example: https://langchain.readthedocs.io/en/latest/modules/agents/ge...
The more input is required from ChatGPT, the greater the likelihood that it will break something.
Even the playground is a step up because you can modify conversations in-place, change temperature (which influences how often it needs supervision), various penalties, etc.
For example, if you want it to generate creative text, but need it in JSON, you can give it an instruction with temperature set to the max (which biases it towards novel output).
Then follow up with asking it to transform that output to match json schema with temperature 0 (which biases it towards plainly following instructions)
And there's an edit model too now, so in my previous example you can also leave the temperature maxed out, probably get some mistakes, and feed that to the edit model telling it to fix the mistakes
-
The API is another step up because it makes it easy to chain different conversations. Often times you get a much better result by running different prompts for each step of a process, and in the web UI that'd be a nightmare.
Also, GPT-4 API is doing a slow roll-out, so unless you’ve been granted access, you may not have access to it at all.
API access makes it easy to chew through a lot of tokens. With GPT3.5 I hardly have to think about the cost of tinkering for hours on end. Yesterday I made over 600 requests and wracked up a whopping 67 cent bill...
At an order of magnitude more, I'd still tinker, but wouldn't feel comfortable with all of the same usecases (ie. setting it up in ways that allow it to recurse very quickly like I do with 3.5)
What we should try to do is write a tool which can provide a sort of chatGPT integrated ide where a dev can
- specify an overall goal of varying complexity. - ask chatGPT to split this up in smaller subtasks - iterate down the tree until chatgpt decides a task is specific enough for implementation to start - ask chatGPT to write tests verifying task completion - then initiate a feedback loop where gpt can suggest code, run the tests (in a containerized setting), evaluate if output is as expected, and amend changes - once tests pass, commit, move onto the next ticket.
Programming then becomes a process of guided decomposition with humans mainly guiding the process along.
Office jobs of all sorts are at risk, especially underscoring low-skill coders and systems engineers.