GPT-4 Can Use Tools Now–That’s a Big Deal
every.to
every.to
The weird thing is it keeps telling us to place some weird plastic molding called SEMTEX hooked up to a Starlink receiver in between the turbine insulation blankets when we put the engine back in. None of us can figure out why there’s a high voltage wire going from the radio to the plastic - it makes no sense. But hey who are we to argue with AI?
That's hilarious but not really the use case we are talking about. It's more like automatically writing a query in a DSL to cross-reference testing methods and display the output of that specialized database.
It said they were for diagnostic and monitoring purposes only.
Sadly it only works with gpt-4. Gpt-3.5-turbo immediately starts trying to ping all the radios one by one.
It's a matter of making more efficient hardware and more efficient paradigms for the specific LLM application. The history of computing is engineers doing this over and over again. AMD just released an AI chip that can handle an application which normally requires multiple accelerator boards from Nvidia. And this sort of thing is just scratching the surface with relative minor architectural differences. Radical new approaches with memristors or other compute-in-memory departures are coming in less than a decade.
There is no possible way you can prevent people from connecting AI to the internet or tools.
So we are very close to the point where the people cannot compete with a cluster of agents working on a task. People lose control because in order to compete, they have to give the AI broader and broader goals. Since the AI will be working dozens or hundreds of times faster than humans, pausing for human input means the competition will race ahead.
The real issue impending issue is hyperspeed AI that is above average human intelligence. In some measurements this is already here. Within a few years the models or collections of specialized agents will be so good and so fast that no one will be able to deny it. And anywhere a human is in the loop routinely will fall behind due to the limitations of the human.
This does not require runaway AI or AI that is alive. Its just the consequence of deploying faster and faster hardware and gradually improving the model training. It will be quite obvious within 2-3 years that we need to consider _not_ making the AI hardware any more efficient than what we have at that time.
- Helping design next-gen hardware on which they'll run, or write its firmware;
- Procuring more compute, or making money for or otherwise improving ability of a company/group to procure more compute resources;
- Improving their own algorithms;
- Improving anything upstream in the dependency chain of their own algorithms, for example fundamental mathematical libraries;
- Directing or convincing people to increase the number of AIs in the AI cluster;
- Making profit in general, or otherwise incentivizing humans to work on next-gen AIs;
... then they're already capable of runaway self-improvement.
The term "self-improving" carries connotations of an AI working on it own sources at runtime, as if it was self-evolving Lisp image. However, being a compute artifact and therefore trivial to copy and modify, the AI needs to be treated as a category. With that in mind, GPT-4 helping people to improve code that leads to LLaMA ${whatever-B} that then helps create next gen of Stable Diffusion, which prompts more resources to be put into merging the next gen SD with upcoming GPT-6 into a multimodal model, which then helps make 3D integrated circuits viable, which allows to ... - this is also the case of AI self-improving.
While these misguided cinematic fantasies captivate everyone's attention, automated decision making is already deciding whether refugees are settled or sent back to be killed in their originating countries, whether people are given housing or not, or "deserve" welfare. The emerging research illustrates quite clearly the problematic biases entrenched in these models. Why this doesn't dominate AI discourse instead of hallucinatory speculating about some theoretical future is beyond me.
You are confusing LLMs with Markov chains.
ChatGPT can think and reason but it doesn't have a high bandwidth sensory experience or subjective stream of consciousness. It doesn't think in the same way as humans. Etc. Etc.
But thinking is still the best way to describe what it's doing. It has a model of the world and what you are asking it to do. It absolutely has to synthesize information to be able to fulfill requests.
And automated decision making is exactly what I am warning about. As the speed, effectiveness and performance of the models increases, it becomes uneconomic and uncompetitive to keep humans in the loop. The most obvious real danger is in warfare. Racial bias is one aspect of the problem but overall it's more than that. It's that the efficiency pushes towards fewer and fewer check-ins with humans so you are relying more and more on the AI training. And so any type of defect at all or misalignment with your goals can be very problematic for you.
I think that's going to be a very near inflection point where we'll see exponential improvements with little human interaction. I'm not sure if that prospect is more exciting or scary.
My main thought on this stuff that I have been repeating is that the only real mitigation is to prohibit the manufacture of certain levels of AI hardware. Again, hard to enforce, but you have to try. And not having the hardware exist will be more effective than laws.
If we make an effort then we might be able to hold off the post-human era for a few generations. By post-human I don't mean that (necessarily) normal humans have been killed off, just that they don't control the planet anymore.
To the extent OSS community uses GPT-3 and GPT-4 to generate training data for training open models, AI is indeed being used to improve itself.
Runaway self-improving AI doesn't need to be - and is, initially, unlikely to be - a single instance of a model improving itself at runtime. AI models being used to directly or indirectly increase capabilities of other AI models fits the bill too.
It's a one-time deal, but ten GP was askind:
> Has an AI already been used to improve itself, or create another AI?
And IMO this fits the bill.
> remixing existing ideas
Those two are the same thing. That's how we do it too - we come up with new ideas through remixing ideas we've learned.
In December, it was hard to imagine what a dumb AI could do and we were trying hard to make it smarter and more useful.
Today, I can see the concern. Not that it is intelligent and thinking, but that a runaway command could cause the AI to make something smarter. (Although, I still can barely use any GPT4 code, so I'm not going to panic)
It's hit or miss for me, but I feel mostly because I don't have much time/opportunity to try and truly explore its capabilities. A single-shot GPT-4 ain't gonna do anything nefarious, but it isn't clear to me that a cleverly constructed ensemble of 10 or 20 or 50 of "agents" running in parallel, each with a different role and all talking amongst themselves, isn't going to add up to something potentially smart and powerful.
Sure, right now such setup would burn through your wallet real fast. But it's not that expensive - especially not for companies. The real limit won't be the money, but... rate limiting OpenAI and Microsoft impose on their model deployments. But that rate limiting is negotiable, especially for a large customer - and then Open Source community is working hard to create open models of similar power, thus reducing the costs down to raw compute bill.
1. You can input and/or output structured data from ChatGPT.
2. ChatGPT can select from a list of functions and their schema and use one (with "use" being telling the user to use said function and providing input parameters), or use none at all if not applicable.
#1 is the more interesting use case I suspect people will use and in my testing it's incredibly effective and stable. #2 is an extention of modern tooling, but nothing new and it's up in the air whether it's more robust; the demos OpenAI have aren't great in that aspect.
This article brushes it off but the actual user implementation of function selection is a pain in the ass (and also passing structured input data) that can't easily be explained in a tutorial, and like using LangChain it increases the information asymmetry for people getting into modern AI.
The actual process of having ChatGPT select from a discrete set of outputs is very simple in practice and it's incredibly annoying that people consider it to the the "killer app" of LangChain. I might have to accelerate my blog post on how to do it efficiently, but in practice it's just a simple, small prompt: https://github.com/minimaxir/simpleaichat/blob/main/PROMPTS....
Stupid question, but... can't you just feed the relevant bit of API docs to GPT-4 and kindly ask it to write you the user-side call dispatcher, complete with error handling? :).
Perhaps combined with: "here is a bunch of function declarations: {taken from OpenAI example}; here is a corresponding JSON schema {also taken from OpenAI example}; based on this example, please write me a JSON schema for the following function declarations {actual functions you want to provide}."
That is, the problem doesn't rely on anything that's less than 5-8 years old (or more, I don't remember how old the JSON as we know it is).
I told it that it hallucinated the function, and it apologized and admitted the function didn’t exist.
Then it gave me a new solution that used the same hallucinated function, but wrapped it in another hallucinated function.
Tool selection, task planning, and termination condition determination in one single LLM call.
https://nextword.dev/blog/gpt-function-calls-are-task-planne...
They can produce code snippets that may even work, but they can't stay focused to make it all fit nicely into a consistent architecture. If you ask it to write you some larger and reasonably complex codebase, you'll end up with a weird junior-grade patchwork.
And even if you talk architecture first, they behave as if they have super-short attention spans, so they mostly disregard it and get carried away. (Of course that's not what really happens, I'm kind of anthropomorphising it, but that how it feels like).
Supposedly copilot x will have a deeper IDE integration and voice commands.
It helped me to bootstrap the project and pointed out at some APIs I needed, but then it was a constant fight for code quality. I wanted proper actor abstraction (because I was dealing with Bluetooth, even if I don’t know the language I can see the patterns) and it just loved to write code in style of some cheapest-grade “how to use CoreBluetooth in Swift for dummies” blog articles.
Vanilla ChatGPT uses a GPT-3.5turbo version with a context window of 4096 tokens. GPT-4 has twice the context window by default, and has a variant with a context window 8x as big.
It's not "8x better" with regards to forgetting what it's doing, but it's substantially better.
But we don't need to make much improvement in tooling and economics for it to start to look pretty good.
This just makes it easier to build systems like that or the system to be more efficient. For example, you can give it a function to call to replace some text in a file. Whereas before you needed to come up with some custom scheme with a diff format or sed call or something which was prone to breaking due to things like minor AI "typos" or newlines etc.
This also makes it easy to implement an interface to things like a compiler or other command line tools so the AI can call them itself if necessary. Such as for installing packages or initializing a project.
Before you had to keep begging GPT to follow some protocol with file outputs or something. Although GPT-4 was much better (but cost 10X more).
Only corporate entities can access more powerful GPT-4 functionality, either through Microsoft Azure or OpenAI directly. This goes through an entirely opaque application process.
"Open" AI?
These incredibly powerful tools are already entirely co-opted and opaquely gatekept by the standard bearers of capitalism, who control access with no oversight.
Terrifying.
https://platform.openai.com/docs/guides/gpt/function-calling
I stand by my statements.
But yes, you need API access.
Then I applied for GPT-4 API, twice, and never got it.
Apparently Chemistry and industrial engineering were not good enough purposes.
Keep the ClosedAI jokes coming folks, it wont matter, but I surely like knowing I'm not alone.