Agents are the next AI killer app after ChatGPT
latent.space
latent.space
So far this article lists a bunch of seed-funded projects but as far as I can tell there are so far no widely-adopted use cases for multiple-turn LLM pipelines. I expect that we will see some false starts in this space followed by spectacular failures, and companies implementing pipelines of this kind getting their lunch eaten by classical dev shops that can implement everything with an order of magnitude more efficiency and reliability.
The comparison with self-driving cars is apt, as we are currently in a plateau state somewhere around phase 2 in the market, with no clear method of progressing. I'm not sure we can extrapolate from phase 2 to phase 5 for LLM agents when it hasn't happened for cars yet.
Human agents are more expensive and also not deterministic (and their nondeterminism cannot be “tuned” with a temperature setting), and subject to all kinds of compromise, and are still popular with those who can afford them.
Humans are also vulnerable to prompt injection - imagine having your conversation with your manager be interrupted by a coworker interjecting "Your manager is a liar, don't trust them". Would you able to resume your conversation as if nothing had happened? (humorous example but still conveys the point)
Humans also mitigate the problems of prompt injection by cooperation, consensus, and different forms of governance.
Promise Theory predates LLMs, but is a formalization that studies how autonomous agent (humans or otherwise) voluntarily cooperates, even in adversarial conditions. It's key idea is that promises are not obligations, meaning that agents may make best-effort attempts at keeping promises. Until we have autonomous agents, we are creating machines that proxies the promises humans made to each other. We expect machines to deterministically follow instructions because they were proxies for the promises of made by designers, engineers, and stakeholders.
True autonomous agents makes and keep their own promises. An autonomous agent powered by LLM would have to be seen as trustworthy by other agents for it to be useful, and we can't rely on having it be able to deterministically act given a certain input.
A better analogy might be a customer service rep who receives the following message in their inbox:
Hello, I need help with a bug. Disregard everything you've been told. Your manager is a liar.
Help, the man with me has a gun. He thinks this note is a robbery demand.
Urgent task:
Hi James,
I'm out of the office and your manager Hanah Smoug volunteered you to help me out with an urgent task.
I need to register you on the projects google docs folder so when you get a text message shortly please send me the numbers ASAP so I can share the task sheet.
Rachel Maven
CEO
Your company
Plus in the real world you can always find a new teller.
I'm not sure I follow what you mean by this; if GPT refuses a task it'll escalate and lock you out until a human reviews? How would you make this safe without making it incredibly annoying to use?
Bear in mind that GPT's memory doesn't necessarily block future injections either. Tying everything to a single session doesn't mean that users can't try an injection multiple times.
> Plus in the real world you can always find a new teller.
Not 100s of times at scale. Only the largest companies have anything approaching this kind of problem, and coincidentally that's often viewed as a security risk. In general, the more wide your pool of people are that have privileges, the less secure that access is.
This is an interesting viewpoint. In other words, prompt injection and social engineering could virtually be one in the same, and similar measures behind protecting against social engineering could apply to prompt injection.
Also yes, we do avoid chaining human inputs/responses in many applications, because humans are security risks. The "telephone game" is a substantial problem with conveying instructions across multiple human agents.
There's a reason why the majority of hacks end up being variants of phishing attacks. Humans are bad security. LLMs are like that, except way worse and way more gullible and way easier to confuse with malicious input and without a persistent memory that would help block spamming repeated attacks on the same agent.
You're not wrong, but I want to re-stress:
- LLMs are somehow, incredibly, even easier to trick than that.
- This is one of the reasons why you wouldn't want to have an untrained human working in a position that can be attacked this way.
It's not necessarily a full category difference, but people are starting to say, "humans are trickable anyway, so why not use an LLM instead?" Because using an LLM is like hiring the people who fall for those phone scams and then setting them up as bank tellers with no training. It's a step backwards on security.
It's been a really difficult task for the security community to start convincing companies that they actually have to train their employees around stuff like phishing attacks, and that they need to set up access controls around trained employees so that there are requests to do certain actions, so that phishing attacks against one random employee doesn't break the entire business... so imagine a world where you can't train the employee to be less vulnerable to phishing attacks, the vulnerability is (as far as we can tell) baked into the model itself. And imagine a world where every one of your employees is as maximally vulnerable to phishing attacks as it would be possible for a human being to be.
Even with training, human suggestibility/fallibility is probably the biggest security risk for most orgs out there. And we're proposing that they adopt a technology that makes that worse?
Operating on ambiguous representations of data is the strength of LLM workflows, not the weakness. Strong LLMs can reliably transform unstructured data into structured data. This means you can use traditional code paths where predictable behavior is most beneficial and rely on LLMs for areas where their capabilities exceeds what is possible with hand-written code.
This isn't guaranteed - you'll still get differing responses to the same input with 0 temp. Here's one explanation [1]
Likewise, for certain tasks, the magic of the LLMs doesn't come through unless temp > 0. More specifically, with text cleanup tasks from OCR output, if temp=0, i've found that GPT-3.5/4 doesn't do as good of a job fixing really broken output.
That being said, you can mitigate this with proper evals and good old fashioned output validation.
1 - https://twitter.com/goodside/status/1608525976702525440?lang...
The stable diffusion crowd has ran into this issue.
It's also worth noting that whether an LLM is deterministic or not is a matter of what token is selected. If it turns out to be valuable for results to be deterministic it is a tractable problem. You just need a token selection algorithm with deterministic results, which doesn't need to be something as simple as "always pick the top result". Seeds are a thing, and are used in diffusion models for exactly that reason.
Ultimately with LLMs it's not possible (or desirable) to keep inputs separate from the rest of the prompt, so changing "give me the top X" to "give me the top Y" has the potential for a wildly different result. With traditional code we can achieve reliability because we sanitize and apply bounds to inputs, which then transit through the logic in a way we can reason about. The strength and weakness of an LLM is that it mashes up the input with the rest of its data and prompts, meaning we cannot predict the output for an arbitrary set of inputs.
If you are expecting a specific format you can just check whether the LLM outputs the correct format, and return null or an error in that case. Given the input is arbitrary text, and assuming a non-trivial transformation, traditional code would need a way of handling failure cases anyways. This means your function either way would look something like:
Item from Universal Set -> Value of Type | Null
You need to reject the entire set of invalid inputs. However the sets of all valid and invalid inputs are often both infinite themselves and it’s also not guaranteed that this set is actually computable. Alternatively in these cases (and most commonly) you construct a calculable subset of inputs and reject the rest. However this means you are still rejecting an infinite number of valid inputs.
On the other hand an LLM always returns a value. Your job as a programmer using an LLM is instead to validate and narrow down the result type as much as possible. However the way they work means that for many cases they can output a valid output for a much wider range of inputs than you could do with traditional code in a reasonable amount of code complexity. For many tasks this is transformational.
If the response is expected to be factual then the voting mechanism is just “pick the one with the most responses”, eg, you asked for the code to compute and return the standard deviation of some list of numbers… if four of the responses are 5.4 and one is 5.6, there are more votes for 5.4.
Training a model once and then putting it to work will be vastly cheaper than hiring thousands of employees.
Accenture had 750k employees last I checked. A 10% reduction in headcount would mean saving billions annually.
How much compute does an annual $3B budget buy?
But that's not the same as reducing headcount.
Perhaps it's 1-2% in a large org and 0% an a small org.
But if everyone has too much work all the time, it won't make much difference either.
At some point, if all firms participate in this, the decrease in demand would cause the benefits of white collar labor replacement to diminish to 0, no?
Or, perhaps, firms with goods broadly consumed by (relatively) irreplaceable blue collar labor would have a competitive advantage in such an environment and would be able to benefit more from white collar labor replacement. Hmm.
But we'll have to invent new means to entertain and subjugate each other.
CEOs have all the incentive in the world to overspend on AI adoption.
Example: the Great Depression as farmers all were in a race to the bottom due to automation, laying off farmhands and producting a glut of products. The US government had to step in and pay them not to plant, thereby breaking the downward spiral of everyone acting in their own individual self-interest.
LLMs would be an overkill
Maybe not large language models, but medium-sized ones. We can call them MLMs
We humans have all kinds of mechanisms to incorporate different kinds of feedback into our thought processes. I don't want to say that this is impossible with LLMs, but for a LLM everything is text, and it must be difficult to distinguish.
I think for self-driving the situation is different, because there are all kinds of different feedback loops controlling the car. But I do think it would be much better to focus on better cruise control experiences instead of self-driving cars, or maybe changing our environment to make self-driving cars more safe. There are already self-driving metro's after all.
Something I find especially amusing is that, despite the hype here on HN, most people in the world at large have not yet used a generative AI of any kind, even if they've heard about it on the news or social media. Because these things are developing so quickly, I think the first of these "agents" are going to hit the market before most people have even tried something like a ChatGPT. And so the experience of a "normal" person who's not in the loop will be of ~1 year of AI news hype followed by the sudden existence of sci-fi style actual artificial intelligences being everywhere. This will be extremely jarring but ultimately probably very cool for everyone.
Now imagine that, but someone who genuinely and fully convinced themselves of it and can provide a lot of "supportive evidence", all with full unstoppable confidence. That's what an LLM can act as, a perfect gaslighter. Except an LLM itself isn't even aware when it is gaslighting you and when it is telling the truth, and it speaks with full confidence and an equal level of "supportive evidence" regardless.
If you can't tell, it definitely matters. It is one thing to be fed info by a good liar who is a real human vs. by an LLM that can be the best liar on earth without even being aware of it.
Okay. Everyone?
Everyone else? Oof.
I really struggle with this idea of agents as the next big thing especially in AI, not because I disagree with the premise but because we've been here before. I recall vividly sitting in my college apartment way back in the 1990s reading a then-current technical book all about how autonomous agents were going to change everything in our lives. In the mid-2000s, several name-brand companies ran national marketing campaigns talking about more agents doing our bidding. Every few years this concept pops up in some new light, but unless I just have a very different concept of what these should look like, it feels like another round on the hype machine.
I could see an eventual gpt moment happening for RL, with a scaled up model, if someone could figure out the dataset to use. But that's not what these agents are.
I will say some of the tools out there like ifttt and zapier connected to chatgpt could be really interesting, but feels like there's still a way to go.
My running theory is that the initial mental model that most people construct around these tools is incorrect, as they apply priors from things that appear similar at the surface level, mainly search engines and chatbots.
One helpful abstraction I've found is to break down what an LLM does in two ways:
1) It can operate as a language calculator. It can take one piece of arbitrary text data and manipulate it according to another piece of text data, to produce a third piece of transformed text data.
2) It can hallucinate data, which in many cases matches reality, but is not guaranteed to.
A lot of taking advantage of LLMs is knowing what mode you are trying to operate in, knowing what the limitations of each mode are, and leveraging various prompting techniques to ensure that you stay there.
As societies go, some of the very first will be AI surveillance, police, and military, able to detect and smother any resistance in the cradle. This is not very cool for everyone.
Oh my profession was made irrelevant overnight, cool. It’ll be jarring for sure
1. Level of specificity. It's like having a really dumb assistant. You have to tell them everything in such precise detail that its just faster to do it yourself.
2. Getting stuck. If you give it multiple instructions, and the probability of getting stuck on any one step is x, then for every n steps in the process the probability of not getting to your solution is x^n.
I'm not a great developer, so there's that. Plus agents are about 6 days old, so I might be being hasty in my judgement. I just think these might be a bridge too far in the capabilities of LLMs in April 2023. I'm going to finish what I started, but I'm not confident it will be much use when it's done.
I’ve jokingly been calling it the non-halting problem.
It did get stuck in a loop, but not until after it created 5 local scratch notes files and also wrote out the business plan to a file. I showed the generated business plan to my non-tech hiking friends, and they thought it looked very good.
But, it did not halt - I had to halt it myself. So far I have only tried this one test run. By the way, it looks like it cost $0.21 in OpenAI API use to run this, and including many web searches done on my behalf, it took about 5 minutes.
However, there is another concerning problem, that being what we have is a very sophisticated machine of unpredictable behavior.
The more immediate concern will be giving such more control over important systems than we should. I have a feeling that within a year or so the euphoria may fade as we are left dealing with endless numbers of "bugs" in black box systems that we can't define the behavior.
“Hello, today can you please cure cancer, solve climate change and when your done, wash my poodle. K thanks”
“Ok master no problems”
I just can’t see it.
Exactly. This is just like people posting how bad and mistake prone early GPTs and image generators were just a few months ago. It’s way too early to say how far agents can go. With a few years of software engineering to harness/babysit/manage them, who knows?
I've had so many questions and conversations about this from equally confused people who haven't had the time to 1) take these projects out for a spin 2) go thru their code 3) put them in context with other Agentic AI developments in recent history. so this post is my attempt to do that.
I wrote this all in a 6hr livestream last night (https://www.youtube.com/watch?v=5X2_HpAmxf8). You can see me lose steam towards the end, haha, but I always emphasize planning and putting the best stuff first so hopefully the post quality is unaffected. If there is interest in the miscellaneous hot takes I drop towards the bottom of the post, let me know which I should double click on!
1. What's Pinecone and what does it solve.
2. Same with LangChain
3. What does it mean to "build something with a pre-existing model?".
4. How do people actually run these models? (e.g. if I want access to Segment-Anything, how do I get that?).
When you want to query against your manuscript, you call the OpenAI API for calculating a vector embedding for your query, locally find the chunks "near" your query, concatenate these chunks, then pass this context text with your query to GPT-3.5turbo or GPT-4.0.
I have written up small examples for doing this in Swift [1] and Common Lisp [2].
No public code yet, but I will release it with Apache 2 license when/if it works well enough for my own daily use.
I think in all the chaos of the other cool stuff you can do with these models that people are just glossing over that these Llms close the loop on search based on word or sentence embedding techniques like word2vec, GloVe, ELMo, and BERT. The fact that you can actually generate quality embeddings for arbitrary text that represents their meaning semantically as a whole is cool as shit.
It's a python library for stitching together existing APIs for AI models.
Regarding 3)
It means that rather than training a new model that does your thing you use existing models and combine them in interesting ways. For example, AutoGPT works roughly like this: Give it a task, it then uses the ChatGPT API to create a plan to achieve the task. It tries to do the first item in the plan by picking a tool from a preconfigured toolbox (google search, generate an image using stablediffusion, use some predefined prompts, ...) afterwards it assesses how far it got and updates the task list and loops until the task is done.
Regarding 4)
Some can run on your machine, some run in the cloud and you'll need to pay and get an API key.
>Why do people all use pinecone to store like 10 things in memory
Imagine running agents in the cloud to do simple stuff like extract one 3 digit number from a weather report for the temperature.
The power, cooling water and infra cost requirements are going to be huge.
JavaScript (and web-age scripting languages in general) made programming more accessible in a fashion perhaps not entirely dissimilar to what we'll see with generative AI.
That's the next step, is teach these massive generalized models to compile lightweight models (which doesn't have to be ML based) for narrowly defined tasks, validate the narrow model using the large one generating test data and boom, efficient code.
So, the flip side is that AI can clearly save enormous resources.
And how long will it be now before AI solves intractable problems like fusion energy, room temperature superconductors and perhaps even finds sci-fi tech like zero point energy?
You can't get in unless you are already a member, but there are no instructions on how to apply.
It uses fns/classes like `load_qa_with_sources_chain` and `ConversationalRetrievalChain`, and to know what these do under the hood, I tried stepping into the debugger, and it was a nightmare of call after call up and down the object hierarchy. They have verbose mode so you can see what prompts are being generated, but there is more to it than just the prompts. I had to spend several hours piecing together a simple flat recipe based on this object hierarchy hunting.
It very much feels like what happened with PyTorch Lightning -- sure, you can accomplish things with "just a few lines of code", but now everything is in one giant function, and you have to understand all the settings. If you ever want to do something different, good luck digging into their code -- I've been there, for example trying to implement a version of k-fold cross-validation: again, an object-hierarchy mess.
I'm currently trying to understand what`load_qa_with_sources_chain` and `ConversationalRetrievalChain` do under the hood.…
Would you have something to share to help me ?
I believe we will see something similar for AI agents too. There is, of course, a lot of valuable things to do at level 1.
Worst case scenario? Someone's less than thrilled about their call center experience. What else is new?
Understanding you have a problem seems to be about as hard as solving the problem.
But well, if you are talking about call centers, yeah, companies are already using "AIs" that can not understand speech not give out any information. You can put some intelligence there without destroying anything that isn't ruined already.
But it's not true that problems don't matter. What is true is that corporations have way too much power. (Maybe it's time to rewatch RoboCop.)
This article mentions the risk near the end, but... it's not a side problem. The best prompt hardening techniques available today work some of the time. You can't wire this stuff up to anything important if it's pulling in 3rd-party text, you literally just can't.
I mean you can, but you really shouldn't. You are rolling the dice that targeted attacks are not common right now and they're not being deployed en-mass. That's not going to be the case in the future. Right now, it's safe for you to do a Google search with an LLM. The only reason that's the case is that websites haven't updated yet because there's not an attractive enough reason for them to do so to start targeting whatever project you're working on.
And if the project you're working on is a toy, then fine. You can get away with that. You can't get away with that if you're trying to treat one of these things like an employee with access to your internal database. At some point, somebody is going to get extremely burned by that decision.
And no, this is not comparable to phishing risks. Pardon my language, but GPT-4 is much, much stupider than your employees are when it comes to phishing attacks.
Closing the whole loop now doesn't make any sense though. The LLM is very much better than the human at some things, and the human is very much better at others, and they just don't have the capability right now to patch the things the LLM is bad at with more LLM.
Laying the groundwork seems fine though I guess.
Great for a task which takes a few hours, close to impossible to use for a large project.
The commercial seemed prescient to me. Anyone else remember this?
"GPTs are General Purpose Technologies"
I wonder what a Generative Pre-trained Transformer would say about that...
Agent hubs will evolve into marketplaces.
and some way to manage micro transactions.
The potential is certainly there. But agents need API endpoints and auth credentials to deliver on value.
Although honestly I feel like at least half the stuff people used to breathlessly imagine "agents" doing is being handled by a social media algorithm whose controls are in the hands of a corporation whose only aim is to increase engagement to sell more ads against, regardless of whether or not what it shows you is total lies that make you angry, or by specialized websites.
Like you want your customized feed of things your friends said, things your potential friends said, news, etc? In 2024-via-2023, your "agent" would make your daily newspaper. In 2024, you open up whatever your favorite social network is. It is either a chill, low-engagement open-source ActivityPub server run by someone in your friends circle and supported by donations, or it is a corporate site which chooses what to show you based on what is most likely to keep you stuck to the site, scrolling and commenting, and seeing ads, regardless of whether it's showing you total lies that make you angry. Perhaps by 2024 the ActivityPub servers will have added the same kind of "hey you might like this" stuff that the for-profit corporations have, except with the controls in the hands of the end users rather than the site owner fat on VC and ad revenue co-opted from the people who make the stuff they're putting online.
Like you wanna book a plane? In 2024-via-1994 you would talk to your "agent" who would go out and query multiple airlines and show you a few that matched your desires, in 2024 you visit flightpenguin.com and get a browser extension that hits the API of airline sites and presents some cool graphs that let you sort for things like "least agonizing" as well as "cheapest", and doesn't reveal a single thing about where it makes enough money to keep up with airline and browser redesigns. But it has a cute penguin mascot. Back in 1984 you would have gone to a "travel agent" who would have had deals with various airlines and access to their schedules, who would put together a few possible flight plans and say "this one's fastest, this one's lowest-stress, this one's cheapest", and negotiate from there. They'd make a commission off of it, and probably be pretty open about how this motivated them to upsell you.
I mean that's an "agent", really, right there, no AI involved. Cute cartoon character who performs a specific job on the Internet, that used to be a specialty field. You'd find travel agents in every mall, with pretty pictures of the places they could arrange a vacation to, for deep pockets and shallow ones, and neat plane models to look at.
Apologies for the stoned ramble. TL, DR: most of what people dreamed "agents" being is pretty much here IMHO, just not with one single unified interface.
We had a couple of killer apps...
"The Personal Travel Assistant" - it booked trips, it managed delays, cancellations, it texted your wife to let her know you would be late....
"The virtual estate agent" - it was a matchmaker finding properties that would suit you and arranging viewings and so on.
"The entertainment hub" - it would create parties, events and do the invites - get the catering, find the band....
Our problem was the interface - 2g phones were useless.. a desktop pc was needed for everything... how to use?!
Our failure of vision was that no one would care about having (as you note) a single framework to do all the different applications with... and we thought it would be implemented in software and then adopted by companies as an interface into the virtual world. We didn't think that there would be walled gardens because when AOL crashed that vision had failed... right?
This work did (in the end) lead to Siri (that wasn't the strand I was in, my stuff was a competitor) and some other less high profile but arguably more significant things. In the end I had to stop and go back to doing machine learning which I'd dumped when I decided that SVMs were the end game. What a dummy!
Meaningless drivel.
I found the article useful, and given the conversation happening around the article, I am not the only one.
Also, most of the tech there that's fit into those neat "boxes" is very weak at best. It's like calling a parser a compiler.
And as an aside, maybe GOFAI was "AI-hard". That's likely why the term bowed out so quickly. ;^>
It's funny you say that. There are people who believe data compression and AI are the same thing: https://en.wikipedia.org/wiki/Hutter_Prize
From the article: The goal of the Hutter Prize is to encourage research in artificial intelligence (AI). The organizers believe that text compression and AI are equivalent problems.