Auto-GPT Unmasked: The Hype and Hard Truths of Its Production Pitfalls
jina.ai
jina.ai
Why should it even work? Cool, now there's thousands of people playing with the same code, glad I'm not getting the github issues or trying to manage the PRs, but lots of smart people that want to play around are playing with the same code and forking it and suggesting things and we have def not reached a point where we need to have a 3 month argument about formatting guidelines, so how can it be a bad thing that a bunch of people are playing with some code and a new tool?
Did we really go from "oh my god deep mind can make things have weird eyeballs" to "WTF this hyper-realistic photograph I created with a paragraph has some shadows that aren't quite in the right place and autogpt gets stuck in loops and thousands of API calls hitting some of the most computationally expensive things humans have ever done cost some money and sometimes go into loops in this very mature codebase?"
humans man....
Why is it crucial? This project has existed for barely a month, and describes itself as an experiment. Who is using this for production?
Or maybe, just maybe, all variables will continue trending in a positive direction in the time it would take to build a production product. And then I wonder what point, exactly, this article is trying to make.
Not how technology works in my experience. My mental model of technology is that it acts as a multiplier on life's peaks and valleys, so the highs get higher and the lows get lower. To my mind that means AGI will bring the ultimate of highs along with the ultimate of lows. Doesn't sound like utopia to me. Sounds like a good time, until it's not.
tho dont see anyone building paywall yet
We have people saying that unfinished toy projects on GitHub are going to revolutionize entire industries on the basis of one contrived demo. We have other people calling for air strikes on data centers and nuclear war to stop... text generators.
What's wrong with people?
https://arktos.com/2023/04/12/chatgpts-diplomatic-gambit-mir...
Boredom?
But yea, I’m seriously concerned about the hype and the stupid things people immediately try to do with it. About others safety and intentions.
“I’ll just hook it up to my credit card and put it on the Internet…”
The internet was a hype wave too. Remember the dotcom bubble? And then also remember how eventually the internet actually did revolutionize communication, business and entertainment?
Yes, there's hype. There's also a tangible sense that we're at this 'internet circa 1995' moment. I feel like the hype is warranted.
There's far more here than cryptocurrency.
This is not the nothingburger of crypto
We literally watched crypto/web3 do the entire speed run from weird nerd project to spam/ponzi/get rich quick/celebrity/vaporware shell game in a couple years. Like All The Way to stadiums being bought and international manhunts and average people losing lots of money with something they didn't understand to most people being annoyed by bots and influencers to virtually 0% of projects having any other purpose but to generate hype
(with actual ongoing interesting tech being built underneath it)
the fact that autogpt in One Week went from AI experiment to "why doesn't this ship money to my bank account or what is github I'm here for the free money" based on youtube videos and spam articles etc really does note bode well for the hype phase of AI
As to the tool itself, I played around with it and it has some cool ideas that I think are valuable, even if it's not AGI
I think this is due to the main loop being a GPT feedback loop. Each loop it has a small chance of something going wrong (or a large chance, depending on the query), so as it loops repeatedly, the chance of failure approaches 100%.
My idea was to replace the core loop, instead of a GPT feedback loop just make it a few lines of Python.
Now the thing actually does what it says it's going to do, "thinking lag" is eliminated, and API usage is reduced 80%.
I turned (parts of) Auto-GPT into a tiny Python library, specialized for internet research.
GPT-3 and GPT-4 are able to use this library to write Python programs that do useful work.
This way they can "crystallize" their plans in code, to ensure that they will run.
Here is the interface:
def search(query, max_results=8) -> List[Dict]: pass # uses duckduckgo
def load(url) -> str | None: pass # uses requests and beautifulsoup
def summarize(text, task) -> str | None: pass # uses gpt-3
def save(filename, text) -> bool: pass
See the comments below the gist for the GPT-3 and 4 versions of main.py (20-30 lines for an internet research agent!).https://gist.github.com/avelican/2d4e718954593e3df9e0e5ee675...
Note: It's currently optimized for my main use-case, which is internet research. So it's not an Auto-GPT in any sense. But it does one thing, and does it fairly well.
P.S. The Holy Grail would be a system where the user enters a query, and the system translates it into Python on top of Auto-GPT's library, and runs that. I haven't tried that yet, and I'm a little afraid to...
My solution to this (not yet integrated) is to use a faster / cheaper model to process all the text first and see if it's actually relevant or not, before running the summarization / task prompt with the main model.
I realized that this is essentially a semantic search engine for text. Using the same principle you can feed a very large text file in, ask it a question, and it finds the page that answers that question.
This would be useful as a layer "beneath" the internet research agent, that it could use to sort through all the noise and answer the question.
https://gist.github.com/avelican/58958a2cf2b7e9f9f555ab94549...
It gives a lot of false positives, I'm not sure if this is a limitation of GPT-3 (have not tried GPT-4 for this yet, a bit too expensive for searching books) or a limitation of my implementation.
It's probably slower and more expensive than just using a vector db, I haven't tried those yet.
All the breathless evangelizing out there is for incredibly simple problems that break down the second we try something complex. Yes a GPT-4 agent can go far in the coming months, but there's a ceiling to how good it can be because there's a ceiling to how well the model can reason. Newer LLMs will be the ultimate answer.
On cost: 3.5 is 1/15th the cost, and a bunch faster though supports less context, so it's worth experimenting with which parts of tasks need GPT-4 and which need 3.5 (perhaps a feature where GPT-4 manages the main tasking / verification and 3.5 handles the individual small tasks - even if GPT has to be called 15 times to get the same result, this approach wins).
On functionality: yeah, this is just some person's plaything. Reading the code it doesn't seem particularly well production grade (like many a python passion program). Something built similarly with a decent architecture (multi-agent + using an approach that automatically optimizes tasks) I can see this getting cheaper and better easily.
Just like in software dev, where we write a spec, write a plan, write acceptance criteria, write code, write tests, run tests, iterate, Auto-GPT type software needs a similar framework to work within that is not just defined by the code, but by generalizing that to an architecture. https://github.com/daveshap/raven is an interesting project exploring some of this.
It saves tons and tons of work though; now that we integrated it into our pipeline after months of tweaking, 3.5 is really starting to remove the need for many of my colleagues. Some of us are needed to say ‘yes or no or redo’ as it were, but everything is far more efficient and much (100x $ or more less per day) cheaper.
Can it? I've played with it a bit and haven't seen that. If someone has some excellent examples, I would love to see.