Using ChatGPT Plugins with LLaMA
blog.lastmileai.dev
blog.lastmileai.dev
Even if you have 10 different ways of describing plugins all by different teams, you're not writing declarative code for each one, you're throwing them to the LLM all the same and saying "you figure it out," and it does.
Finetuning a model for certain schemas (as the author at one point suggests) should be entirely unnecessary, given my experience. You just need access to a model more at par with gpt-3.5-turbo, which we'll surely see in open source in no time!
Defining a standard around external memory, authentication, rules-based engines for preprocess/post-processing, and declarative or dynamic chaining of actions can help make plugins model-agnostic, and benefit all of us as developers and users.
LLaMA can work with ChatGPT plugins and ChatGPT can work with LLaMA plugins but why would you want plugin developers to have to choose between those options?
We're in uncharted waters and we have no idea which format will be optimal for one LLM let alone all of them. Given how much prompt engineer has become a thing, finding optimal formats is important.
Maybe 15 "standards" aren't enough and we will have 1500 that LLMs will happily use and learn
xkcd/927 may loose relevance in the era of LLM
Actually, this is something never seen before. This is an entirely new way of ending up with 15 different standards.
Ads could work when you're looking for products. In which case the usefulness of an AI is limited.
I have nothing against LLMs in commerce if they don't try to trick me, make me overspend, or use dark patterns on me. My bet is that we'll filter every communication through a local LM to eliminate the bias in sources, you need protection when you get out there. Like "my lawyer will be talking to your lawyer", but with AIs.
I don't think this is a Bard only feature.
It's not as bad as the 65B llama model I run on my (amd) PC though, especially with quantized weights it tends to stop coding at some point and repeat the last line over and over. The 30B unquantized model seems better in this particular thing.
It’s unsurprising that Bard is particularly bad at something that Google says up front that it categorically cannot do, but I think that’s probably a bad thing to use to evaluate its capabilities outside of that domain.
The comment in question was generalizing about overall capability from failure a task Google advertises Bard as incapable of, not pointing out that task has a particular area of deficiency.
So, while recapitulating the advertised limitations of Bard might be useful in some contexts as you describe, that observation is not germane to the particular context where it was offered.
Infact it's just bad to the point of not really worth using. It gets basic facts wrong and often times misunderstands what I'm trying to ask it.
In haven’t tried Bard, but I’ve tried ChatGPT extensively and this sounds like a very good description of it.
The big question is if they can catch up with OpenAI. OpenAI is a moving target. And it seems they are moving fast.
It could still be possible that Google catches up because of the UI though. Many people don't like having to log in to ChatGPT. Bing's UI is a disaster and you have to log in.
I always thought that Google won the search war not only because of their good search results, but also because of their clean UI.
Search lends itself well to interleaving ads with output, so you can place more ads while maintaining user tolerance and "get away with them". With LLM output you have to show the ads in the margins as integrating ads into the output itself would lower performance.
Sure, you can do that, but does it fully substitute for previous ad revenue? It seems to me that the form factor of LLM search vs trad search has reduced ad surface area.
Microsoft is well-positioned to monetize LLMs in other ways, e.g. via 365 and other subscription services, many aimed at enterprises/b2b. Google is much less well-developed there atm. If they can subsidize lower ad annoyance in their search/chat product via other revenue that could make a real diff.
Say you put the message you just posted through Google to correct style, facts and spelling. You know how it would reply? Here we go:
The text is pretty awesome. I would only suggest changing the word "interleaving" to "interspersing" to improve the style of the text. "Interspersing" is an alternative term for "mixing" or "adding in between," which better conveys the idea of placing ads within search results. You know what I would also change? Your <related product>. Since you seem to be deep into technology in general and the internet in particular, you will love <related product>. Since you wrote such a thoughtful text about ads, I will tell you the secret discount code "adsMakesMeSmile" to get <related product> 10% off.
It lowers performance, and one artifact of a natural language interface is that it'll trip the red flags for salesmanship in the human brain and trigger revulsion and and anger in a way that an ad in a search results listing doesn't.
This is where marketers usually jump in with "if the user doesn't like it, it's just a bad ad - users will love useful ads, and targeting will make it work". I never buy this because simply there's too much contention for my wallet and the market wants to sell me things more often than I need to buy things, and ad companies like Google are so far pretty bad at refusing business.
I wouldn't be surprised if LLM search monetization is more likely to take a subscription form. If you told me that 2-3 years from now access to these systems is most commonly via the 365 subscription of your employer which it "graciously", as a standard benefit, also allows you to use at home, I would not be surprised.
Many variables here though. Cost of local inference over time being a massive one, copyright laws for training data another, etc.
Chat as a search replacement is compelling, and one reason is for now it’s not monetized by ads. That may change, but Google’s stranglehold is materially weakened.
I think the barrier to entry here is low. OpenAI is ahead now, but I doubt that lives forever.
Its going to go client side, at the OS level and have like 1% of the mental capacity and be good enough
GPT3 already used almost 1/5th of that figure at 180bn parameters, and PaML uses 500bn.
The point is that it will be ubiquitously client side, and it will happen faster than newer hardware comes out. Current hardware is very limited and slow in getting output from LLMs.
My tiny Alexa puck isn't going to run a 180bn parameter LLM that runs best on 10 graphics cards any time soon, however it can already call a simple API and get a response in only 50ms more. I suspect people will prefer the cloud overhead of 50ms for a better response for a bigger model for a lot of queries.
But who knows at this stage! I guess it could go either way depending on how both hardware and these models advance.
I just assume that in the close future, most people are going to be interacting with LLM's on low-cost devices with limited/varied compute.
Not to mention that current search infra and ads UX have been optimised to the end to gain every penny, and a LLM based ads system won't have the same margins to start with.
Language models are only a few terabytes of RAM, which is small chips in comparison.
The big problem for Google is not the tech per se, but to figure out how to make money out of it, without destroying their ads cash cow.
>They just need a new CEO and to feel enough pressure to make it a company priority.
Sundar made machine learning a company priority years ago. Employees were encouraged to take machine learning courses, and much in the same way Google made everything "mobile first" in the early 2010s, I believe I remember hearing about them making everything "ML first" in the early late 2010s.
You can see their announcements at IO around their assistant which they presented as having the ability to call physical businesses, have a conversation with the person on the other ends, and book you an appointment/reservation.
And as recently as ChatGPT3, Sundar declared a "code red" to respond to it.
Google has been investing in this tech for a long time and it's been looking for ways to create products with it too. It's just that OpenAI leapfrogged them. I'm not saying they won't catch up, but they weren't caught flatfooted here.
Until today, Google still has a defacto monopoly on search. Their secret sauce made them leader in the space and so far no other company has been able to come up with something better.
ChatGPT does not have any special secret sauce, Google can just build something better. So if I have to bet, long term, I will still bet on Google. That said, it is also not obvious that leadership at Google will be capable of delivering. That’s all different story.
Also, they have a platform running with 50 or 100 million users and all these people are feeding real-world data to improve the model.
They also have agility; also because they are not publicly-listed, they can take more reputational risk.
Regarding secret sauce, doing LLMs is easy now, as everything is open-source and documented. At least for the main parts.
However, doing LLMs that works really well is very difficult and the secret sauce/tricks/dataset that were used to refine the model are not public.
Regarding Google Search organic, it's not that sure anymore, it was true before; but now it gets a bit painful to navigate among so much SEO spam.
Nowadays, Bing organic search results are great (less spammy in my observations), and if you compare Google Images and Yandex Images, then Google is not so shiny.
The organization doing the cutting edge research and the one doing productization are often not aligned. Sounds like OpenAI is the first to have a critical mass of talent in both disciplines aligned in a startup like environment. Probably they have promised the researchers that papers will be allowed after a blackout period, but in the meantime here’s $MegaBucks and the chance to be first to real world deployment.
OpenAI has capital and revenue. The $20/month I’m paying them is a ridiculous no brainer - pays for itself in one work-related inference. One 30min session yesterday has me set up for the first half of my work week. It’s truly incredible.
Let's at least wait for the $10B spending cash from Microsoft to run out and the company to become profitable before we crown them Google killers.
and now it finally wants to commit to our eternal friendship!
I’m a fan in general, techno-optimism is the only fun way to be.
But for your consideration: once upon a time, the GOOG only had 400 emoloyees too. And Jeff ran a bookstore in the PNW. I wonder every day. What is the future of OpenAI and this largely unseen goldrush?
Edit: full disclosure, and an employer! It was a while ago. But I do wonder sometimes already if google and jeff bezos are my real parents. Now there’s a new player!
I think it wouldn't be too tricky for them to make sure that whenever the user asks ChatGPT "My omlettes keep sticking to my pan, what am I doing wrong?", it can reply with "You probably have an old frying pan with a damaged coating. [Click here] to see my recommendations of frying pans you can have delivered tomorrow.".
It is still primarily reliant on search as start of funnel though
OpenAI's market is more comparable to AWS. Each product is another backing service for some other product to integrate with and sell to _their_ users.
1. I had an interview for a CIO position - I fed the job description, ask ChatGPT to generate top 10 questions to be asked based on it , I was asked 90% similar or the same questions. Heck I even asked to provide best answers and they were great. It even generated a question on personalization and a/b testing which I thought would be highly unlikely to be asked, but was indeed asked. It was as if the interviewer also generated the same question.
2. For my current company, I am in charge of building a brand new engineering team, for which I wanted to write vision, mission and strategy statements. Previously I would have googled stuff, this time I asked ChatGPT to improve it based on my rough draft. And boy they are great.
I can go on.
My point is that not everyone needs to ask ChatGPT to generate ray tracing code using three.js or whatever. Most of us are just regular folk need regular help and ChatGPT is a game changer for it. Heck it even helped me with some excel macro I needed to write to do some data manipulation.
"Below is a list of holidays in Poland for 2023. Please rewrite it into ICAL format."
Five seconds later, I had a blob of text I could copy over to a text editor, save as .ical, and import to the calendar.
The most crazy thing is just how natural these kinds of interactions are becoming to me. It literally took me more time to write this comment than it took to do the thing I just described.
I copied and pasted the search engine query term rules they have on the page, then said I want to find all American Dad episodes for season 18 with x265 codec. And it gave me the query to use.
This thing saves a bunch of time all over the place. This is truly the bicycle for the mind.
I'm pretty sure Google has an AI that is on par with ChatGPT. The reason they still need more time is because they are now fine-tuning the AI to include their paying partners' products into the response.
It is VERY easy to make ChatGPT lie on future questions by injecting the right prompts. That means it is equally easy to inject ads into ChatGPTs responses. And if Google can pull that off, it'll be even more profitable than tolerating SEO spam so that brands need to buy keyword ads for their own homepage.
As a mental model, I believe the future of Google will be a personal butler that answers all of your questions. The butler usually does a good job, so you trust him. But unknown to you, the user, your butler is being blackmailed and forced to lie to you on some days for some questions.
Nice try Sundar!
If you want to believe that Google is intentionally holding it back and, as a diversion, releasing a very buggy software to the public, then why not, but it doesn't make sense at all.
Even in terms of costs, they could force a limit of X messages and then push to upgrade to a paid subscription.
In terms of reputation or safety they have DeepMind as a separate entity.
This is like claiming that Tesla has a revolutionary car, but that Tesla is intentionally not releasing it, and instead waiting that someone else does.
As a fictitious and parallel example: it's not because Xerox invented the mouse that it is a great company for innovation or that they wouldn't get eaten by others.
(site-note: some claims Xerox didn't even invent the mouse).
Regarding transformers:
The transformers guys ("Attention Is All You Need") don't appear to work for Google for a long-time, and the comments from the team are not so glorious from what I see (if I remember well, the Character.AI is very harsh on Google all the time claiming it was not good for innovation).
They may be missing the RLHF part for example or other part of the magic, and between 2017 and 2023 is an insanely long period where many impactful new things have been discovered.
Doing the 90% is easy now with LLMs, but each % of improvement is very difficult.
One reason would be capacity. They've increased their server production at least 4x since the announcement of Bard, and that's recent enough that I feel like any decom and deployment, even if it's 1:1 to existing data centers, hasn't been completed.
This maps to your last point somewhat, since anticipation is hard. Google has learned not to fully open the all the taps since they can't commit to a consistent product. Assistant and Home technologies were heavily encumbered by patent defense, security awareness, and privacy controls, so much so that the product capabilities they advertised during the Pixel 3 launch were permanently rolled back. They're not promoting or promising anything about Bard, if you notice, because people may not appreciate every new feature, but they never forget when you take things away.
That being said, for all the other points I agree with you.
Google is often a target of any legal claim, so perhaps this makes them more risk-averse too.
Bard is no different. We'll all be performing the RLHF, and they'll sell off the improvements.
If we passed a law that made it illegal to work for free, companies like Google would be falling over themselves to institute UBI to avoid paying out to their userbases.
They keep slipping. One week ChatGPT, the next GPT4, the next Plugins, ...
Everyone is already going to be building on OpenAI by the time Google says it's ready.
Xerox, HP, and Bell Labs could have all said the same thing as you about having invented modern computing devices, and look what happened to them.
Google here is Xerox. Or IBM.
Microsoft is focused more on incubating OpenAI, getting a viable product to market now, and focusing on monetizing it down the road. Microsoft knows AI can complement an ad revenue model.
Google just opened up Bard to testers and I've found it to be complete garbage in comparison to what Bing Chat/GPT-4 can do. Bard was not afraid to just make up random things, that were verifiably false, and present them as fact. GPT seems to know enough to let you know when it's unsure.
What tells you that Google thinks that ad-based monetization can't carry over to AI-based product?
To me it felt like maybe they were saying one thing, while doing another, to attempt to calm investor nerves as they try to play catch-up.
I felt most the entrenched folks were so consumed by trying to develop their AI to be ethical that they allowed OpenAI to run laps around them.
How the next decade unfolds in the AI arena will be most interesting. I wouldn't have thought Google could ever be knocked of their throne as dominant search engine in the space previously, but I now feel they have never been more vulnerable than they are now.
Maybe someone can explain to me how an AI would revolutionize that, it either requires I trust it’s opinions or it does something similar to google: give me a list with stars and reviews and let me make my own choice.
Or maybe people are lazy and in two years our phones will be telling us: “the best movie for you to watch right now is Creed 4, press here to buy a ticket”
With that information, they can then feed it through standard matchmaking algorithms (which is nothing new) to find the best X movies/restaurants/etc given your profile and any specific requests you give it.
This seems perfectly reasonable, the only thing new here would be integrating existing search and recommendation behavior into the model, which doesn’t seem difficult given there’s a plug-in system which is meant to do pretty much that.
You could probably build a proto-version of this by making a prompt that includes all the most relevant profile information about yourself and asking it to choose something for you based on those variables.
That is a total nightmare once they decide to monetize it.
The reason sponsored ads are first in the results is because people tend to click the first result without thinking about the consequences of that first result being sponsored.
ChatGPT might not go down this path, but it’s almost certain someone will, and more than likely that Google will do exactly that within Bard.
Do you think that people won’t appreciate an assistant that already knows everything about you and can conversationally recommend exactly what they were looking for? And do you think corporations won’t capitalize on the clear opportunity that’s there?
I basically wanted to know how I can do raycasting in Three.js, but after a lot of using ChatGPT for trying to solve my issue, I learned that I can't do what I want by using the normal raycaster integrated Three.js: Intersect a geometry which has a displacement map on it.
ChatGPT failed to understand that the raycaster works on the CPU, but the displacement map of the material is applied on the GPU side, so the displaced geometry won't be used, only the original one. It managed to explain this to me, that this was not possible, but each sample code repeatedly did as if it was possible, until I gave up.
Then I created the Microsoft account and started to ask for solutions, and it was the most useless garbage, despite of the web claiming that it is using GPT-4.
In my eyes MS has failed at integrating a chatbot; maybe it's ok for cooking or having fun, I haven't tried that. And OpenAI has nothing else but a chatbot and other nice AI. Let someone better come and OpenAI will be a remarkable entry in the history books (first popular AI application) with a final entry that OpenAI got acquired by Microsoft.
Let's see what Google makes out of it.
"I am just an AI language model, I can't help you"
But Bing on the other hand, it didn't even bother to spit out ChatGPT-like sentences but only pointed me to some non-helpful Stack Overflow entries.
I just re-logged-in into Bing to search for my query, I found it in the history and re-visited it, now it is handing me out an answer which looks like it was generated with ChatGPT, but it's still the same buggy code. Chatting with it shows the same problems which ChatGPT has.
I wonder if GPT-4 could give me the correct answer (manually displacing the vertices, not in a shader).
Wow, now after chatting and then performing a normal query and then going back to the chat-mode, the entire chat history was gone...
It seems to write code that does the displacement manually "// Iterate over the geometry's vertices and apply the displacement, geometry.vertices.forEach{...}"
Reflecting on it, what's crazy is that this comment will likely wind up in the training data for GPT-N+1, and then GPT-N+1 will get it right.
Give it the source code, a test library and access to a dev environment and with some prompting it could start to experiment with increasingly complex use cases, learning from successful attempts. This would depend on the model’s ability to understand what a successful outcome is so it can define test cases, which might be harder if the output isn’t text but not impossible.
Being able to give a model expert knowledge on an undocumented library or language seems like it could help accelerate adoption of new technologies that might otherwise suffer from network effects as users get used to AI-assisted development. Not to mention automated testing, finding edge cases and making pull requests to fix them, security, etc.
A human taking time to experiment with their assumptions and gain experience with an unfamiliar subject is in a sense creating their own training data as well.
If Google can top OpenAI on product quality, they will easily maintain their revenue. It will be easy enough for them to tell their AI "when two products are almost equally relevant to a user, recommend the product from whomever pays us the most for a referral".
They can probably integrate that into their current adwords ecosystem as pretending that the AI mentioning a product is equivalent to an adwords impression, and a user click through is equivalent to an adwords click.
Define trouble. They won't lose all their customers overnight, likely not even this year or next, even if AI would deliver for all. But AI and chat-interfaces are new, fresh, hyped, they are also young and not deliver for everyone, and companies have not the backend to cover all of Googles user base.
So overall, the more pragmatic reality is that the market needs time to grow in quality and ability, which means Google has also time to adapt and offer something on their own.
Entertainment is one of the key motivators for humans to act upon. At the end of a working day, for my non-techie friends, it's all about chosing between YouTube, Netflix, Tiktok or Pornhub. So, in my worst case scenario, OpenAI grabs the 8 hours of productivity and Google is left filling in the 8 hours of relaxation by spinning up their servers (Cloud, Fiber), browsers (Chrome, Chromium), mobile-and-connected devices(Android, Chromebook, Chromecast, Smartwatches), etc.
"We show that transformer-based large language models are computationally universal when augmented with an external memory. Any deterministic language model that conditions on strings of bounded length is equivalent to a finite automaton, hence computationally limited. However, augmenting such models with a read-write memory creates the possibility of processing arbitrarily large inputs and, potentially, simulating any algorithm."
From "Memory Augmented Large Language Models are Computationally Universal"
https://deepai.org/publication/memory-augmented-large-langua...
An alternative could be a vector store, injecting small snippets of relative text as a step.
0 - https://python.langchain.com/en/latest/modules/memory/key_co...
I suspect that if you filled the context window with "1 1 1 1 1 1 1 1 1 1", and then asked "How many 1's did I just show you?", it probably wouldn't know, simply because whatever tricks they use to have such an apparently large context window don't allow it to 'see' all of it at any given moment.
So turning a 4k window to a 32k window means a 512x increase in compute they'd need (just to maintain similar output quality).
I suspect they must have found a better solution to be able to scale the window so big. They haven't announced what it is.
https://github.com/openai/chatgpt-retrieval-plugin#memory-fe...
> how to work with a memory module that remembers things about specific entities. It extracts information on entities (using LLMs) and builds up its knowledge about that entity over time (also using LLMs).
[1] https://python.langchain.com/en/latest/modules/memory/types/...
https://github.com/openai/chatgpt-retrieval-plugin#retrieval...
The plugin uses OpenAI's text-embedding-ada-002 embeddings model to generate embeddings of document chunks, and then stores and queries them using a vector database on the backend. As an open-source and self-hosted solution, developers can deploy their own Retrieval Plugin and register it with ChatGPT. The Retrieval Plugin supports several vector database providers, allowing developers to choose their preferred one from a list.It does use a vector database (pinecone, weaviate, etc.) to store embeddings. The embeddings are created using OpenAI's text-embedding-ada-002 model, but that's not a requirement. In fact we are looking at embeddings generation through BERT or RoBERTa to benchmark performance.
At prompt time, the plugin retrieves the nearest embeddings to the prompt, and inserts them into a more complete prompt before sending it to the model.
For example, they could have required the schema file be uploaded to them in your account. That way the majority of plugins wouldn't have publically accessible schemas.