They get task performance by doing a lot more than just feeding a prompt straight to an llm, and then we performance compare them to raw local options.
The problem is, as this secret sauce changes, your use case performance is also going to vary in ways that are impossible for you to fix. What if it can do math this month and next month the hidden component that recognizes math problems and feeds them to a real calculator is removed? Now your use case is broken.
Feels like building on sand.
No one is doing secret math in the backend people are building on. The OpenAI API allows you to call functions now, but even that is just a formalized way of passing tokens into the "raw LLM".
All the features in the comment you replied to only apply to the web interface, and here you're being given an open interface you can introspect.
If you did you'd also know what evals are.
Although performance has varied over time https://arxiv.org/pdf/2307.09009.pdf I also notice that the API allows you to use a frozen version of the model which avoids the worries I mentioned.
https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim...
Overall evals and pinning against checkpoints are how you avoid those worries, but in general, if you solve a problem robustly, it's going to be rare for changes in the LLM to suddenly break what you're doing. Investing in handling a wide range of inputs gracefully also pays off on handling changes to the underlying model.
How do you know that? With SaaS you are at the mercy of the vendor.
It's not particularly more fluid than anything you couldn't whip up yourself (and the repo linked proves that) but there's also not much value in trying to compete with ChatGPT's frontend.
For most products ChatGPT's frontend is the minimal level of acceptable performance that you need to beat, not an maximal one really worth exploring.
If you're letting people do fun long-form roleplay adventures using summarization alongside some sort of named entity K-V store driven by the LLM would be a good strategy.
If you're building a tool that's mostly for internal data, something that leans heavily into detailed answers with direct verbatim citations and having your frontend create new threads when there's a clear break in the topic of a request is a clever strategy since quality drops with context length and you want to save tokens for citations.
People who are saying LLMs suck or are X or are Y are mostly just completely underutilizing them because LLMs make it super easy to solve problems superficially: when it comes to actually scaling those solutions to production you need more than random RAG vector database wrappers.
I'd be curious to hear more about how exactly this works. You do NER on the prompt (and maybe on the completion too) and store the entities in a database and then what? How does the LLM interact with it?
Let's say we want to let our chat remember the character slammed the door last time they were in Village X with the mayor in their presence and have the mayor comment next time they see the player.
Every X tokens we can fire a prompt with a chunk of conversation and a list of semantically similar entities that already exist, letting the LLM return an edited list along the lines of:
entity: mayor
location: village X
priority: HIGH
keywords: town hall, interact, talk
"memory, likelyEffect"[]: door slammed in face, anger at player
Now we have:- multiple fields for similarity search
- an easy way to manage evictions (sweep up lowest priority)
- most importantly: we're providing guidance for the LLM to help it ignore irrelevant context
When the user goes back to village X we can fetch entities in village X and whittle that list down based on priority and similarly to the user prompt.
None of this has any determinism: instead you're optimizing for the illusion of continuity and trading off predictability.
You're aiming for players being shocked that next time they talk to the mayor he's already upset with them, and if they ask why he can reply intelligently.
And to my original point while this works for a game-like experience, you wouldn't want to play around with this kind of fuzzy setup for your companies internal CRM bot or something. You're optimizing for the exact value proposition of your use-case rather than just trying to throw a raw RAG setup at it
I’ve been working on short and long term memory windows at allofus.ai for about 6 months now and it’s way more complex than I had originally thought it would be.
Even if you can magically extend the content window, the added data confuses and waters down the reasoning of the LLM. You must do layered abstraction and compression with goal based memory for it to continue to reason without distraction of irrelevant data.
It’s an amazing realization, almost like a proof that memory is a kind of layered reasoning compression system. Intelligence of any kind can’t understand everything forever. It must cull the irrelevant details, process the remains and reason on a vector that arises from them.
I don't really know what any sort of "big leap" beyond this people are expecting, incremental performance for sure. But what else?
When you enable it, it is pretty shocking. And it’s pretty simple to enable. You just give it a meta instruct to decide when to message you and what to store to introspect on.
*Edit Azure chatgpt, would be amazed/disappointed if chatgpt used langchain.
Also this is the azure repo from OP, nothing to do with the actual ChatGPT front-end that was asked about. I highly doubt the official ChatGPT front-end uses langchain, for example.