OpenAI doesn't have some sort of egress feed for your database.
OpenAI doesn't have some sort of egress feed for your database.
That's what they're trying to incentivize, especically with being able to upload files for their own implementation of RAG. You're not getting the vector representation of those files back, and switching to another provider will require rebuilding and testing that infrastructure.
The developer experience is lacking vs. other vector database providers and the performance doesn't match those that prioritize performance rather than devex. You're also spending time writing plumbing around postgres that isn't really transferrable work.
For some people already in the ecosystem it will make sense.
You're thinking of traditional apps and APIs.
In an AI application, most of the work is in prompt engineering, not wiring up the API to your app. Prompts that work well for one model will fail horribly for another. People spend months refining their prompts before they're safe to share with users, and switching platforms will require doing most of that refinement over again.
Remember: a good model with a good prompt will generate bad outputs sometimes.
A bad model with a bad prompt will generate a good output sometimes.
That is simply a fact with these non deterministic models.
You have to do many iterations for each prompt to verify they are working correctly.
> I’ve not had much problems moving between LLMs…
If you want to move your prompts to a different model, you’re effectively replacing one:
f(prompt + seed) => output
With different black box implementation.
Unless you’re measuring the output over multiple iterations of (seed) and verifying your prompt still does the right thing, it’s actually very likely that what you’ve done if take an application with a known output space and converted it to an application with an unknown output space…
…that partially overlaps the original output space!
So it looks like it’s the same.
…but it isn’t, and the “isn’t” is in weird edge cases.
Unless you’re measuring that, you simply now have an app that does “eh, who knows?”
So yes. Porting is trivial if you don’t care if you have the same functionality.
…but reliably porting is much harder (or longer).
I think we're exiting the phase where people can launch an AI app and have people use it just because of the initial "wow factor" and moving into the phase where users will start churning and businesses will need to make sure that their AI agent is performing and they they understand how well it's performing.
BTW its much faster and cheaper to artive at a good prompt if you sample the model in deterministic mode (ie temperature=0)
By default you have to guess if the difference is due to the prompt change or due to the dice roll, as you’ve noticed, but you don’t need to!
This is degenerate (greedy) behaviour, and not representative of the what the prompt will behave like at a higher temperature.
(At least, that’s my understanding; it’s a complex topic but broadly speaking there no specific reason, as far as I’m aware, to expect that a particular combination of params/prompt is representative of any other combination of params/prompt for the same model; it may be, but it may not. Certainly on models like GPT4 it is not, for reasons that are not clear to anyone. So… take care with your prompt testing. setting temperature to 0 is basically meaningless unless you expect to use a temperature of 0 in production. The results you get from your prompts at temp 0 are not generally reflective of the results you will get at temp > 0).