Swap OpenAI with any open-source model
postgresml.org
postgresml.org
We built an open-source AI SDK (Python & JavaScript) that provides a drop-in replacement for OpenAI’s chat completion endpoint. We'd love to know what you think so we can make switching as easy as possible and get more folks on open-source.
You can swap in almost any open-source model on Huggingface. HuggingFaceH4/zephyr-7b-beta, Gryphe/MythoMax-L2-13b, teknium/OpenHermes-2.5-Mistral-7B and more.
If you haven't seen us here before, we're PostgresML, an open-source MLOps platfom built on Postgres. We bring ML to the database rather than the other way around. We're incredibly passionate about keeping AI truly open. So we needed a way for our customers to easily escape OpenAI's clutches. Give it a go and let us know if we're missing any models, or what else would help you switch.
You can check out the blog post for the details, but here's the git diff:
- from openai import OpenAI + from pgml import OpenSource AI
- client = OpenAI(openai_api_key=openai_api_key) + client = OpenSourceAI(database_url=database_url)
messages = [{"role": "system", "content" : "You are a helpful assistant"}, {"role": "user", "content" : "What is 1+1?"}]
- response = client.chat.completions.create(..) + response = client.chat_completions_create(..)
- return response.choices[0].message.content + return response["choices"][0]["message"]["content"]
Can you explain to me what that means? Maybe I've been having this problem recently.
* 12-factor: support setting via env vars, config files, api, etc
* decoupling auth config from model config
* supporting registration of multiple auths & models, not just one, including via 12-factor
* streamlining ability of llm apps to negotitate which llm models
* inferring & validating matchup of what your model provider gives and your app configures & requests, ideally at config or load time, and in a testable way
* transparent native support for each provider, as each provider & api has annoying deviations & useful features that end up being relevant. Ex: even openai vs azure openai has differences like the notion of 'deployments' and around rate limits that should be handled but also exposed
* observability: introspection hooks, including configuration for opentelemetry metrics, telemetry, & logs, including tenant/user dictionaries
Without that kind of stuff, a third-party dependency is more annoying than useful for 'serious' implementations, b/c we ended up fighting the library vs using
(And we'd be happy to OSS etc if relevant.. such a bear!)
Does openAI support that stuff, or is that a more general MLOps (or whatever you'd call it) list of functionality you need? The API call aspect seems like a pretty easy part that's not a barrier to switching models. It's whatever customizations for a given application that I'd expect to be more work. But re your text I quoted, I'm always wary of "one line of code" solutions that are harder to customize than it would be to just write the the thing.
all that stuff is what we needed to do to support multiple providers vs hard-coding a single one. none of it is specific to our codebase. if we were going to take on a third-party dependency, we'd need it to be serious about this kind of thing, else we have an external dependency to work around on our critical path
Edit: As an example, if the switching layer doesn't implement model negotiation and doesn't expose key model details (which vary by model provider service, and often require REST calls to introspect), we can't add model negotiation / retries / etc on top, and would have to edit the innards to enable that, which opens up all sorts of questions
For me the killer feature of a library like this would be if it implemented function calling. Even if it was for a very restricted grammar - like the traditional ReAct prompt:
Solve a question answering task with interleaving Thought, Action, Observation usteps. Thought can reason about the current situation, and Action can be three types:
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
(3) Finish[answer], which returns the answer and finishes the task.
After each observation, provide the next Thought and next Action. Here are some examples:
One parameter functions would be enough for the start (but maybe something that would work with multi-line strings - this line oriented grammar is kind of restricted).That said, it's not that hard to fine-tune a model to understand function calling -- we do that as part of all of our OpenPipe fine tunes, and you can see the serialization method we use here: https://github.com/OpenPipe/OpenPipe/blob/main/app/src/model...
It isn't particularly difficult, and I'd expect more general-purpose fine-tunes will start doing something similar as they get more mature!
Obviously, if there are things that depend on specific model capabilities (e.g., image interpretation or other multimodal use), you'll need a model that supports that, but there are multimodel open models. You'll need a way to recognize and error on model/capability combos that are invalid though.
Yes, OSS models -- from my understanding, this is true in theory of any model over which you have fine-grained control of inference, which you do if the inference code isn't a black box, but this has been generally implemented in the frameworks for running OSS models -- can be constrained to generate according to a grammar of arbitrary specificity.
Here is an plugin for one of the popular tools for running multiple local models (it is a tool that presents an OpenAI-style API for downstream consumption, as well as presenting its own Web UI for direct interaction with the models.): https://github.com/hallucinate-games/oobabooga-jsonformer-pl...
Unless everybody is going to follow OpenAIs lead for every feature they implement, general compatibility is going to be challenging. Basically, the only reason this is possible now is due to the fact that there isn’t yet a lot of diversity in the way models are being used.
SillyTavern technically supports this, though its not a central use case, so the UI isn't particular convenient for mid-conversation change of models, but it works.
But I don't know of anything where it is a central feature and convenient (and if it was, you could also implement things like parallel requests and choosing which response to keep.)
PanicException: Error getting DATABASE_POOLS for writing: PoisonError { .. }
I can't find anything on google or your docs on DATABASE_POOLS
We are still actively building out our documentation. You can find the current docs here: https://postgresml.org/docs/introduction/machine-learning/sd...
Thank you.