OpenAI compatibility
ollama.ai
ollama.ai
1: Quite literally hours ago: https://euri.ca/blog/2024-llm-self-hosting-is-easy-now/
I just played around with this tool and it works as advertised, which is cool but I'm up and running already. (For anyone reading this though who, like me, doesn't want to learn all the optimization work... I might see which one is faster on your machine)
'llama.cpp-based' generally seems like the norm.
Ollama is just really easy to set up & get going on MacOS. Integral support like this means one less thing to wire up or worry about when using a local LLM as a drop-in replacement for OpenAI's remote API. Ollama also has a model library[1] you can browse & easily retrieve models from.
Another project, Ollama-webui[2] is a nice webui/frontend for local LLM models in Ollama - it supports the latest LLaVA for multimodal image/prompt input, too.
It's also possible to connect to OpenAI API and use GPT-4 on per token plan. I cancelled my chatGPT subscription since. But 90% of the usage for me is Mistral 7B fine-tunes, I rarely use OpenAI
re: cancelling ChatGPT subscription: I am tempted to do this also except I suspect that when they release GPT-5 there may be a waiting list, and I don’t want any delays in trying it out.
No special flags or anything, just the standard format. Do take care of the spaces and end of lines. sharing a gist of the function I use for formatting it: https://gist.github.com/theskcd/a3948d4062ed8d3e697121cabd65... (hope this helps!)
Straight up giving large files tends to degrade performance (so you need to do some reranking on the snippets before sending them over
they separate serving heavy weights from model definition and usage itself.
what that means is weights of some model, let's say mixtral are loaded on the server process (and kept in memory for 5m as default) and you interact with it by using modelfile (inspired by dockerfile) - all your modelfiles that inherit FROM mixtral will reuse those weights already loaded in memory, so you can instantly swap between different system prompts etc - those appear as normal models to use through cli or ui.
the effect is that you have very low latency, very good interface - for programming api and ui.
ps. it's not only for macs
open weight models + (llama.app) as ollama + ollama-webui = real openai.
You can mix models in a single model file, it's a feature I've been experimenting with lately
Note: you don't have to rely on their model Library, you can use your own. Secondly, support for new models is through their bindings with llama.cpp
> A few pip install X’s and you’re off to the races with Llama 2! Well, maybe you are, my dev machine doesn’t have the resources to respond on even the smallest model in less than an hour.
I never tried to run these LLMs on my own machine -- is it this bad?
I guess if I only have a moderate GPU, say a 4060TI, there is no chance I can play with it, then?
I also have the 16GB version, which I assume would be a little bit better.
Unfortunately, having tried this and a bunch of other models, they are all toys compared to GPT-4.
I still need GPT-4 for some tasks, but in daily usage it's replaced much of ChatGPT usage, especially since I can import all of my ChatGPT chat history. Also curious to learn about what people want to do with local AI.
Does anyone know why this would be?
- get a Mac Mini or Mac Studio - just run ollama serve, - run ollama web-ui in docker - add some coding assitant model from ollamahub with the web-ui - upload your documents in the web-ui
No code needed, you have your self hosted LLM with basic RAG giving you answers with your documents in context. For us the deepseek coder 33b model is fast enough on a Mac Studio with 64gb ram and can give pretty good suggestions based on our internal coding documentation.
[1] https://docs.google.com/document/d/1OpZl4P3d0WKH9XtErUZib5_2...
wonder what pain points people have around the API becoming a standard, and if anyone has taken a crack at any alternative standards that people should consider.
The power of open source!
I'm fine with it emerging as a community standard if there's a REALLY robust specification for what the community considers to be "OpenAI API compatible".
Crucially, that standard needs to stay stable even if OpenAI have released a brand new feature this morning.
So I want the following:
- A very solid API specification, including error conditions
- A test suite that can be used to check that new implementations conform to that specification
- A name. I want to know what it means when software claims to be "compatible with OpenAI-API-Spec v3" (for example)
Right now telling me something is "OpenAI API compatible" really isn't enough information. Which bits of that API? Which particular date-in-time was it created to match?
To consume them, just assume that every field is optional and extra fields might appear at any time.
- `tool_choice={type: "function", function: {name: "getWeather"}}`, where the developer can force a specific tool to be called. - `tool_choice="none"`, where the developer can force the model to address the user, rather than call a tool.
If you have any other feedback, please feel free to email me at atty@openai.com. Thanks!
To confirm, the `functions` parameter will continue to be supported.
We renamed `functions` to `tools` to better align with the naming across our products (Assistants, ChatGPT), where we support other tools like `code_interpreter` and `retrieval` in addition to `function`s.
If you have any other feedback for us, please feel free to email me at atty@openai.com. Thanks!
OpenAI compatible just seems to mean 'you can format your prompt like the `messages` array'.
I've been needing something exactly like this to test against in local dev environments :) Ollama having this will make my life / testing against the myriad of LLMs we need to support way, way easier.
Seems everyone is centralizing behind OpenAI API compatibility, e.g. there is OpenLLM and a few others which implement the same API as well.
- We first struggled with token limits [solved]
- We had issues with consistent JSON ouput [solved]
- We had rate limiting and performance issues for the large 3rd party models [solved]
- We wanted to reduce costs by hosting our own OSS models for small and medium complex tasks [solved]
It's like your product becomes automatically cheaper, more reliable, and more scalable with every new major LLM advancement.
Obivously you still need to build up defensibility and focus on differentiating with everything “non-AI”.
How has this been solved in your opinion? Do you mean with recent versions with much bigger limits but also heaps more expensive?
It's nice that you have the role and content thing but that was always fairly trivial to implement.
When it gets to agents you do need to execute actions. In the agent hosting system I started, I included a scripting engine, which makes me think that maybe I need to set up security and permissions for the agent system and just let it run code. Which is what I started.
So I guess I am not sure I really need the function/tool calling. But if I see a bunch of people actually am standardizing on tool calls then maybe I need it in my framework just because it will be expected. Even if I have arbitrary script execution.
Function calling/tool choice is done at the application level and currently there's no standard format, and the popular ones are essentually inefficient bespoke system prompts: https://github.com/langchain-ai/langchain/blob/master/libs/l...
Is this true for open ai - or just everything else?
Anyway, probably best that they didn't release support that doesn't work.
curl https://ollama.ai/install.sh | sh
However, that script asks for root-level privileges via sudo the last time I checked. So, if you want the tool, you may want to download the script and have a look at it, or modify it depending on your needs.https://www.debian.org/doc/debian-policy/ch-archive.html#the...
The whole thing, actually:
[0] https://github.com/ollama/ollama/blob/main/docs/linux.md#man...
it's great to have some standard API even if that's isn't perfect, but having second API that allows you to use full potential (like B2 for backblaze) is also fine
so there isn't one model fits all, and if your model have different capabilities, then imo you should provide both options
It's a little bit easier to use if you want to do this without an HTTP API, directly in Python.
If Ollama doesn't have a cli flag that disables auto updating and networking altogether, I'm not letting it anywhere near my production environments. Period.
I know not everyone uses LangChain, but I thought that was one of the primary use-cases for it.
Its what I ended up doing.
(Better for certain use cases, that is, I’m not saying LangChain doesn't have uses.)
Also Autogen seems popular and well-ish liked https://microsoft.github.io/autogen/
LangChain definitely has the most market-/mind- share. For example, GCP has a blog post on supporting it: https://cloud.google.com/blog/products/ai-machine-learning/d...
The main reason is langsmith. (But there are other reasons too). Because of langchain we got "free" (as in no development necessary) langsmith integration and now I can debug my llm.
Before that it was trying to make sense of whats happening inside my app within hundreds and hundreds of lines of text which was extremely painful and time consuming.
Also, lc people are extremely nice and very open/quick to feedback.
The abstractions are too verbose, and make it difficult, but the value we've been getting from lc as a whole cannot be overstated.
other benefits:
* easy integrations with vector stores (we tried several until landing on one but switching was easy)
* easily adopting features like chat history, that would've taken us ages to determine correctly on our own
people that complain and say "just call your llm directly": If your usecase is that simple, of course. using lc for that usecase is also almost equally simple.
But if you have more complex use cases, lc provides some verbose abstractions, but it's very likely that you would've done the same.
I'm building a React Native app to connect mobile devices to local LLM servers run with these programs.
import OpenAI from 'openai'
const openai = new OpenAI({
baseURL: 'http://localhost:11434/v1',
apiKey: 'ollama', // required but unused
})
const chatCompletion = await
openai.chat.completions.create({
model: 'llama2',
messages: [{ role: 'user', content: 'Why is the sky blue?' }],
})
console.log(completion.choices[0].message.content)
I am getting the below error: return new NotFoundError(status, error, message, headers);
^
NotFoundError: 404 404 page not foundLlama.cpp is not far behind, but I find the well structured python code of transformers easy to modify and extend(with context free grammars, function calling etc) than just waiting for your favourite alternate runtime support a new model.
I never liked ollama, maybe because ollama builds on llama.cpp (a project I truly respect) but adds so much marketing bs.
For example, the @ollama account on twitter keeps shitposting on every possible thread to advertise ollama. The other day someone posted something about their Mac setup and @ollama said: "You can run ollama on that Mac."
I don't like it when +500 people are working tirelessly on llama.cpp and then guys like langchain, ollama, etc. rip off the benefits.
I don't know who is behind Ollama and don't really care about them. I can agree with your disgust for VC 'open source' projects. But there's a reason they become popular and get investment: because they are valuable to people, and people use them.
If Ollama was just a wrapper over llama.cpp, then everyone would just use llama.cpp.
It's not just marketing, either. Compare the README of llama.cpp to the Ollama homepage, notice the stark contrast of how difficult getting llama.cpp connected to some dumb JS app is compared to Ollama. That's why it becomes valuable.
The same thing happened with Docker and we're just now barely getting a viable alternative after Docker as a company imploded, Podman Desktop, and even then it still suffers from major instability on e.g. modern macs.
The sooner open source devs in general learn to make their projects usable by an average developer, the sooner it will be competitive with these VC-funded 'open source' projects.
It takes literally one line to install it (git clone and then make).
It takes one line to run the server as mentioned on their examples/server README.
./server -m <model> <any additional arguments like mmlock>Sorry, I'm new to ollama 'ecosystem'.
From llama.cpp readme, I ctrl-F-ed "Node.js: withcatai/node-llama-cpp" and from there, I got to https://withcatai.github.io/node-llama-cpp/guide/
Can you explain how ollama does it 'easier' ?
I've already got a web UI that "should" work with anything that matches OpenAI's chat API, though I'm sure everyone here knows how reliable air-quotes like that are when a developer says them.
You also don't need to actually install my web UI, as it runs from the github page and the endpoint and API key are both configurable by the user during a chat session.
Also (a) the ollama command line interface is good enough for what I actually want, (b) my actual problem was not realising I'd only installed the python and not the underlying model.
> pip install ollama
- https://ollama.ai/blog/python-javascript-libraries
is just the python libraries, not ollama itself, which the libraries need, and without which they will just…
> httpx.ConnectError: [Errno 61] Connection refused
Install the main app from the big friendly download button, and this problem fixed itself: https://ollama.ai/download
Example use case would be to support a web application with, say, 100k DAU.
https://github.com/triton-inference-server/tensorrtllm_backe...
It’s used by Mistral, AWS, Cloudflare, and countless others.
vLLM, HF TGI, Rayserve, etc are certainly viable but Triton has many truly unique and very powerful features (not to mention performance).
100k DAU doesn’t mean much, you’d need to get a better understanding of the application, input tokens, generated output tokens, request rates, peaks, etc not to mention required time to first token, tokens per second, etc.
Anyway, the point is Triton is just about the only thing out there for use in this general range and up.
What I like about vLLM is the following:
- It exposes AsyncLLMEngine, which can be easily wrapped in any API you'd like.
- It has a logit processor API making it simple to integrate custom sampling logic.
- It has decent support for interference of quantized models.
Mistral[0]:
"Acknowledgement We are grateful to NVIDIA for supporting us in integrating TensorRT-LLM and Triton and working alongside us to make a sparse mixture of experts compatible with TRT-LLM."
Cloudflare[1]: "It will also feature NVIDIA’s full stack inference software —including NVIDIA TensorRT-LLM and NVIDIA Triton Inference server — to further accelerate performance of AI applications, including large language models."
Amazon[2]: "Amazon uses the Text-To-Text Transfer Transformer (T5) natural language processing (NLP) model for spelling correction. To accelerate text correction, they leverage NVIDIA AI inference software, including NVIDIA Triton™ Inference Server, and NVIDIA® TensorRT™, an SDK for high performance deep learning inference."
There are many, many more results for AWS (internally and for customers) with plenty of "case studies", and "customer success stories", etc describing deployments. You can also find large enterprises like Siemens, etc using Triton internally and embedded/deployed within products. Triton also runs on the embedded Jetson series of hardware and there are all kinds of large entities doing edge/hybrid inference with this approach.
You can also add at least Phind, Perplexity, and Databricks to the list. These are just the public ones, look at a high scale production deployment of ML/AI in any use case and there is a very good chance there's Triton in there.
I encourage you to do your own research because the advantages/differences are too many to list. Triton can do everything you listed and often better (especially quantization) but off the top of my head:
- Support for the kserve API for model management. Triton can load/reload/unload models dynamically while running, including model versioning and config params to allow clients to specify model version, require specification of version, or default to latest, etc.
- Built in integration and support for S3 and other object stores for model management that in conjunction with the kserve API means you can hit the Triton API and just tell it to grab model X version Y and it will be running in seconds. Think of what this means when you have thousands of Triton instances throughout core, edge, K8s, etc, etc... Like Cloudflare.
- Multiple backend support with support for literally any model: TF, Torch, ONNX, etc with dynamic runtime compilation for TensorRT (with caching and int8 calibration if you want it), OpenVINO, etc acceleration. You can run any LLM (or multiple), Whisper, Stable Diffusion, sentence embeddings, image classification, and literally any model on the same instance (or whatever) because at the fundamental level Triton was designed for multiple backends, multiple models, and multiple versions. It operates on a in/out concept with tensors or arbitrary data. Which can be combined with the Python and model ensemble support to do anything...
- Python backend. Triton can do pre/post-processing in the framework for things like tokenizers and decoders. With ensemble you can arbitrary chain together inputs/outputs from any number of models/encoders/decoders/custom pre-processing/post-processing/etc. You can also, of course, build your own backends to do anything you need to do that can't be done with included backends or when performance is critical.
- Extremely fine grained control for dispatching, memory management, scheduling, etc. For example the dynamic batcher can be configured with all kinds of latency guarantees (configured in nanoseconds) to balance request latency vs optimal max batch size while taking into account node GPU+CPU availability across any number of GPUs and/or CPU threads on a per-model basis. It also supports loading of arbitrary models to CPU, which can come in handy for acceleration of models that can run well on unused CPU resources - things like image classification/object detection. With ONNX and OpenVINO it's surprisingly useful. This can be configured on a per model basis, with a variety of scheduling/thread/etc options.
- OpenTelemetry (not that special) and Prometheus metrics. Prometheus will drill down to an absurd level of detail with not only request details but also the hardware itself (including temperature, power, etc).
- Support for Model Navigator[3] and Performance Analyzer[4]. These tools are on a completely different level... They will take any arbitrary model, export it to a package, and allow you to define any number of arbitrary metrics to target a runtime format and model configuration so you can do things like:
- p95 of time to first token: X
- While achieving X RPS
- While keeping power utilization below X
They will take the exported model and dynamically deploy the package to a triton instance running on your actual inference serving hardware, then generate requests to meet your SLAs to come up with the optimal model configuration. You even get exported metrics and pretty reports for every configuration used/attempted. You can take the same exported package, change the SLA params, and it will automatically re-generate the configuration for you.
- Performance on a completely different level. TensorRT-LLM especially is extremely new and very early but already at high scale you can start to see > 10k RPS on a single node.
- gRPC support. Especially when using pre/post processing, ensemble, etc you can configure clients programmatically to use the individual models or the ensemble chain (as one example). This opens up a very wide range of powerful architecture options that simply aren't available anywhere else. gRPC could probably be thought of as AsyncLLMEngine on steroids, it can abstract actual input/output or expose raw in/out so models, tokenizers, decoders, clients, etc can send/receive raw data/numpy/tensors.
- DALI support[5]. Combined with everything above, you can add DALI in the processing chain to do things like take input image/audio/etc, copy to GPU once, GPU accelerate scaling/conversion/resampling/whatever, pipe through whatever you want (all on GPU), and get output back to the network with a single CPU copy for in/out.
vLLM and HF TGI are very cool and I use them in certain cases. The fact you can give them a HF model and they just fire up with a single command and offer good performance is very impressive but there are an untold number of reasons these providers use Triton. It's in a class of its own.
[0] - https://mistral.ai/news/la-plateforme/
[1] - https://www.cloudflare.com/press-releases/2023/cloudflare-po...
[2] - https://www.nvidia.com/en-us/case-studies/amazon-accelerates...
[3] - https://github.com/triton-inference-server/model_navigator
[4] - https://github.com/triton-inference-server/client/blob/main/...
[5] - https://github.com/triton-inference-server/dali_backend
It doesn't even need to be very accurate because my own estimations aren't either :)
[1]: https://msty.app
Is it just ease of use or is there something I’m missing?
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
No direct embedding model can be cross-compatable. (exception: constrastive learning models like CLIP)
Reminds me of how Zoom got it start with the “growth hacking” of the installation. Not enough to keep me from using it, but enough for me to keep from using it for anything serious or secure.
If you can't find in foul play in code you can't prove.