Nitro: A fast, lightweight inference server with OpenAI-Compatible API
nitro.jan.ai
nitro.jan.ai
Side-note: I really don't get the C++ in AI thing at all. It's becoming a meme. I wonder why they didn't go with Rust instead. It'll be easier to deploy to the edge as WASM, too.
[1]: https://github.com/ggerganov/llama.cpp/graphs/contributors
Rust is really interesting for stuff besides running the actual llm though. Python can be a huge performance/debugging pain when your frontend and such get huge.
That's the reason why llama.cpp is so attractive though. It just runs the LLM.
easiest : ollama
best for servering : vLLM
OpenAI Compatible : LocalAI
Thanks, I was looking exactly for this!
Llamafile doesnt support an OpenAI compatible embedding endpoint yet but I filed an issue and they seem to be interested in adding it.
To this end, I'm working on a simple library (that sadly doesn't have a shiny website like this post) and will make it public soon.
- OpenAI APIs expose the extra features (like llama.cpp's grammar) anyway, as extra parameters.
- There's a huge productivity benefit to making your backend swappable, without having to redo all the API calls.
Prompt formatting in particular is a huge sticker with OpenAI endpoints though. Something does need to be done about that.
I wanted to be able to call llama.cpp from Python, but I didn't want to use the llama-cpp-python wrapper because it automatically downloads and builds llama.cpp, which I didn't want to do. I like the simplicity of llama.cpp and its unix philosophy of doing one thing well, so I prefer to build llama.cpp myself and then call it from Python.
As in, after you've installed the package, it will pause your application at runtime to download and build llama.cpp? That's not how bindings are supposed to work! That sounds seriously cursed.
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
...Unfortunately they have issues. The C++ version straight up ignores parameters like temperature, the python implementation does not support batching.
Obviously only relevant for non-inference API calls.
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
...But still, these are not very common API calls? Generally an OpenAI endpoint is mostly inference calls, right? And llama.cpp's slowness is going to blow that advantage away.
(Someone correct me if I'm wrong, I'm only familiar with building llama.cpp, can it run from just the binary without cuda?)
Fully possible the post was done in a hurry and you didn't think about it prior, but it's looking hand wavy which is not a good signal for experienced candidates you're trying to attract.
Because if you want to comply with AGPL then you can't build a product around this unless you negotiate a different license. Unless you want to release all of your own code. Which might be possible also.
But I would go with llama.cpp, ollama, candle (Rust), or whatever. Non of those are copyleft.
...So I would not recommend a batched llama.cpp server to anyone, TBH, when there are a grab back of splended OpenAI endpoints like TabbyAPI, Aphrodite, LiteLLM, VLLM...
It’s a basic value prop and I’m way behind the times on local inference (so barely a few months lol), but it seems different enough that I’m excited to try it. Other tools have tried this path but I’ve seen none with a demo page quite this convincing…
Not that it isn't a nice idea -- making it easy to test existing OpenAI-based apps against competing open models is a pretty good thing.
I see the value in being able to locally run apps designed to talk to OpenAI apis.
I'm all for cool innovations, but... I'm not really seeing the advantage here beyond the shiny website to attract VCs.
It's not a model, it puts an API on local models.
I found that to be incredibly impractical. I tried to do it for my project AIMD, but the cost and quality just made absolutely no sense even with the top models.
I definitely see your specific point tho, and have found the same for high-level usecases. Local models become really useful when you need smaller models for ensemble systems, to give one class of use case you might want to try out —- e.g. proofreading, simple summarization, tone detection, etc.
> Built on top of the cutting-edge inference library llama.cpp, modified to be production ready.
It's not. It's literally just llama.cpp -> https://github.com/janhq/nitro/blob/main/.gitmodules
Llama.cpp makes no pretense at being a robust safe network ready library; it's a high performance library.
You've made no changes to llama.cpp here; you're just calling the llama.cpp API directly from your drogon app.
Hm.
...
Look... that's interesting, but, honestly, I know there's this wave of "C++ is back!" stuff going on, but building network applications in C++ is very tricky to do right, and while this is cool, I'm not sure 'llama.cpp is in c++ because it needs to be fast' is a good reason to go 'so lets build a network server in c++ too!'.
I mean, I guess you could argue that since llama.cpp is a C++ application, it's fair for them to offer their own server example with an openai compatible API (which you can read about here: https://github.com/ggerganov/llama.cpp/issues/4216, https://github.com/ggerganov/llama.cpp/blob/master/examples/...).
...but a production ready application?
I wrote a rust binding to llama.cpp and my conclusion was that llama.cpp is pretty bleeding edge software, and bluntly, you should process isolate it from anything you really care about, if you want to avoid undefined behavior after long running inference sequences; because it updates very often, and often breaks. Those breaks are usually UB. It does not have a 'stable' version.
Further more, when you run large models and run out of memory, C++ applications are notoriously unreliable in their 'handle OOM' behaviour.
Soo.... I know there's something fun here, but really... unless you had a really really compelling reason to need to write your server software in c++ (and I see no compelling reason here), I'm curious why you would?
It seems enormously risky.
The quality of this code is 'fun', not 'production ready'.
I hope this is a typo/mistake in the docs, because if not, that's a terrible idea. Nitro cannot serve GPT-3.5, so it should absolutely not pretend to. "Drop-in replacement" doesn't mean lying about capabilities. If that model is specifically requested, Nitro should return an error, not silently do something else instead.
But... I do agree, this should be feature-gated behavior
Why?
Your mindset would mean that Windows would have next to no backwards compatibility, for instance.
The most important factor is inference speed. For something called Nitro, I really expected speed benchmarks. I'd be interested in CPU, CUDA, and MPS at different batch sizes.
Most llama.cpp openai servers are pretty close to vanilla llama.cpp, albeit without the batching support.
Imagine I want to test something fully offline, but my app was built using OpenAI's API's ideally I can do a quick swap now and have an offline compatible model.
Additionally, let's say I want to offer a free version of AI to my non-paying users, but for my paying users I can afford to allow them to use OpenAI, this will allow for an easier swap without having to rewrite everything.
There are many of us who are running solo operations and the only alternative to quick a/b testing in the broader scheme of your entire application follow is to either write your own wrappers or use some humongous horrendous abstraction layer like langchain, which is going to cost you 2x time in the future.
I’ve added support for multiple apis in my app but it’s definitely a non trivial task to test new ones, especially with so many on the market.
@Linear, you should give that designer a cool 50 g's bonus.