Ollama v0.1.33 with Llama 3, Phi 3, and Qwen 110B
github.com
github.com
[1] https://devboard.gitsense.com/ggerganov?r=ggerganov%2Fllama....
Full disclosure: This is my tool
The chance onnx becomes significantly relevant here went from 1% to 15% this week. They're demo'ing ~2x faster inference with Phi-3. There's been fits and starts on LLMs in ONNX for a year, but, with Wintel's AI PC™ push, and all the constituent parts in place (4 bit quants! adaptive quants!), I'd put very good money on it.
Microsoft isn't going to pay for something that amounts to a useful setup script wrapped around an inefficient convenience library intended for people to be able to run AI on consumer hardware. There's no exploitable value proposition, whereas building their own closed source AI systems that are tightly coupled to the Windows ecosystem and favor cloud services allows them to extract maximum rent.
All of this of course acknowledging that llama.cpp is an incredible project with competitive performance and support for almost any platform.
[1] https://github.com/ml-explore/mlx
Edit: this would be useful because in many cases some workloads can be local, but others cannot... e.g. if you really need gpt4 for specific queries.
The examples show image input though: https://github.com/ollama/ollama/blob/main/docs/api.md#reque...
Maybe you can file an issue here: https://github.com/ollama/ollama/issues
You can use Llama3 70B with aider via Ollama [0]. It's also available for free via Groq [1] (with rate limits). And OpenRouter has it available [2] for low cost on their paid api.
[0] https://aider.chat/docs/llms.html#ollama
For those looking to create their own benchmarks, promptfoo[0] is one way to do this locally:
prompts:
- "Write this in Python 3: {{ask}}"
providers:
- ollama:chat:llama3:8b
- ollama:chat:phi3
- ollama:chat:qwen:7b
tests:
- vars:
ask: a function to determine if a number is prime
- vars:
ask: a function to split a restaurant bill given individual contributions and shared items
Jumping in because I'm a big believer in (1) local LLMs, and (2) evals specific to individual use cases.Looks like its a vector DB used for creating and looking up embeddings (vectors). LLM is the second part of RAG, the first part is having a good embedding model.
Jina AI is doing different things but one of them is having embeddings and I use their English/ German embeddings as in one demo I am working with German data.
You can use pip as well but yes, let me add something about Poetry in case people don't know about it :)
ex. I could condition on "\n\n<|assistant|>||<|system|>||<|user>", but it'd still be wrong.
Pretty much everything Phi 3 feels like it needed to all come out within 48 hours a month too early. The ONNX genai library doesn't work on Mac, at all, the mobile SDKs don't support it...sigh
I think the articles are just not really upvoted unless it's really big news, makes sense because HN is for more than just AI.
But I don't think it's anti-AI like most people here would be pretty anti-cryptocurrency (and for good reason IMO)
Ollama allows you to run those models.
Different things.
Or do you mean commercial deployment of models for inference?
My understanding is the modern quantization algorithms are typically implemented in Pytorch.
The only thing I know (from using it) that with quantization I can fit models like llama2 13b, in my 24GB of VRAM when I use q8 (16GB) instead of fp16 (26GB). This means I can get nearly the full quality of llama2 13b's output while still being able to use only my GPU, without the need to do very slow inference on only CPU+RAM.
And the models are quantized before inference, so I'd only download 16GB for the llama2 13b q8 instead of the full 26GB, which means it's not done on the fly.
I don't know why I still do it, but everytime I read so many comments how good model X is, and how it outperforms anything else, and then I want to see it for myself.
In my home I have a large gaming rig that sometimes runs Ollama+Open WebUI, then I also have a bunch of other services running on a smaller server and a Raspberry Pi which reach out to Ollama for their LLM inference needs.
HF is the biggest provider of llms, and I guess I haven’t run into it’s limitations yet.
Built-in model recommendations are also handy.
Very friendly tool!
However it's not open-source.