Getting Started with Mistral-7b-Instruct-v0.1
secondstate.io
secondstate.io
No... no, that is not what Wasm means at all.
wasi_nn is just an abstraction layers for wrappers around llama.cpp/gguf.cpp/onyx/others. It is not any more cross-platform than these libraries/apps deployed themselves, or rust apps using bindings for them.
https://github.com/jmorganca/ollama/blob/main/cmd/cmd.go#L95...
The default registry for Ollama is https://ollama.ai/library (I don't know about the API endpoint). I'm not sure if/how you can actually use it with another registry endpoint.
When people say that something "uses Docker" that usually involves utilizing the Docker daemon with it's sandboxing, networking, etc.. Ollama doesn't do or need to do that as it can just do inference natively based on GGUF files (+ the other configuration that comes as part of their model file).
However, as you have noticed, it does borrow heavily from Docker on a conceptual level. E.g. apart from the registry mechanism its Modelfile heavily mimics a Dockerfile.
[0]: https://github.com/jmorganca/ollama/blob/aabd71aede99443d585...
I'm as big of a fan of Rust and WASM as the next person, but throwing around claims like that without benchmarks is one of the quickest ways to get your product dismissed.
There’s hardly any / basically no overhead to running an application in a container if that’s what you mean by overhead. If you mean the image size - well you only need to add the things you need - the problem is people tend to abuse images and install a lot of packages in the final image which absolutely aren’t required.
Rust or golang at the core would be nice as Python can be slow at times. I do wish more folks would give Tauri a go for GUI apps.
Will assess myself but wonder if anyone tried.
and also Baichuan https://www.secondstate.io/articles/baichuan2-13b-chat/
There is an Arabic one called Jais mentioned by Satya a few days ago on the Microsoft dev day would like to try out
In my experience most open source models do quite well with Python and Javascript, but don't perform so well when it comes to other languages and the usage of their special characteristics (e.g. it will still be able to write simple control flow in Rust, but probably not use async/await correctly).
I've tried using LLMs in general with IaC (K8s resources/Helm charts), and they all did well when asking it a very specific thing about it (e.g. "Is there an alternative way to accomplish X?"), but also never had it perform well when outputting YAML directly (either from scratch of modifying it).
Give zephyr a try (available in ollama and similar places)
It's a fine tune of mistral and works quite well.
But as you point out, these models have less general knowledge compared to their massive siblings. Knowledge based queries are going to be lower quality than using gpt-4.
Pretty good at instructions
Uncensored, no RHLF bs
Anyway, to answer your question: https://ollama.ai/
I read somewhere that Mistral 7B had a similar performance to GPT-3, but it seems to be miles behind it unfortunately.
To be fair to the developers of Mistral, it's still amazing that I can run this on my computer, and it's certainly better than the first LLaMA.