HNHacker News
TopNewBestAskShowJobs

jmorgan

1,780 karma · joined January 31, 2014

https://github.com/ollama/ollama
submissionscomments
jmorgan··on Ollama now supports AMD graphics cards
Ah, this is probably from missing ROCm libraries. The dynamic libraries are available as one of the release assets (warning: it's about 4GB expanded) https://github.com/ollama/ollama/releases/tag/v0.1.29 – dropping them in the same directory as the `ollama` binary should work.
jmorgan··on Ollama now supports AMD graphics cards
The compatibility matrix is quite complex for both AMD and NVIDIA graphics cards, and completely agree: there is a lot of work to do, but the hope is to gracefully fall back to older cards.. they still speed up inference quite a bit when they do work!
jmorgan··on Gemma: New Open Models
Hi! This is such an exciting release. Congratulations!

I work on Ollama and used the provided GGUF files to quantize the model. As mentioned by a few people here, the 4-bit integer quantized models (which Ollama defaults to) seem to have strange output with non-existent words and funny use of whitespace.

Do you have a link /reference as to how the models were converted to GGUF format? And is it expected that quantizing the models might cause this issue?

Thanks so much!

jmorgan··on Ollama is now available on Windows in preview
Ollama should support anything CUDA compute capability 5+ (P3000 is 6.1) https://developer.nvidia.com/cuda-gpus. Possible to shoot me an email? (in my HN bio). The `server` logs should have information regarding GPU detection in the first 10-20 lines or so that can help debug. Sorry!
jmorgan··on Ollama is now available on Windows in preview
This is a really interesting question. I think there's definitely a world for both deployment models. Maybe a good analogy is database engines: both SQLite (a library) and Postgres (a long-running service) have widespread use cases with tradeoffs.
jmorgan··on Ollama is now available on Windows in preview
AMD GPU support is definitely an important part of the project roadmap (sorry this isn't better published in a ROADMAP.md or similar for the project – will do that soon).

A few of the maintainers of the project are from the Toronto area, the original home of ATI technologies [1], and so we personally want to see Ollama work well on AMD GPUs :).

One of the test machines we use to work on AMD support for Ollama is running a Radeon RX 7900XT, and it's quite fast. Definitely comparable to a high-end GeForce 40 series GPU.

[1]: https://en.wikipedia.org/wiki/ATI_Technologies

jmorgan··on Ollama is now available on Windows in preview
Indeed, WSL has surprisingly good GPU passthrough and AVX instruction support, which makes running models fast albeit the virtualization layer. WSL comes with it's own setup steps and performance considerations (not to mention quite a few folks are still using WSL 1 in their workflow), and so a lot of folks asked for a pre-built Windows version that runs natively!
jmorgan··on Ollama releases Python and JavaScript Libraries
Once you have a custom `client` you can use it in place of `ollama`. For example:

  client = Client(host='http://my.ollama.host:11434')
  response = client.chat(model='llama2', messages=[...])
jmorgan··on Ollama releases Python and JavaScript Libraries
Persistent model loading will be possible with: https://github.com/ollama/ollama/pull/2146 – sorry it isn't yet! More to come on filesize and API improvements
jmorgan··on Ollama releases Python and JavaScript Libraries
Sorry this isn't easier!

You can enable mlock manually in the /api/generate and /api/chat endpoints by specifying the "use_mlock" option:

{“options”: {“use_mlock”: true}}

Many other sever configurations are also available there: https://github.com/ollama/ollama/blob/main/docs/api.md#reque...

jmorgan··on The Orange Pi 5 Plus
Next upcoming Ollama version will support non-AVX CPUs
jmorgan··on Many options for running Mistral models in your terminal using LLM
Wow, as an author of the project I'm so sorry about you having to restart your computer. The memory management in Ollama needs a lot of improvement – will be working on this a bunch going forward. I also have a M1 32GB Mac and it's unfortunately just below the amount of memory Mixtral needs to run well (for now!)
jmorgan··on Llamafile lets you distribute and run LLMs with a single file
This is a great point. Context size has a large impact on memory requirements and Ollama should take this into account (something to work on :)
jmorgan··on Yi-34B-Chat
There are other options! Here's a few:

ollama run yi:34b-chat-q2_K # 2-bit ollama run yi:34b-chat-q4_0 # 4-bit ollama run yi:34b-chat-q8_0 # 8-bit

jmorgan··on Yi-34B-Chat
The 6B model is unfortunately still a base text completion model. I've been waiting for the Chat version it to be open-sourced :). The 01-ai team is working on it! https://github.com/01-ai/Yi/issues/173
jmorgan··on Mistral releases ‘unmoderated’ chatbot via torrent
https://ollama.ai/library/mistral should be the instruct/uncensored model, let me know if it that doesn't seem to be the case! You may have to run `ollama pull mistral` to download the latest version
jmorgan··on Mistral releases ‘unmoderated’ chatbot via torrent
It's possible we had tagged `mistral` to be the original model until they had released the instruct (uncensored) model. Re-running `ollama pull mistral` and `ollama run mistral` should respond much differently than above now as it's been updated to default to the latter :-)
jmorgan··on Ollama for Linux – Run LLMs on Linux with GPU Acceleration
There's a ton of cool opportunity in the runtime layer. I've been keeping my eye on the compiler-based approaches. From what I've gathered many of the larger "production" inference tools use compilers:

- https://github.com/openai/triton

- https://github.com/NVIDIA/TensorRT

TVM and other compiler-based approaches seem to really perform really well and make supporting different backends really easy. A good friend who's been in this space for a while told me llama.cpp is sort of a "hand crafted" version of what these compilers could output, which I think speaks to the craftmanship Georgi and the ggml team have put into llama.cpp, but also the opportunity to "compile" versions of llama.cpp for other model architectures or platforms.

jmorgan··on Ollama for Linux – Run LLMs on Linux with GPU Acceleration
Thanks Dang and sorry!!
jmorgan··on LLM Falcon 180B Needs 720GB RAM to Run
A 4-bit quantized version will require (still a whopping) ~100GB of RAM using tools like Ollama and Llama.cpp

If you want to try it with Ollama on macOS (keeping in mind you'll need the new Mac Studio with 192GB of memory) this command will work:

  ollama run falcon:180b
Right now it gets about ~5 t/s.. there's quite a bit of work being done to improve inference speeds like speculative sampling which uses a smaller model for a subset of tokens
jmorgan··on Wasmer – Run, Publish and Deploy any code, anywhere
Thanks! This is really helpful.
jmorgan··on Wasmer – Run, Publish and Deploy any code, anywhere
Could anyone speak to the core differences between Web Assembly (Wasm) and WebAssembly System Interface (WASI)? I noticed Go 1.21 supports WASI as a target, and given how new that is, I haven't seen many example projects that use it yet. Hoping it could make deploying to places like Cloudflare workers more feasible for Go projects now.
jmorgan··on Run LLMs at home, BitTorrent‑style
This is neat. Model weights are split into their layers and distributed across several machines who then report themselves in a big hash table when they are ready to perform inference or fine tuning "as a team" over their subset of the layers.

It's early but I've been working on hosting model weights in a Docker registry for https://github.com/jmorganca/ollama. Mainly for the content addressability (Ollama will verify the correct weights are downloaded every time) and ultimately weights can be fetched by their content instead of by their name or url (which may change!). Perhaps a good next step might be to split the models by layers and store each layer independently for use cases like this (or even just for downloading + running larger models over several "local" machines).

jmorgan··on Exllamav2: Inference library for running LLMs locally on consumer-class GPUs
That's fast. It's exciting to see more ways to run these models locally. How does this compare to llama.cpp – both in speed and approach?
jmorgan··on Asking 60 LLMs a set of 20 questions
I'd be interested to see how models behave at different parameter sizes or quantization levels locally with the Ollama integration. For anyone trying promptfoo's local model Ollama provider, Ollama can be found at https://github.com/jmorganca/ollama

From some early poking around with a basic coding question using Code Llama locally (`ollama:codellama:7b` `ollama:codellama:13b` etc in promptfoo) it seems like quantization has little effect on the output, but changing the parameter count has pretty dramatic effects. This is quite interesting since the 8-bit quantized 7b model is about the same size as a 4-bit 13b model. Perhaps this is just one test though – will be trying this with more tests!

jmorgan··on Asking 60 LLMs a set of 20 questions
:-) that's awesome. Thanks! Nice work on this.
jmorgan··on Asking 60 LLMs a set of 20 questions
This is very cool. Sorry if I missed it (poked around the site and your GitHub repo), but is the script available anywhere for others to run?

Would love to publish results of running this against a series of ~10-20 open-source models with different quantization levels using Ollama and a 192GB M2 Ultra Mac Studio: https://github.com/jmorganca/ollama#model-library

jmorgan··on Llama 2 on togetherAI is as bad of a privacy nightmare as OpenAI
Q4_0 is often mentioned as being the "tried and true" quantization level to try first. I've heard folks have had good results with 3-bit quantization (Q3_K_M) as well
jmorgan··on Run ChatGPT-like LLMs on your laptop in 3 lines of code
I work on Ollama. It's a good question since there are quite a few tools emerging in this space.

The focus for Ollama is to make downloading and serving a model easy – there's an included `ollama` CLI but it's all powered by a REST API. Hopefully, it's a way to support really cool applications of LLMs like OP's onprem tool.

OP's tool is more focused on ingesting and analyzing data. There seems to be quite a bit of interesting opportunity as an application of LLMs – e.g. analyzing not only local docs but data in a remote data store.

jmorgan··on Run ChatGPT-like LLMs on your laptop in 3 lines of code
Love how simple of an interface this has. Local LLM tooling can be super daunting, but reducing it to a simple ingest() and then prompt() is really neat.

By chance, have you checked out Ollama (https://github.com/jmorganca/ollama) as a way to run the models like Llama 2 under the hood?

One of the goals of the project is to make it easy to download and run GPU-accelerated models, ideally with everything pre-compiled so it's easy to get up and running. It's API that can be used by tools like this – would love to know if it would be helpful (or not!)

There's a LangChain model integration for it and a PrivateGPT example as well that might be a good pointer on using the LangChain integration: https://github.com/jmorganca/ollama/tree/main/examples/priva.... There's also a LangChain PR open to add support for generating embeddings, although there's a bit more work to do to support the major embedding models.

Best of luck with the project!

← PreviousPage 2 of 5Next →