HNHacker News
TopNewBestAskShowJobs

jmorgan

1,780 karma · joined January 31, 2014

https://github.com/ollama/ollama
submissionscomments
jmorgan··on Refact Code LLM: 1.6B LLM for code that reaches 32% HumanEval
This is true! Although I'm also really excited at the potential speed (both for loading the model and token generation) of a 1B model for things like code completion.
jmorgan··on ChatGPT Enterprise
That's quite interesting. I hadn't thought of sparsity in the weights as a way to compress models, although this is an obvious opportunity in retrospect! I started doing some digging and found https://github.com/SqueezeAILab/SqueezeLLM, although I'm sure there's newer work on this idea.
jmorgan··on ChatGPT Enterprise
Ah, gotcha! I thought you probably meant something else. I've been wondering this too, and it's something I've been meaning to look at.

On a related note it doesn't seem like many local runners are leveraging techniques like PagedAttention yet (see https://vllm.ai/) which is inspired by operating system memory paging to reduce memory requirements for LLMs.

It's not quite what you mentioned, but it might have a similar effect! Would love to know if you've seen other methods that might help reduce memory requirements.. it's one of the largest resource bottlenecks to running LLMs right now!

jmorgan··on ChatGPT Enterprise
There are different levels of quantization available for different models (if that's what you mean :). E.g. here are the versions available for Llama 2: https://ollama.ai/library/llama2/tags which go down to 2-bit quantization (which surprisingly still happens to work reasonably well).
jmorgan··on ChatGPT Enterprise
I believe it would have also kept its stars, issues and other data.
jmorgan··on ChatGPT Enterprise
Seemed like a great project. Hope to see it come back!

There are some great open-source projects in this space – not quite the same – many are focused on local LLMs like Llama2 or Code Llama which was released last week:

- https://github.com/jmorganca/ollama (download & run LLMs locally - I'm a maintainer)

- https://github.com/simonw/llm (access LLMs from the cli - cloud and local)

- https://github.com/oobabooga/text-generation-webui (a web ui w/ different backends)

- https://github.com/ggerganov/llama.cpp (fast local LLM runner)

- https://github.com/go-skynet/LocalAI (has an openai-compatible api)

jmorgan··on Continue with LocalAI: An alternative to GitHub's Copilot that runs locally
Continue has a great guide on using the new Code Llama model launched by Facebook last week: https://continue.dev/docs/walkthroughs/codellama

Continue also works with various backends and fine-tuned versions of Code Llama. E.g. for a local experience with GPU acceleration on macOS, continue can be used with Ollama (https://github.com/jmorganca/ollama):

  ollama pull codellama

  from continuedev.src.continuedev.libs.llm.ollama import Ollama

  config = ContinueConfig(
    models=Models(
      default=Ollama(model="wizardcoder:34b-python")
    )
  )
jmorgan··on Show HN: Dumbar, a not so smart menubar app
Super cool. I just realized who you were from your handle and wanted say that your work on electron over the years has been absolutely amazing. Thank you for all you've done & built.
jmorgan··on Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
Thanks for creating an issue! And sorry for the error folks… working on it!
jmorgan··on Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
Yes you'll need the latest version (0.0.16) to run the 34B model. It should run great on that machine!

The download (both as a Mac app and standalone binary) is available here: https://github.com/jmorganca/ollama/releases/tag/v0.0.16. And I will work on getting that brew formula updated as well! Sorry to see you hit an error!

jmorgan··on Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
It really is good. Surprisingly it seems to answer instruct-like prompts well! I’ve been using it with Ollama (https://github.com/jmorganca/ollama) with prompts like:

  ollama run phind-codellama "write c code to reverse a linked list"
To run this on an m1 Mac or similar machine, you'll need around 32GB of memory for the 4-bit quantized version since it's a 34B parameter model and is quite big (20GB).
jmorgan··on Code Llama, a state-of-the-art large language model for coding
A few folks and I have been working on an open-source tool that does some of this (and hopefully more soon!) https://github.com/jmorganca/ollama

There's a "PrivateGPT" example in there that is similar to your third point above: https://github.com/jmorganca/ollama/tree/main/examples/priva...

Would love to know your thoughts

jmorgan··on Code Llama, a state-of-the-art large language model for coding
Indeed! After pulling a model with "ollama pull codellama" you can access it via the REST API:

  curl -X POST http://localhost:11434/api/generate -d '{                        
    "model": "codellama",
    "prompt":"write a python script to add two numbers"
  }'
jmorgan··on Code Llama, a state-of-the-art large language model for coding
Sorry, this should be fixed now! To update you'll have to run:

  ollama pull codellama:7b-instruct
jmorgan··on Code Llama, a state-of-the-art large language model for coding
> managed to get infinite streams of near nonsense

This should be fixed now! To update you'll have to run:

  ollama pull codellama:7b-instruct
jmorgan··on Code Llama, a state-of-the-art large language model for coding
To run Code Llama locally, the 7B parameter quantized version can be downloaded and run with the open-source tool Ollama: https://github.com/jmorganca/ollama

   ollama run codellama "write a python function to add two numbers"
More models coming soon (completion, python and more parameter counts)
jmorgan··on SeamlessM4T, a Multimodal AI Model for Speech and Text Translation
I'm curious about this too. Lately I've been building an open source tool to help bring make pulling + running models easier locally – https://github.com/jmorganca/ollama – right now we work with the awesome llama.cpp project, however, other model types have definitely come up. LLMs are a small section of what's available on huggingface for example.

It's especially interesting how you could combine different model types - e.g. translation + text completion (or image generation) – it could be a pretty powerful combination...

jmorgan··on How Is LLaMa.cpp Possible?
This project's been a blast to work with. While it's written in C++, it provides a C interface to compile against which makes it especially easy to extend with Go, Python and other runtimes.

A few folks and I have been building a tool with it in Go for pulling & running multiple models, and serving them on a REST API: https://github.com/jmorganca/ollama

In similar light, you haven't checked it out, llama.cpp also has a pretty extensive "server" tool (in its examples directory in the repo) with a web ui and support for grammar (e.g. forcing the output to be JSON)

jmorgan··on Why host your own LLM?
One benefit of self-hosting LLMs is the wide range of fine-tuned models available, including uncensored models. A popular one over the last weeks was Llama 2 Uncensored by George Sung: https://ollama.ai/blog/run-llama2-uncensored-locally

A few more:

- Wizard Vicuna 13B uncensored

- Nous Hermes Llama 2

- WizardLM Uncensored llama2

jmorgan··on Show HN: AI-town, run your own custom AI world SIM with JavaScript
If you haven't yet checked out the Generative Agents project referenced by OP, definitely give it a look, it's open source: https://github.com/joonspk-research/generative_agents

Over the weekend Lance Martin got it working with local models using llama.cpp and ollama.ai which saves $ on longer sims since all inference happens locally https://twitter.com/RLanceMartin/status/1690829179615657985. It's neat how the AI agents interface with each other – e.g. one will host a party and invites will be sent throughout the group

jmorgan··on Azure ChatGPT: Private and secure ChatGPT for internal enterprise use
RE 2 - neat! What are some tasks you've been using smaller models (with perhaps larger context sizes) for?
jmorgan··on Azure ChatGPT: Private and secure ChatGPT for internal enterprise use
It's early, and this definitely isn't customer facing in the traditional sense, but a team member of mine set up a Discord bot running Llama 2 70B on a Mac studio and we've been quite impressed by its responses to folks who test it.

IIRC chat bots are central the vision Facebook has with LLMs (e.g. every instagram account has a personal chat bot), so I would expect the Llama models to get increasingly better at this task.

That said the 7B and 13B models definitely don't quite seem ready yet for production customer interaction :-)

jmorgan··on Azure ChatGPT: Private and secure ChatGPT for internal enterprise use
This appears to be a web frontend with authentication for Azure's OpenAI API, which is a great choice if you can't use Chat GPT or its API at work.

If you're looking to try the "open" models like Llama 2 (or it's uncensored version Llama 2 Uncensored), check out https://github.com/jmorganca/ollama or some of the lower level runners like llama.cpp (which powers the aforementioned project I'm working on) or Candle, the new project by hugging face.

What's are folks' take on this vs Llama 2, which was recently released by Facebook Research? While I haven't tested it extensively, 70B model is supposed to rival Chat GPT 3.5 in most areas, and there are now some new fine-tuned versions that excel at specific tasks like coding (the 'codeup' model) or the new Wizard Math (https://github.com/nlpxucan/WizardLM) which claims to outperform ChatGPT 3.5 on grade school math problems.

jmorgan··on Beginner's Guide to Llama Models
There is! While not easy to use yet, there's a sort-of-hidden way models can be listed with:

  curl https://ollama.ai/v2/_catalog | jq 
Then to list "tags" for a given model (e.g. llama2):

  curl https://ollama.ai/v2/library/llama2/tags/list | jq
jmorgan··on Beginner's Guide to Llama Models
I really enjoyed Anrej Kaparthy's llama2.c project (https://github.com/karpathy/llama2.c), which runs through creating and running a miniature Llama2 architecture model from scratch.
jmorgan··on Beginner's Guide to Llama Models
An easy way to try many of the fine-tuned Llama 2 models is https://github.com/jmorganca/ollama.

A maintainer of the project has been collecting a full list here (with different quantization levels), most of which are Llama 2-based: https://gist.github.com/mchiang0610/b959e3c189ec1e948e4f6a1f...

Since the release of Llama 2 the number of models based on it has been growing significantly.. some popular ones:

- codeup (A code generation model - DeepSE)

- llama2-uncensored (George Sung)

- nous-hermes-llama2 (Nous Research)

- wizardlm-uncensored (WizardLM)

- stablebeluga (Stability AI)

The article also recommends oobabooga's text-generation-webui which includes a full web dashboard.

jmorgan··on Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching
I do think this is an apt analogy. I've heard a counterpoint that there won't be enough "destinations" for this to work, but then it's not hard to imagine a single order of magnitude more LLM "destinations" than today (all of which were launched in the last 12 months).

There's also the fact that the data being sent to these LLM "destinations" could be significantly more valuable (or contain significantly more sensitive information) than the average segment identity or track objects.

jmorgan··on Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching
Great. Let's chat!
jmorgan··on Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching
The idea of an LLM proxy is super compelling. There's a lot of powerful ideas baked into the proxy form factor – I think you've listed out quite a few of them. It reminds me a bit of what Cloudflare did for the web: both making it faster and safer/easier. Have you considered local LLMs at all for Llama 2? A few people and I have been working on https://github.com/jmorganca/ollama/ and was thinking it would be helpful to be able to augment it with a proxy layer like this. Not only that, but it might help folks dynamically choose to run locally (vs against a cloud LLM) for certain prompts.
jmorgan··on Show HN: Chat with your data using LangChain, Pinecone, and Airbyte
LangChain supports local LLMs like Llama 2 with Ollama (https://github.com/jmorganca/ollama) as of this morning, in both their Python and Javascript versions:

https://python.langchain.com/docs/integrations/llms/ollama

This can be a great option if you'd like to keep your data local versus submitting it to a cloud LLM, with the added benefit of saving costs if you're submitting many questions in a row (e.g. in batches)

← PreviousPage 3 of 5Next →