LM Studio – Discover, download, and run local LLMs
lmstudio.ai
lmstudio.ai
If you're looking to do the same with open source code, you could likely run Ollama and a UI.
https://github.com/jmorganca/ollama + https://github.com/ollama-webui/ollama-webui
It looks like a lot of the tooling is heavily engineered for a set of modern popular LLM-esque models. And looks like llama.cpp also supports LoRA models, so I'd assume there is a way to engineer a pipeline from LoRA to llama.cpp deployments, which probably covers quite a broad set of possibilities.
Beyond llama.cpp, can someone point me to what the broader community uses for general PyTorch model deployments?
I haven't quite ever self-hosted models, and am really keen to do one. Ideally, I am looking for something that stays close to the PyTorch core, and therefore allows me the flexibility to take any nn.Module to production.
[1]: https://github.com/jmorganca/ollama/blob/main/docs/import.md.
As far as I know, ollama doesn’t support exllama, qlora fine tuning, multi-GPU, etc. Text-generation-webui might seem like a science project, but it’s leagues ahead (Like 2-4x faster inference with the right plugins) of everything else. Also has a nice openai mock API that works great.
ooba's worth keeping an eye on, but koboldcpp is more stable, almost as versatile, and way less frustrating. It also still supports GGML.
https://github.com/oobabooga/text-generation-webui/issues/41...
- A local model runtime
- A model catalog
- A UI to chat with the models easily
- An openAI compatible API
And it has several plugins such as for RAG (using ChromaDB) and others.
Personally I think the positioning is very interesting. They're well positioned to take advantage of new capabilities in the OS ecosystem.
It's still unfortunate that it is not itself open-source.
Personally I use a locally served frontend to use ChatGPT via API.
Feature set seems like a decent amount of overlap. One limitation of FastChat, as far as I can tell, is that one is limited to the models that FastChat supports (though I think it would be minor to modify it to support arbitrary models?)
The other is some python thing that requires python knowledge of creating virtual environments, installing dependencies, etc...
Want to try uncensored models.
I have a question, looking for the most popular "uncensored" model I just find "TheBloke/Luna-AI-Llama2-Uncensored-GGML", but it has 14 files to download between 2 to 7 GB, I just download the first one: https://imgur.com/a/DE2byOB
I try the model and it works: https://imgur.com/a/2vtPcui
I should download all the 14 files to get better results?
Also, asking how to make a bomb it looks that at least this model isn't "uncesored": https://imgur.com/a/iYz7VYQ
teknium/openhermes-2.5-mistral-7b is a good one
You don't need all 14 files, just pick one that is recommended with a slight loss of quality - hover over the little (i)icon to find out
Only need 1 of the files, but recommend checking out the GGUF version of the model (just replace GGML in the URL) instead of GGML. Llama.cpp no longer supports GGML, and not sure if TheBloke still uploads new GGML versions of models.
Again, not questioning your motives or anything, just straight up curious. To use your example, any of us can find bomb building info online fairly easily, and has been a point of social contention since the Anarchist's cookbook. Nobody needs an uncensored LLM for that, of course.
Earlier I was looking for “a phrase that is used as an insult for someone who writes with too much rambling” and all I got was some bullshit about how it’s sorry but it can’t do that because it’s allegedly against its OpenAI rules.
So I asked again “a phrase negatively used to mean someone that writes too much while rambling” and it worked.
I simply cannot be bothered to deal with stupid insipid “corporate friendly language” and other dumb restrictions.
Imagine having a real conversation with someone and they freaked out any time anything negative was discussed?
TLDR: Thought police ruining LLMs
I'm not using models censored or otherwise, but I imagine I'd feel the same way there - I won't be offended, so don't try to be too clever, just give me the unfiltered results and let me decide what's correct.
(Bing AI actually banned me for trying to generate images a rough likeness of myself by combining traits of minor celebrities - in combination it shouldn't have looked like those people either, so I don't think should violate ToS, certainly it didn't in intention (I wanted myself, but it doesn't know what I look like and I couldn't provide a photo at the time, idk if you can now (banned!)) so it does happen, 'false positive censoring' if you like.)
i've been trying to see how much i can get away with before it suspends or bans me using one of my throwaway accounts. it takes some doing, but if you convince it you aren't doing something shady right before doing something shady it'll play along with you a fair bit more than if you just say "hi bing, i wanna do some shady shit." unfortunately you have to engineer your "i'm not gonna do (insert something shady)" prompt on a per shade basis.
It doesn't make sense for me personally. It does make sense if you're offering an LLM publicly, so that you doesn't get bad PR if your LLM says some politically incorrect or questionable things.
Unless we are talking about gods and not flawed humans like me I prefer to have the say in what is right and what is wrong for things that are entirely personal and only affect me.
[1] https://www.brookings.edu/articles/the-politics-of-ai-chatgp...
When you ask a local LLM, at worst you get no useful info. When you ask online, at worst you spend the rest of your life in a government black site without any chance of due process.
- The chatbox field has a normal "write here" state, when no chat is really selected. I thought my keyboard broke until I discovered that
- I didn't find a way to set cuda acceleration before loading a model, only managed to set gpu offloaded layers and using "relaunch to apply"
- Some HugginFace models are simply not listed and there's no indication about why. I guess models are really curated, but somehow presented as a HuggingFace browser?
- Scrolling in the accordion parts of the interface seems to be responding to mouse wheel scroll only. I have a mouse with a damaged one and couldn't find a way to reliably navigate to bottom drawers
That said, I really liked the server tab, which allowed for initial debugging very easily
I did find it quite useful for opening a socket for remote tooling (played around withe the continue plugin).
The quirky UI did slow me down,but nothing really showstopping
1. People who understand LLMs and know how to run them and have access to run them on the cloud. 2. People who understand LLMs well enough but don’t have access to cloud resources - but still have a decent MacBook Pro. Or maybe access to cloud resources is done via overly tight pipelines. 3. People who are interested in LLMs but don’t have enough technical chops/time to get things going with Llama CPP. 4. People who are fans of LLMs but can’t even install stuff in their computer.
This is clearly for #3 and it works well for that group of people. It could also be for #2 when they don’t want to spin up their own front end.
The fact it has configuration is good, as long as it has some defaults.
I'm more than capable of compiling/installing/running pretty much any software, but all I want is the ability to chat with a LLM of my choice without spending an afternoon tabbing back to a 30 step esoteric GitHub .md full of caveats, assumptions, and requiring dependencies to be installed and configured according to preferences I don't have.
Being able to swap out models is also handy. This probably saved a couple of hours of my life, which I appreciate.
Tl;DR It's for Mac users
“Deep understanding of what is a computer, what is computer software, and how the two relate.”
Right after the senior ML role that requires people understand how to write “algorithms and programs.”
Kinda hard to take those kinds of requirements seriously.
They seem to have even lowered expectations a bit. Two months ago [1] they were already hiring for that role (or a very, very similar one), but back then you needed experience with "mission-critical code in C++17", now just "production code in C++14".
1: http://web.archive.org/web/20230922170941/https://lmstudio.a...
I wouldn't put C++ devs on too high of a pedestal. I got away with writing shitty C++ code for years before I really knew what I was doing. It still worked though.
Seems like a joke, but many developers do not really understand what's going on behind the scenes. This gets straight to the point. They don't care about HR keyword matching on your CV, or how many years of experience you have of being a mediocre developer with language X or framework Y. I guess during the interview they will investigate whether you truly understand the fundamentals.
Do you have recommendations about this? or blog posts to get started? What would be a decent hardware configuration?
Together with ollama-webui, it can replace ChatGPT 3.5 for most tasks. I also use it in VSCode and nvim with plugins, works great!
I have been meaning to write a short blog post about my setup...
These are heavily optimized for more efficient memory usage, performance, and responsiveness when serving large numbers of concurrent requests/users in addition to things like model versioning/hot load/reload/etc, Prometheus metrics, things like that.
One major difference is at this level a lot of the more aggressive memory optimization techniques and support for CPU aren't even considered. Generally speaking you get GPTQ and possibly AWQ quantization + their optimizations + CUDA only. Their target users and their use cases are often using A100/H100 and just trying to need fewer of them. Support for lower VRAM cards, older CUDA compute architectures, etc come secondary to that (for the most part).
The bad news is because "low VRAM cards" like the 24GB RTX 3090 and RTX 4090 aren't really targetted by these frameworks you'll eventually run into "Yeah you're going to need more VRAM for that model/configuration. That's just how it is." as opposed to some of the approaches for local/single session serving that emphasize memory optimization first and tokens/s for a single session next. Often with no consideration or support at all for multiple simultaneous sessions.
It's certainly possible that with time these serving frameworks will deploy more optimizations and strategies for low VRAM cards but if you look at timelines to even implement quantization support (as one example) it's definitely an after-thought and typically only implemented when it aligns with the overall "more tokens for more users across more sessions on the same hardware" goals.
Loading a 70B model on CPU and getting 3 tokens/s (or whatever) is basically seen as an interesting yet completely impractical and irrelevant curiosity to these projects.
In the end "the right tool for the job" always applies.
The setup of connecting to Ollama is a bit clunky, but once it's set up it works well!
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
Have a look at the localllama subreddit
In short though dual 3090 is common, single 4090 or various flavours of M123 macs. Alternatively p40 can be jury-rigged too but research that carefully. In fact anything with more than one gpu is going to require careful research
Like telling someone interested in 3D printing minis to build a 3D printer instead of buying one. Obviously that helps them get to their goal of printing minis faster right?
But yes, TheBloke tends to have conversions up very quickly as well and has made a name for himself for doing this (+more)
So learn to cook.
Killed the LM studio process and re-opened it and the ghost background usage is down to about 5%.
The demo on this video is from Intel Mac https://youtu.be/C0GmAmyhVxM?si=puTCpGWButsNvKA5
It also supports openai compatible api and completely open-source unlike LM studio
As a contrived example, what happens if you feed the LoTR books, the Hobbit, the Silmarillion, and whatever else is germane, into an LLM?
Is there a base, empty, “ignorant” LLM that is used as a seed?
Do you end up with a Middle Earth savant?
Just how does all this work?
Such a model (LLaMa is a good example) is not "ignorant," but rather a generalized model capable of a wide range of language tasks. This base model does not have specialized knowledge in any one area but has a broad understanding based on the diverse training data.
If you were to "feed" Tolkien's specific books into this general LLM, the model wouldn't become a Middle Earth savant. It would still provide responses based on its broad training. It might generate text that reflects the style or themes of Tolkien's work if it has learned this from the broader training data, but its responses would be based on patterns learned from the entire dataset, not just those books.
https://magazine.sebastianraschka.com/p/understanding-large-...
https://magazine.sebastianraschka.com/p/finetuning-large-lan...
I expected it to not let me run this. I have an intel Macbook, was expecting that I'd need Apple Silicon... am I misunderstanding something? I get fairly fast results at the prompt with the default model. How's this thing running with whatever shitty GPU I have in my laptop?
I include a universal binary of llama.cpp's server example to do inference. What's your machine? The lowest spec I've heard it running on is a 2017 iMac with 8GB RAM (~5.5 tokens/s). On my m1 with 64GB RAM I get ~30 tokens per second on the default 7B model.
I'm now delving into getting this running in Terminal... there are a few things I want to try that I don't think the simple interface allows.
Also, I've noticed that when chats get a few kilobytes long, it just seizes up and can't go further. I complained to it, it spent a sentence apologizing, started up where it left off... and got about 12 words further.
Thanks for trying it!
I have an HP z440 with an E5-1630 v4 and 64GB DDR4 quad channel RAM.
I run LLMs on my CPU, and the 7 billion parameter models spit out text faster than I can read it.
I wish it supported LMMs (multi modal models.)
2. Llama 2
3. Code Llama
4. Orca Mini
5. Vicuna
What can I do with any of these models that won't result in 50% hallucinations/it recommending code with APIs that don't exist/it recommending basically regurgitated StackOverflow historical out of date answers (that it was trained on) for libraries that have had their versions/APIs change, etc?
Can somebody share one real use case they are using any of these models for?
That's why I wanted to try to understand, what am I missing about local-toy LLMs. How are they not just noise/nonsense generators?
They're bad at generative tasks. Don't have it write code or scientific papers from scratch, but you can have it review anything you've written. You can also do summaries, keyword/entity extraction, and the like safely. Any reductive task works pretty well.
On a whim, I asked Zephyr 7B (Mistral based) “what’s the name of that Ruby method that does <insert code>” and it gave me 3 different correct ways of doing what I wanted, including the one I couldn’t remember. That was a real “oh wow” moment.
So offline situations is the most likely use case for me.
Privacy.
If you use llm's on private documents you don't want others to see, you'll likely prefer local models strongly.
Local might still be there after an online service is no longer available.
There are many times when I am searching for a solution to a problem, and I would be perfectly happy with a possible answer I could test that has a 50% chance of being correct.
Humans should test the output of all AIs.
Humans should test the validity of everything we hear, see, and read.
I will buy this with so much enthusiasm if it holds up. Argh, this has been such a pain point.
I hope it supports my (3060) GPU, though.
And when I do try them, it's with Little Snitch blocking outgoing connections.
EDIT: The problem is I'm on macOS 13.2 (Ventura). According to a message in Discord, the minimum version for some (most?) models is 13.6.
https://github.com/nlpxucan/WizardLM/tree/main/WizardCoder
https://huggingface.co/WizardLM/WizardCoder-Python-34B-V1.0
Top of the line consumer machines can run this at a good clip, though most machines will need to use a quantized model (ExLlamaV2 is quite fast). I found a model for that as well, though I haven't used it myself:
https://huggingface.co/oobabooga/CodeBooga-34B-v0.1-EXL2-4.2...
GPT-4 has an estimated 1.8 trillion parameters. Orders of magnitude beyond open source models and ~10x GPT-3.5 which has 175 billion parameters.
https://the-decoder.com/gpt-4-architecture-datasets-costs-an...
The decent chat ones are based on gpt data and they’re basically shitty distilled models.
The best use case is a narrow one that you decide and can create adequate fine-tuning data around. Plenty of real production ability here.