How I run LLMs locally
abishekmuthian.com
abishekmuthian.com
I like this. If we insist on pushing forward with GenAI we should probably at least make some digital or physical monument like "The Tomb of the Unknown Creator".
Cause they sure as sh*t ain't gettin paid. RIP.
that can and should also be good enough for the AI era.
Megacorp scraping the entire internet vs individuals sharing knowledge on a platform are a completely different scale.
Also, at a social level - the worst kind of user has always been a help vampire, and LLMs are really good at increasing the number of help vampires.
Maybe we should hype up LLMs more so the help vampires and LLMs can keep each other busy?
Your intentions might not have been to make money, but you were creating social credit that could be redeemed for a higher and better paying job. With GitHub, Stack Overflow, etc., you are adding to your resume, but with AI, you literally get nothing in return for contributing.
I pay for o1-pro cheerfully, but I wouldn't pay anything at all for Stack Overflow. ChatGPT certainly generates its share of BS, but I have yet to have a question rejected because somebody who was using a different language or OS asked about something vaguely similar 8 years ago.
Sure, if AI was made free for everybody (or only be charged for cost to run).
With Stack Overflow, GitHub and others, there is a mutual understanding that contributing can benefit the contributor. What is the incentive to continue contributing if the social agreement is, you get to help define a statistical weight for the next token and nobody will know?
I think the future business model may require AI companies to pay people to contribute, or it might not be a technology roadblock, but rather a data roadblock that prevents further advancement.
I think lots of people contribute everywhere without getting any benefit at all.
I don't doubt that some who do contribute hoping for, or expecting, some ancillary benefit.
I'd suggest that pretty much the only tangible benefit I can see is those searching for a job. Contributing in public spaces is a good technique for self-promotion as being skilled in an area.
Then again I'd suggest that the majority of people participating on those sites are already employed, so they're not doing it for that benefit. I'd even argue that their day job accomplishments are likely to be more impressive than their github account when it comes to their next interview.
So perhaps I can reassure you. I'm pretty sure people will continue to absorb information, and skills, and will continue to share that with others. This has been the way for thousands of years. It has survived the inventions of writing, printing, radio, television and the internet. It will survive LLMs.
Guilds used to jealously guard their secrets. Metallurgy techniques were lost when their creators died, or were silenced.
And most crucially - the audience has always been primarily humans. There has never been an audience composition, where authors have to worry about plagiarism as the default.
The idea of free exchange of ideas is something that we enjoyed only recently.
This isn’t naysaying or doom and gloom - this is simply reality. Placing our hopes on the wrong things leads to disappointment, anger and resentment when reality decides our hopes are an insufficient argument to change its ways.
But the company benefits from less confusion and a better user experience. Companies are literally paying employees to provide content as it benefits the company.
I do believe there are people who freely choose to contribute with no strings attached, and I guess we'll learn in the coming years if people will contribute their time and effort for benevolent reasons.
Heck - how many people will go to stack overflow when they can get pseudo good answers from GenAI in the first place?
And stack overflow is filled with bots, and not humans?
Why would they contribute, when every action they take will simply mean OpenAI or someone will benefit, and some random bot will answer?
Signal vs Noise is what the internet is all about. Plastic was a godsend when it was invented. It’s a plague found at the bottom of the Mariana Trench today.
I’m not like anti what’s happening or for it, it’s just, that social credit depends on those institutions surviving.
I guess I never really thought about having a dog in this race (effectively being a tradesman as I am, and not a "content creator"). I did write a ton on ServerFault.com. I guess I am in this unwittingly.
I'm a little salty about the LLM training on my Stack Exchange answers but I knew what I was getting in to when I signed-up. I don't really subscribe to notions of "intellectual property" so I don't feel strongly on that front.
It just feels impolite and rude. More like plagiarism and less like copyright infringement. A matter of tact between people, versus a legal matter.
The way LLMs turn the collective human expression into "slop" that "they" then "speak" with a tone of authority about feels scummy. It feels like a person who has read a few books and picked up the vernacular and idiom of a trade confidently lying about being an expert.
I can't attribute that scumminess to the LLM itself, since it's just a pile of numbers. I absolutely attribute that scumminess to the companies making money from them.
re: Stack Exchange social credit and redeeming it - I'm not a good self-promoter, and admittedly ServerFault.com is a much smaller traffic Stack Exchange site than Stack Overflow, but begin the top-ranked user on the site for 5+ years didn't confer much in the way of real-world benefits. I had a ton of fun though.
(I got a tiny bit of name recognition from some IRL people and a free trip to the Stack Overflow offices in NYC one time. I definitely got a boost of happiness every time a friend related a story to the effect of: "I ran into an issue, search-engined it, and came up with something you wrote on Server Fault that solved my problem.")
so if they weren't making money (or weren't planning on making any), then would it still be "scummy"?
In other words, do you feel that they're only scummy because they're able to profit off the work (where as you didn't or couldn't)? Why isn't this sour grapes?
The reality is many wikipedia, stack overflow, etc contributors want information to be free and correct, and don't want money, so it's not sour grapes, it's rather annoyance at a perversion of the intent and vision.
I contributed to wikipedia because I want a free reservoir of human knowledge to benefit all, I want the commons to be rich with information. Anyone making money off it is scummy not because I couldn't figure out how to, but because they are perverting the intention of information to be free.
Instead, we've ended up with one of the main interfaces to wikipedia being a paid often inaccurate chatbot for a for-profit company which doesn't attribute wikipedia and burns down forests as a side-effect.
This isn't sour grapes, this is recognizing exploitation of the commons.
Equally I'd suggest that the commons is not free. It has to be paid for by someone. Wikipedia exists by begging for donations. Google sells advertising (as does StackOverflow as job listings) etc.
I mean, the first carpenter who took "common knowledge" and wrote a (paid for) book did the same thing. Knowledge is definitely not free, and it costs money to spread it.
(As an aside, I've been using LLMs for free all year.)
All through history people have exploited the commons. The printing press, books, universities, education, radio, television, through computers, Google, sites like SO. LLMs are just the latest step in a long long line of history.
There is no such thing as a free lunch. Over grazing common pasture land, results in its decimation.
There are national and international level bodies required to ensure we dont kill all the rhinos. Hell - that we dont kill all the people.
The printing press, universities, education - these are NOT commons in many places, nor do they function as commons. Let alone function as LLMs.
Common knowledge is not the same as the commons.
does it diminish, if the commons is knowledge based, such as online sources? Those sources does not truly disappear after the information is extracted and placed into an LLM.
Unlike a physical commons, which has limitations on use, informational commons don't.
So the fact that someone else is able to gain more value out of the knowledge than others is not a reason to make them scummy - as if they alone don't deserve access to the knowledge that you claim should be free.
If contributors, after seeing how someone else is able to make profits off previously freely available knowledge, feel that they somehow now suddenly deserve to be paid after the fact, then i dont know how to say it but to call it sour grapes.
My contributions to Stack Overflow have all been done anonymously, and I haven't ever felt even the slightest bit of desire to link that identity to my real one, add it to my resume, brag to friends.
Having and sharing knowledge to me is its own reward, and I have no intention of profiting from it in any way.
Your own personal way of thinking about the world shines through when you search for ulterior motivations and state as fact that those factors, like a desire for fame or money, must be present.
I have spent the last year in a new area (sql) and I've written a lot of questions to LLMs, which it gas been able to answer well enough for me to make speedy progress.
I'm a big fan of StackOverflow, and Google, and before that reference books to gather and learn.
Each technology builds on the layer before. Information is disseminated, repackaged, reauthored.
I get that some people feel like their contribution should be the end of the line. Despite perhaps that they got that knowledge from somewhere. Do they credit their college professor when posting on Reddit?
So again, thank you for your contribution. Your willingness to answer questions, and the answers you provided, will exist long after you and I do not.
I tip my hat.
Given that, as a thought experiment, would you be OK with your answers/comments being attributed to others, or that once you've submitted them, there is no linkage to you at all (e.g. you can't even know what you submitted)? Would it be as satisfying of a process if your contributions are just dissolved into a soup of other data irreversibly?
That doesn't sound like a system I'd be as keen to contribute to. Maybe the ulterior motive is at least being able to find my body of work as a source of personal fulfillment. Where is my work in the various LLMs? I have no idea, and will likely never know.
Yes. Wikipedia _almost_ operates like this. I have no expectations of anyone digging into who wrote what, it turns into a soup of information. I still do know I contributed, but I don't care if what I wrote gets rewritten, replaced, improved.
4chan does operate like this, and back in the days that /prog/ had meaningful discussions, I enjoyed participating in threads there.
So am I.
> I didn't start contributing to make money, just to share information and learn from each other. that can and should also be good enough for the AI era.
Sure; my content was contributed under CC-BY-SA, and if AI honors the rather simple terms of that license, then it's also good enough for the AI era, just as I had the same expectations of human consumers.
I stayed away from using LLMs for a long time but then came a point when it became a disadvantage to not use one, especially as a person belonging to disadvantaged section of the society I'm forced use any & all forms of technology which helps me & others like me to have some equity.
Besides I don't want a land acknowledgement, I want a statue!!
Why should training be subject to copyright? (and at what stages of the process). People learn most of what they know from copyright media, giving copyright owners more control over that might be a bad idea
But AI companies are paying content owners for for access recently (partially to reduce legal risk, partially to get access to material not publicly available) but then giving deep-pocket Incumbent a monopoly might be a bad idea.
Perhaps requiring that deep-pocket companies actually compensate copyright holders would be a starting point for a fairer system?
the other option is the holders "win" and these models must only be trained on owned material, in which case the market will collapse into a handful of models controlled by companies that already own huge swaths of intellectual property. basically think DisneyDiffusion or RandomHouse-LLM. Nobody is getting paid more but it's all above board since it's been trained on all the data they have rights to. You might see some holders benefit if they have a particularly large and useful dataset, like Reddit or the Wall Street Journal.
People with power and money, can get paid. Artists who have no reach and recognition get exploited. Especially those from countries which arent in north America and Europe.
Should we extend that model to text books? If I learn about a topic from a book, can I never write a book of my own on that topic?
Should we extend that model to the web? If I learned CSS and JavaScript reading StackOverflow, am I banned from writing a book, giving classes, or indeed even answering questions on those topics?
I ask this in seriousness. I get that LLM training is new, and it's causing concerns. But those concerns have existed forever - the dissemination of information has been going on a long time.
I'm sure the same moral panic existed the first time someone started making marks in clay to describe what us the best time of year to plant the crops.
Time was if a person read multiple books to write a new text, they either purchased them, or borrowed them from a library which had purchased them, and then acknowledged them in a footnote or reference section.
At least one author noted that there was a concern that writing would lead to a diminishment of human memory/loss of oral tradition (Louis L'Amour in _The Walking Drum_).
Really? I can't tell if you're joking, so I'll take it at face value.
See, I associate the earliest famous (I thought) expression of that concern with Plato, and before today I couldn't remember any other associated details enough to articulate them with confidence. ChatGPT tells me, using the above quote without the citation as a prompt, that it was in Plato in his dialogue Phaedrus, and offers additional succinct contextual information and a better quote from that work. I probably first learned to associate that complaint about writing with Plato in college, and probably got it from C.D.C. Reeve, who was a philosophy professor and expert on Plato at the college I attended. But I feel no need to cite any of Reeve's works when dropping that vague reference. If I were to use any of Reeve's original thoughts related to analysis of Plato, then a reference would be merited.
It seems to me that there are different layers of abstraction of knowledge and memory, and LLMs mostly capture and very effectively synthesize knowledge at layers of abstraction that are above that of grammar checkers and below that of plagiarism in most cases. It's true that it is the nature of many of today's biggest transformers that they do in some cases produce output that qualifies as plagiarism by conventional standards. Every instance of that plagiarism is problematic, and should be a primary focus of innovation going forward. But in this conversation no one seems to acknowledge that the bar has been moved. The machine looked upon the library, and produced some output, therefore we should assume it is all theft? I am not persuaded.
Nb: if you’re still paying $20/mo for a feature-poor chat experience that’s locked to a single provider, you should consider using any of the many wonderful chat clients that take a variety of API keys instead. You might find that your LLM utilization doesn’t quite fit a flat rate model, and that the feature set of the third-party client is comparable (or surpasses) that of the LLM provider’s.
edit: included repo link; note on API keys as alternative to subscription
Just for AI image generation i would rather recommend krita with the https://github.com/Acly/krita-ai-diffusion plugin.
Does anyone know of a lighter/minimalist version?
It's cross platform, works with most LLMs, and has extension support.
Another alternative is LibreChat, it's more fully featured but its heavier weight and spins up quite a few docker instances.
https://github.com/danny-avila/LibreChat
If you're going through the trouble of running models locally, it doesn't make sense to couple that with a proprietary closed source interface (like LM Studio).
Jan and Librechat are both less popular than either open-webui or ollama, how do their features compare?
In the case of AlphaGeometry, I made AlphaGeometryRE to get rid of tensorflow/jax/flax, and 100+ Python packages.
One of the best places to get the latest info on what people are doing with local models is /lmg/ on 4chan's /g/
Edit: outside the HPC community specifically, I mean
You don't want an A100 unless you've already got datacenter provisioning at your house and an empty 1U rack. I genuinely cannot stress this enough - these are datacenter cards for a reason. The best bang-for-your buck will be consumer-grade cards like the 3060 and 3090, as well as the bigger devkits like the Jetson Orin.
https://www.reddit.com/r/LocalLLaMA/
Search for terms like hardware build, running large models, multiple GPU’s, etc. Many people there have multiple, consumer GPU’s. There’s probably something about running multiple A100’s.
HuggingFace might have tutorials, too.
Warning: If it’s A100’s, most people say to just rent them from the cloud as needed cuz they’re highly costly upfront. If they’re normally idle, then it’s not as cost effective to own them. Some were using services like vast.ai to get rentals cheaper.
If you need a guide, the A100 class cards are certainly not what you want. Running data center level clusters in the home is no walk in the park. You wouldn’t even be able to power a decent sized cluster without getting an electrician involved for proper power. Networking is expensive. Even second-hand it all adds up fast. Don’t underestimate how much it would cost.
Getting some 3090s in a case with a big gaming PSU is the way to go for home use. If you really want A100 level, rent it from a cloud provider. Trying to build your own cluster would be like lighting money on fire for a rapidly depreciating asset.
Or rephrasing for somebody that never tried to experiment with models: which one do you need to have something like ChatGPT-4o (text only) at home?
If your model does fit into VRAM, if its getting ejected there will be a startup pause. Try setting OLLAMA_KEEP_ALIVE to 1 (see https://github.com/ollama/ollama/blob/main/docs/faq.md#how-d...).
> pair of ebay RTX 3090s
So... 1700 USD?
With llama3.3 70B, two RTX3090s gives you 48GB of VRAM and the model uses about 44Gb; so the first start is slow (loading the model into VRAM) but after that response is fast (subject to comment above about KEEP_ALIVE).
As I mentioned in a comment upthread, I find ~30B models to be the minimum for getting somewhat reliable output. Though even ~70B models pale in comparison with popular cloud LLMs. Local LLMs just can't compete with the quality of cloud services, so they're not worth using for most professional tasks.
Seem to remember a similar point in late 90s and having to build boxes to run NT/SQL7.0 for local dev
Expect there will be a swing back to on-prem once enterprise starts moving faster and the legal teams begin to understand what is happening data-side with RAG, agents etc.
This seems extremely overpriced.
https://www.ctoservers.com/nvidia-dgx-station-a100---quad-nv...
£ 80k
not to mention all of the other benefits of delegating all the work of setting up the GPUs, public HTTP server, designing the API, security, keeping the model up-to-date with the state of the art, etc
reminds me of the people in the 2000s / early 2010s who would build their own linux boxes back when the platform was super unstable, constantly fighting driver issues etc instead of just getting a mac.
roll-your-own-LLM makes even less sense. at least for those early 2000s linux guys, even if you spent an ungodly amount of time going through the arch wiki or compiling gentoo or whatever, at least those skills are somewhat transferrable to sysadmin/SRE. i dont see how setting up your own instance of ollama has any transferable skills
the only way i could see it making sense is if you're doing super cutting edge stuff that necessitates owning a tinybox, or if you're trying to get a job at openAI or anysphere
But I do.
There are a lot of things to criticize about LLMs (the answer is quite likely to ignore what you're actually asking, for example) but your speed problem looks like a config issue instead. Are you calling the API in streaming mode?
I was running a 5 bit quantized model of codestral 22b with a Radeon RX 7900 (20 gb), compiled with Vulkan only.
Eyeball only, but the prompt responses were maybe 2x or 3x slower than OpenLLMs gpt-4o (maybe 2-4 seconds for most paragraph long responses).
I have a couple of 3090s and have tested most of the popular local LLMs (Llama3, DeepSeek, Qwen, etc.) at the highest possible settings I can run them comfortably (~30B@q8, or ~70B@q4), and they can't keep up with something like Claude 3.5 Sonnet. So I find myself just using Sonnet most of the time, instead of fighting with hallucinated output. Sonnet still hallucinates and gets things wrong a lot, but not as often as local LLMs do.
Maybe if I had more hardware I could run larger models at higher quants, but frankly, I'm not sure it would make a difference. At the end of the day, I want these tools to be as helpful as possible and not waste my time, and local LLMs are just not there yet.
Next major step up is 48GB and then hundreds of GB. But a lot of ML models target 16-24gb since that's in the grad student price range.
from https://www.asacomputers.com/nvidia-l40s-48gb-graphics-card....
nvidia l40s 48gb graphics card Our price: $7,569.10*
Not arguing against 'great', but cost efficiency is questionable. for 10% you can get two used 3090. The good thing about LLMs is they are sequential and should be easily parallelized. Model can be split in several sub-models, by the number of GPUs. Then 2,3,4.. GPUs should improve performance proportionally on big batches, and make it possible to run bigger model on low end hardware.
Kind of one of the reasons AMD is a sleeper stock for me. If people only knew.
Have a look through Hugging Face to see which models interest you. A rough estimate for the amount of VRAM you need is half the model size plus a couple gigs. So, if using the 70B models interests you, two 4080s wouldn't fit it, but two 3090s would. If you're just interested in the 1B, 3B and 7B models (llama 3B is fantastic), you really don't need much at all. A single 3060 can handle that, and those are not expensive.
If you use speculative decoding (a small model generates tokens verified by a larger model, I'm not sure on the specifics) you can get past 20 tokens per second it seems. You can also fit 32B models like Qwen/Qwen Coder at Q6 with lots of context this way, with spec decoding, closer to 40+ tks/s.
Buy a MacMini or MacbookPro with RAM maxed out.
I just bought an M4 mac mini for exactly this use case that has 64GB for ~2k. You can get 128GB on the MBP for ~5k. These will run much larger (and more useful) models.
EDIT: Since the request was for < $1600, you can still get a 32GB mac mini for $1200 or 24GB for $800
[Edit: OK I see I am adding cost when checking due to choosing a larger SSD drive, so $5,000 is more of a fair bottom price, with 1TB of storage.]
Responding specifically to this very specific claim: "Can get 128GB of ram for a reasonable price."
I'm open to your explanation of how this is reasonable — I mean, you didn't say cheap, to be fair. Maybe 128GB of ram on GPUs would be way more (that's like 6 x 4090s), is what you're saying.
For anyone who wants to reply with other amounts of memory, that's not what I'm talking about here.
But on another point, do you think the ram really buys you the equivalent of GPU memory? Is Apple's melding of CPU/GPU really that good?
I'm not just coming from a point of skepticism, I'm actually kind of hoping to be convinced you're right, so wanting to hear the argument in more detail.
It's not for the casual user, but for somebody who derives significant value from running it locally.
Personally I use the MacMini as a hub for a project I'm working on as it gives me full control and is simply much cheaper operationally. A one time ~$2000 cost isn't so bad for replacing tasks that a human would have to do. e.g. In my case I'm parsing loosely organized financial documents where structured data isn't available.
I suspect the hardware costs will continue to decline rapidly as they have in the past though, so that $5k for 128GB will likely be $5k for 256GB in a year or two, and so on.
We're almost at the inflection point where really powerful models are able to be inferenced locally for cheap
I don't see prices falling much in the near term, a Mac Studio M2 Max or Ultra has been keeping its value surprisingly well as of late (mainly because of AI?). Just like 3090s/4090s are holding their value really well also.
If you already own a PC, it makes a hell of a lot more sense to spend $900 on a 3090 than it does to spec out a Mac Mini with 24gb of RAM. Plus, the Nvidia setup can scale to as many GPUs as you own which gives you options for upgrading that Apple wouldn't be caught dead offering.
Oh, and native Linux support that doesn't suck balls is a plus. I haven't benchmarked a Mac since the M2 generation, but the figures I can find put the M4 Max's compute somewhere near the desktop 3060 Ti: https://browser.geekbench.com/opencl-benchmarks
You can easily use the MacMini as a hub for running the LLM while you do work on your main computer (and it won't eat up your system resources or turn your primary computer into a heater)
I hope that more non-mac PCs come out optimized for high RAM SoC, I'm personally not a huge Apple fan but use them begrudgingly.
Also your $900 quote is a used/refurbished GPU. I've had plenty of GPUs burn out on me in the old days, not sure how it is nowadays, but that's a lot to pay for a used part IMO
e.g. LLM as one step of an ETL-style pipeline
Latency of the response really only matters if that response is user facing and is being actively awaited by the user
The only advantage is the M4 Max's ability to have way more VRAM than a 3060 Ti. You won't find many M4 Maxes with just 8 or 16 GB of RAM, and I don't think you can do much except use really small models with a 3060 Ti.
There's no doubt in my mind that the PC is the better performer if raw power is your concern. It's far-and-away the better value if you don't need to buy new hardware and only need a GPU. $2,000 of Nvidia GPUs will buy you halfway to an enterprise cluster, $2,000 of Apple hardware will get you a laptop chip with HBM.
If I could spend $4k on a non-Apple turn key solution that I could reasonably manage in my house, I would totally consider it.
Unless you're specifically speccing out a computer for mobile use, the price premium you spend on a Mac isn't for better software or faster hardware. If you can tolerate Linux or Windows, I don't see why you'd even consider Mac hardware for your desktop. In the OP's position, suggesting Apple hardware literally makes no sense. They're not asking for the best hardware that runs MacOS, they're asking for the best hardware for AI.
> If I could spend $4k on a non-Apple turn key solution that I could reasonably manage in my house, I would totally consider it.
You can't pay Apple $4k for a turnkey solution, either. MacOS is borderline useless for headless inference; Vulkan compute and OpenCL are both MIA, package managers break on regular system updates and don't support rollback, LTS support barely exists, most coreutils are outdated and unmaintained, Asahi features things that MacOS doesn't support and vice-versa... you can't fool me into thinking that's a "turn key solution" any day of the week. If your car requires you to pick a package manager after you turn the engine over, then I really feel sorry for you. The state of MacOS for AI inference is truly no better than what Microsoft did with DirectML. By some accounts it's quite a bit worse.
> You can't pay Apple $4k for a turnkey solution, either.
I've seen/read plenty of success stories of Metal ports of models being used via LM Studio without much configuration/setup/hardware scavenging, so we can just disagree there.
Or live in europe where any wall-socket can give you closer to 3kW. For crazier setups like charging your EV you can have three-phase plugs with ~22kW to play with. 1m2 of floor-space isn't that substantial either unless you already live in a closet in middle of the most crowded city.
That's still a lot of power, though, and does not invalidate your point.
nVidia GPUs actually have similar efficiency, despite all of Apple’s marketing. The difference is they the nVidia GPUs have a much higher ceiling.
Meagre VRAM in these Nvidia consumer GPUs is indeed painful but with increasing performance of smaller LLMs & fine tuned models I don't think 12GB, 14GB, 16GB Nvidia GPUs offering much better performance over a Mac can be easily dismissed.
I don’t have experience with AMD cards so I can’t vouch for them.
The tokens per second is very slow though, but at least it can execute it.
I think the future will be increasingly more powerful NPU's built into future CPU's. That will need to paired with higher bandwidth memory, maybe HBM, or silicon photonics for off chip memory.
Embarrassing and any VCs reading this can contact me to talk about how to fix that. lm-studio is today the closest competition (but not close enough) and Adobe or Microsoft could do it if they fired their current folks which prevent this from happening.
If you're not using Oobabooga, you're likely not playing with the settings on models, and if you're not playing with your models settings, you're hardly even scratching the surface on its total capabilities.
That also supports running local models that you already installed using ollama and things like, adding pdf and preserving chat history. So what am I looking at here with ooba?
Thanks!
Mysty is like obsidian in that it’s not a true prosumer or pro level tool. This is not photoshop for text. It’s Lightroom for text.
Not going to lie, getting Oobabooga to work on my laptop was no small feet. Conda didn't work, python modules didn't work. cmd_macos failed with an error... But eventually I got there by hand. :D So the "Getting Started" was quite steep.
I also deploy text-generation-webui for clients on k8s with gpu for similar reasons.
Last I checked, llamafile / ollama are not as optimised for gpu use.
For image generation I moved from automatic webui to comfyui a few months ago - they're different beasts, for some workflow automatic is easier to use but for most tasks you can create a better workflow with enough comfy extensions.
Facefusion warrants a mention for faceswapping
Is there somewhere I can find a computer like this pre-built?
Like LocalLLama subreddit is extremely helpful. You have to figure out what you want, what tools you want to use (by reading posts and asking questions), and then find some guides on setting up these tools. As you proceed you will run into issues and can easily find most of your answers in the same subreddit (or subreddits designated to the particular tools you're trying to set up) because hundreds of people had the same issues before.
Not to say that the article is useless, it's just semi-organized set of links.
#!/usr/bin/env bash
set -eu
set -o errexit
set -o nounset
set -o pipefail
readonly SCRIPT_SRC="$(dirname "${BASH_SOURCE[${#BASH_SOURCE[@]} - 1]}")"
readonly SCRIPT_DIR="$(cd "${SCRIPT_SRC}" >/dev/null 2>&1 && pwd)"
readonly SCRIPT_NAME=$(basename "$0")
# Avoid issues when wine is installed.
sudo su -c 'echo 0 > /proc/sys/fs/binfmt_misc/status'
# Graceful exit to perform any clean up, if needed.
trap terminate INT
# Exits the script with a given error level.
function terminate() {
level=10
if [ $# -ge 1 ] && [ -n "$1" ]; then level="$1"; fi
exit $level
}
# Concatenates multiple files.
join() {
local -r prefix="$1"
local -r content="$2"
local -r suffix="$3"
printf "%s%s%s" "$(cat ${prefix})" "$(cat ${content})" "$(cat ${suffix})"
}
# Swapping this symbolic link allows swapping the LLM without script changes.
readonly LINK_MODEL="${SCRIPT_DIR}/llm.gguf"
# Dereference the model's symbolic link to its path relative to the script.
readonly PATH_MODEL="$(realpath --relative-to="${SCRIPT_DIR}" "${LINK_MODEL}")"
# Extract the file name for the model.
readonly FILE_MODEL=$(basename "${PATH_MODEL}")
# Look up the prompt format based on the model being used.
readonly PROMPT_FORMAT=$(grep -m1 ${FILE_MODEL} map.txt | sed 's/.*: //')
# Guard against missing prompt templates.
if [ -z "${PROMPT_FORMAT}" ]; then
echo "Add prompt template for '${FILE_MODEL}'."
terminate 11
fi
readonly FILE_MODEL_NAME=$(basename $FILE_MODEL)
if [ -z "${1:-}" ]; then
# Write the output to a name corresponding to the model being used.
PATH_OUTPUT="output/${FILE_MODEL_NAME%.*}.txt"
else
PATH_OUTPUT="$1"
fi
# The system file defines the parameters of the interaction.
readonly PATH_PROMPT_SYSTEM="system.txt"
# The user file prompts the model as to what we want to generate.
readonly PATH_PROMPT_USER="user.txt"
readonly PATH_PREFIX_SYSTEM="templates/${PROMPT_FORMAT}/prefix-system.txt"
readonly PATH_PREFIX_USER="templates/${PROMPT_FORMAT}/prefix-user.txt"
readonly PATH_PREFIX_ASSIST="templates/${PROMPT_FORMAT}/prefix-assistant.txt"
readonly PATH_SUFFIX_SYSTEM="templates/${PROMPT_FORMAT}/suffix-system.txt"
readonly PATH_SUFFIX_USER="templates/${PROMPT_FORMAT}/suffix-user.txt"
readonly PATH_SUFFIX_ASSIST="templates/${PROMPT_FORMAT}/suffix-assistant.txt"
echo "Running: ${PATH_MODEL}"
echo "Reading: ${PATH_PREFIX_SYSTEM}"
echo "Reading: ${PATH_PREFIX_USER}"
echo "Reading: ${PATH_PREFIX_ASSIST}"
echo "Writing: ${PATH_OUTPUT}"
# Capture the entirety of the instructions to obtain the input length.
readonly INSTRUCT=$(
join ${PATH_PREFIX_SYSTEM} ${PATH_PROMPT_SYSTEM} ${PATH_PREFIX_SYSTEM}
join ${PATH_SUFFIX_USER} ${PATH_PROMPT_USER} ${PATH_SUFFIX_USER}
join ${PATH_SUFFIX_ASSIST} "/dev/null" ${PATH_SUFFIX_ASSIST}
)
(
echo ${INSTRUCT}
) | ./llamafile \
-m "${LINK_MODEL}" \
-e \
-f /dev/stdin \
-n 1000 \
-c ${#INSTRUCT} \
--repeat-penalty 1.0 \
--temp 0.3 \
--silent-prompt > ${PATH_OUTPUT}
#--log-disable \
echo "Outputs: ${PATH_OUTPUT}"
terminate 0
map.txt: c4ai-command-r-plus-q4.gguf: cmdr
dare-34b-200k-q6.gguf: orca-vicuna
gemma-2-27b-q4.gguf: gemma
gemma-2-7b-q5.gguf: gemma
gemma-2-Ifable-9B.Q5_K_M.gguf: gemma
llama-3-64k-q4.gguf: llama3
llama-3-64k-q4.gguf: llama3
llama-3-1048k-q4.gguf: llama3
llama-3-1048k-q8.gguf: llama3
llama-3-8b-q4.gguf: llama3
llama-3-8b-q8.gguf: llama3
llama-3-8b-1048k-q6.gguf: llama3
llama-3-70b-q4.gguf: llama3
llama-3-70b-64k-q4.gguf: llama3
llama-3-smaug-70b-q4.gguf: llama3
llama-3-giraffe-128k-q4.gguf: llama3
lzlv-q4.gguf: alpaca
mistral-nemo-12b-q4.gguf: mistral
openorca-q4.gguf: chatml
openorca-q8.gguf: chatml
quill-72b-q4.gguf: none
qwen2-72b-q4.gguf: none
tess-yi-q4.gguf: vicuna
tess-yi-q8.gguf: vicuna
tess-yarn-q4.gguf: vicuna
tess-yarn-q8.gguf: vicuna
wizard-q4.gguf: vicuna-short
wizard-q8.gguf: vicuna-short
Templates (all the template directories contain the same set of file names, but differ in content): templates/
├── alpaca
├── chatml
├── cmdr
├── gemma
├── llama3
├── mistral
├── none
├── orca-vicuna
├── vicuna
└── vicuna-short
├── prefix-assistant.txt
├── prefix-system.txt
├── prefix-user.txt
├── suffix-assistant.txt
├── suffix-system.txt
└── suffix-user.txt
If there's interest, I'll make a repo.Take for instance, self-hosting your website may have all these considerations, but you're getting information from the LLMs. It would be helpful to know that the LLM is in your control.
Same with LLMs, you can use providers who don’t log requests and SOC2 compliant.
Small models that run locally is a waste of time as they don’t have adequate value compared to larger models.
It depends on how important the web site is and what the purpose is. Personal blog (Ghost, WordPress)? File sharing (Nextcloud)? Document collaboration (Etherpad)? Media (Jellyfin)? Why _not_ run it from your room with a reverse proxy? You're paying for internet anyway.
> Same with LLMs, you can use providers who don’t log requests and SOC2 compliant.
Sure. Until they change their minds and decide they want to, or until they go belly-up because they didn't monetize you.
> Small models that run locally is a waste of time as they don’t have adequate value compared to larger models.
A small model won't get you GPT-4o value, but for coding, simple image generation, story prompts, "how long should I boil an egg" questions, etc. it'll do just fine, and it's yours. As a bonus, while a lot of energy went into creating the models you'd use, you're saving a lot of energy using them compared to asking the giant models simple questions.
Millions of businesses trust Google and Microsoft with their spreadsheets, emails, documents, etc… The EULAs are very similar.
If any major provider was caught violating the terms of these agreements they’d lose a massive amount of business! It’s just not worth it.
https://youtu.be/pvcK9X4XEF8?si=9j2ELkbXxHGoZEnZ&t=1157
> ...if you're a Silicon Valley entrepreneur, which hopefully all of you will be, is if it took off then you'd hire a whole bunch of lawyers to go clean the mess up, right. But if nobody uses your product it doesn't matter that you stole all the content.
Yeah. Because I have tons and tons of fear, uncertainty and doubt about how these giga-entities control so many aspects of my life. Maybe I shouldn't be spreading it but on the other hand they are so massively huge that any FUD I spread will have absolutely no effect on how they operate. If you want to smell real FUD go sniff Microsoft[1].
[1] https://en.wikipedia.org/wiki/Fear,_uncertainty,_and_doubt#M...
But it turns out they did /do.
You're commenting on a post about how the author runs LLMs locally because they find them useful. Do you think they would run them and write an entire post on how they use it if they didn't find it useful? The author is seriously using them.
Can you expand a bit on how you might do this?
The same applies (to smaller extent) to Google, and more recently to Alibaba.
A lot of free (not open-source, mind you) LLMs are either from Meta or built by other smaller outfits by tweaking Meta's LLMs (there are exceptions, like Mistral).
One thing to keep in mind is that while training a model requires A LOT of data, manual intervention, hardware, power, and time, using these LLMs (inference, or forward pass in the neural network) is really not. Right now it typically requires NVidia GPU with lots of VRAM but I have no doubt in just a few months/maybe a year someone will find much easier way to do that without sacrificing the speed.
While it does appear that a large number of people are using these for RP and... um... similar stuff, I do find code generation to be fairly good, esp with some recent Qwens (from Alibaba). Full disclaimer: I do use this sparingly, either to get the boilerplate, or to generate a complete function/module with clearly defined specs, and I sometimes have to re-factor the output to fit my style and preferences.
I also use various general models (mostly Meta's) fairly regularly to come up with and discuss various business and other ideas, and get general understanding in the areas I want to get more knowledge of (both IT and non-IT), which helps when I need to start digging deeper into the details.
I usually run quantized versions (I have older GPU with lower VRAM).
Wordsmithing and generating illustrations for my blog articles (I prefer Plex, it's actually fairly good with adding various captions to illustrations; the built-in image generator in ChatGPT was still horrible at it when I tried it a few months ago).
Some resume tweaking as part of manual workflow (so multiple iterations with checking back and forth between what LLM gave me and my version).
So it's mostly stuff for my own personal consumption, that I don't necessarily trust the cloud with.
If you have a SaaS idea with an LLM or other generative AI at its core, processing the requests locally is probably not the best choice. Unless you're prototyping, in which case it can help.
For me though, this would be all upside because I have largely explored technical topics with language models that would only be impressive to an employer.
At this point, it is like asking what does someone use a computer for? The use cases are so varied.
I can see how it would be interesting for myself to setup a local model just for the fun of setting it up. When it comes down to it for me though it is just so much easier to pay $20 a month for Sonnet that it isn't even close or really a decision point.
They’re giving away these expensive models because it doesn’t hurt Meta’s ad business but reduces risk that competitors could grow moats using closed models. The old “commoditize your complements” play.
The multimillion (or billion) dollar collections of hardware assemble the datasets that we (people who run LLMs locally) run. The non-open datasets that we-host-LLMs-for-money companies do the same, and their data isn't all that much fancier. Open LLMs are catching up to closed ones, and the competition means everyone wins except for the victims of Nvidia's price gouging.
This is a bit like confusing the resources that go in to making a game with the resources needed to run and play that game. They're worlds different.
TRAINING an LLM requires a lot of compute. Running inference on a pre-trained LLM is less computationally expensive, to the point where you can run LLAMA (cost $$$ to train on Meta's GPU cluster) with CPU-based inference.