876 karma · joined October 28, 2019
The 14B Q4_K_M needs 9GB, but Q3_K_M is 7.3GB. But you also need some room for context. Still, maybe using `--override-tensor` in llama.cpp would get you a 50% improvement over "naively" offloading layers to the GPU. Or possibly GPT-OSS-20B. It's 12.1GB in MXFP4, but it’s a MOE model so only a part of it would need to be on the GPU. On my dedicated 12GB 3060 it runs at 85 t/s, with a smallish context. I've also read on Reddit some claims that Qwen3 4B 2507 might be better than 8B, because Qwen never released a "2507" update for 8B.
For image generation, Flux and Qwen Image work with ComfyUI. I also use Nunchaku, which improves speed considerably.
# Runs the IP change check on Mon - Sun at 09:00, 10:30, 12:00, 20:00 -- If the power goes out or the router reboots I get a new IP. On the server I use fail2ban and if I log into the admin panel I might get banned for making too many requests. So my IP needs to be "blessed".
# Runs Letsencrypt certificate expiry check on Sundays at 11:00 and 18:00 -- I still have a server where I update the certificates by hand.
# Runs the "daily" backup -- Just rsync
# Download Godaddy auction data every day at 19:00 -- I don't actively do this anymore but I used to check, based on certain criteria, for domains that were about to expire.
# Download the sellers.json on the 1st of every month at 19:00 -- I use this to collect data on websites that appear and disappear from the Mediavine and Adthrive sellers.json
The test I ran was to ask the LLM about an expired domain of a doctor (obstetrician). I no longer remember the exact domain, but it was similar to annasmithmd.com. One LLM would tell me it used to belong to a doctor named Megan Smith. Another got the name right, Anna Smith, but when I asked it what kind of a doctor, which specialty, it answered pediatrician.
So the LLM had no clue, but from the name of the domain it could infer (I guess that's why they call it inference) that the "md" part was associated with doctors.
By the way, newer LLMs are very good at making domains more human readable by splitting them into words.
I have a Ryzen 3 4100. Just tested Qwen2.5-Coder-32B-Instruct-Q3_K_S.gguf with llama.cpp.
CPU-only:
54.08 t/s prompt eval
2.69 t/s inference
---
CPU + 52/65 layers offloaded to GPU (RTX 3060 12GB):
166.79 t/s prompt eval
6.62 t/s inference
I started with a local system using llama.cpp on CPU alone and for short questions and answers it was OK for me. Because (in 2023) I didn't know if LLMs would be any good, I chose cheap components https://news.ycombinator.com/item?id=40267208.
Since AWS was getting pretty expensive, I also bought an RTX 3060(16GB), an extra 16GB RAM (for a total of 32GB) and a superfast 1TB M.2 SSD. The total cost of the components was around €620.
Here are some basic LLM performance numbers for my system:
I have an RTX 3060(12GB) and 32GB RAM. Just ran Qwen2.5-14B-Instruct-Q4_K_M.gguf in llama.cpp with flash attention enabled and 8K context. I get get 845t/s for prompt processing and 25t/s for generation.
For a while I even ran llama.cpp without a GPU (don't recommend it for diffusion) and with the same model (Qwen2.5 14B) I would get 11t/s for processing and 4t/s for generation. Acceptable for chats with short questions/instructions and answers.
You can get some inspiration from businesses for sale on Empire Flippers https://empireflippers.com/marketplace/.
As a rule of thumb for choosing the niche, pick from one of these https://support.google.com/admob/answer/3150953?hl=en
- B550MH motherboard
- Ryzen 3 4100 CPU
- 32GB (2x16) RAM cranked up to 3200MHz (prompt generation in memory bound)
- 256GB M.2 NVMe (helps with loading models faster)
- Nvidia 3060 12GB
Software-wise, I use llamafile because on the CPU it's faster by 10-20% for prompt processing than llama.cpp.
Performance "Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf":
CPU-only: 23.47 t/s (processing), 8.73 t/s (generation)
GPU: 941.5 t/s (processing), 29.4 t/s (generation)
As another option, you can rent a machine with a decent GPU on vast.ai. An Nvidia 3090 can be rented for about $0.20/hr.
Don't use any tools. I run it from the command line:
./main -f ~/Desktop/prompts/multishot/llama3-few-shot-prompt-10.txt -m ~/Desktop/models/Meta-Llama-3-8B-Instruct-Q8_0.gguf --temp 0 --color -c 1024 -n -1 --repeat_penalty 1.2 -tb 8 --log-disable 2>/dev/null
I prefer `main` to the new `llama-cli` because when searching history for "llama" I want to get commands that contain the "llama" models, not "mistral" ones, for example.
if (req.http.host ~ "^(?i)(example.com|www.example.com)") { #redirect to https } else { return(synth(403, "Not allowed.")); }
It basically checks if the host is my domain. I don’t know know what the equivalent of `req.http.host` is on the web server you use. This "solution" might run into issues with Google Translate, but I’m not sure.
I looked today on Newegg and similar PC components would cost $220-230.
From a performance perspective, I get about 9 tokens/s from mistral-7b-instruct-v0.2.Q4_K_M.gguf with a 1024 context size. This is with overclocked RAM which added 15-20% more speed.
The Mac Mini is probably faster than this. However the custom built PC route gives you the option to add more RAM later on to try bigger models. It also lets you add a decent GPU. Something like a used 3060, as one of comments says.
I even found the page listing one of the themes I sold with Wayback Machine - "Winning bid: $125". Some people were buying these themes, change the "Theme by" link in the footer and give them away for free. Probably for SEO reasons.
The /etc/hosts.work file contains a list of websites I don’t want to check during work. But HN is not in this list, yet.
Another route would be AWS SES. They have a free tier. For managing the subscribe/unsubscribe process you could still use MailChimp or a self-hosted solution like Keila https://www.keila.io/, although I have no experience with this.
I noticed they also offer an RSS to Email option, but I couldn’t find if it’s available on the Free plan. https://mailchimp.com/features/rss-to-email/
As a teenager I’ve been "traumatized" by Pascal, but I was lucky to find a refuge in HTML and CSS. After I got a job as a web designer I discovered Python and I realized that I could actually program.
The people who came up with this new syntax have different needs that I do and probably have a different kind of brain than mine.