978 karma · joined June 10, 2010
At that point it’s irreverent because your eyeballs are not watching a long sequence of pretty still pictures, but rather your brain is watching a story in a way similar to reading a good book.
Its easy to feed an RBL to unbound to do pi-hole type work, I use pf to transparently redirect all external DNS requests to my local unbound server but I get the bind automation around things like DNSSEC, DHCP ddns and ACME cert renewals.
I'm surprised this isn't a more common stack.
R1 starts at about 10t/s on an empty context but quickly falls off. I'd say the majority of my tokens are generating around 6t/s.
Some of the other big MoE models can be quite a bit faster.
I'm mostly using QwenCoder 480b at Q8 these days for 9t/s average. I've found I get better real-world results out of it than K2, R1 or GLM4.5.
It really is amazing what ggerganov and the llama.cpp team have done to democratize LLMs for individuals that can't afford a massive GPU farm worth more than the average annual salary.
My gut feeling is that there's not enough benefit to outweigh the risk of putting a middleman in the chain of custody from the original model to my nvme.
However, I can't know for sure without more testing than I have the time or inclination for, which is why I was hoping there had been some analysis you could point me to.
If we are truly worried about climate change and are unable to curb our consumption, then we should plant as many trees as we can and aggressively shift as much of our long-lived infrastructure to using wood products as possible.
Grow it, use it, maintain it.
If you have a network monitoring or asset system you can export IP addresses from, you should use a small glue script to automatically build a smokeping configuration. I've got one for our LAN and one for the WAN at each of our sites so I can track down issues at either level.
The LAN connection charts are a great daily sanity check, and the WAN connections (I have every-to-every for each site so any and all inter-site issues can be seen) can help keep your ISP honest with the service they're delivering.
Filename: demo.py
```python
...python code here...
```If that isn't "ancient" in terms of AI workstation build guides, then I don't know what is.
Its funny to see people independently "discover" these builds that are a year plus old.
Everyone is sleeping on these guides, but I guess the stink of 4chan scares people away?
This is what I run at home. I built it just over a year ago and have run every single model that has been released.
You need something with about 800GB to run the full model with context. You'd still need 400GB to even run a half-sized Q4 quant of R1, so there is no reasonable way that it would work.
It's a bit harder when they've provided the safetensors in FP8 like for the DS3 series, but these smaller distilled models appear to be BF16, so the normal convert/quant pipeline should work fine.
Or, even farther off the deep-end: have you considered open-sourcing any old versions of your prompts or pipeline? Say one year after they are superseded in your production system?