HNHacker News
TopNewBestAskShowJobs

DrPhish

978 karma · joined June 10, 2010

submissionscomments
DrPhish··on macOS 27 Golden Gate – Review
Yes, the disrespect Apple/Google/MS has for their users (on equipment they have purchased!) is palpable. They seem to have a real sense of entitlement these days. It used to be universally accepted that losing or corrupting user data was a (if not THE) cardinal sin in development. In many ways we get what we tolerate, but the lack of viable alternatives in some spaces (esp cellphones and TVs) makes it very hard to hold the big players to any kind of account economically. At least no one is holding a gun to our heads to use MacOS or other proprietary systems on our desktops/laptops/servers for the most part.
DrPhish··on 216M Spy TVs – The LG Smart TV Problem [video]
Every speaker is also a microphone
DrPhish··on The case for physical media ownership
When you watch on vhs or laserdisc the loss of resolution only bothers you til the movie sucks you in.

At that point it’s irreverent because your eyeballs are not watching a long sequence of pretty still pictures, but rather your brain is watching a story in a way similar to reading a good book.

DrPhish··on Framework's 10G Ethernet module exposes USB-C's complexity
I redid everything that matters in my house/homelab with DAC cables for exactly that reason. Order of magnitude difference in watts and heat
DrPhish··on xAI joins SpaceX
Wouldn’t you also need to include the Ancient Greek phryctoriae military fire signalling system by that logic? It probably wasn’t the first, at that.
DrPhish··on Antirender: remove the glossy shine on architectural renderings
Very “futurological congress” thought
DrPhish··on Spherical Snake
Also s4nake, the concept in a 4k binary from the demoscene circa 2013

https://www.pouet.net/prod.php?which=61035

DrPhish··on Sick of smart TVs? Here are your best options
Just use a commercial signage display
DrPhish··on Btop: A better modern alternative of htop with a gamified interface
It’s not a process monitor, really, but to me the AWS Lightsail monitor tab feels like this. The “sustainable” line hits me right in the OCD to keep me grinding on cpu usage of the workload to keep extra spend at zero.
DrPhish··on Mistral raises 1.7B€, partners with ASML
Model back doors feel like baseless fearmongering. Something like https://rentry.org/IsolatedLinuxWebService should provide a good guarantee of privacy and security.
DrPhish··on My Own DNS Server at Home – Part 1: IPv4
I have this as well, but run a heavily locked down and isolated BIND server with NSD and Unbound for external authoritative and internal caching DNS respectively.

Its easy to feed an RBL to unbound to do pi-hole type work, I use pf to transparently redirect all external DNS requests to my local unbound server but I get the bind automation around things like DNSSEC, DHCP ddns and ACME cert renewals.

I'm surprised this isn't a more common stack.

DrPhish··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Thanks, it was a bit of a gamble at the time (lots of dodgy ebay parts), but it paid off.

R1 starts at about 10t/s on an empty context but quickly falls off. I'd say the majority of my tokens are generating around 6t/s.

Some of the other big MoE models can be quite a bit faster.

I'm mostly using QwenCoder 480b at Q8 these days for 9t/s average. I've found I get better real-world results out of it than K2, R1 or GLM4.5.

DrPhish··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
No, I’m running the unquantized 120b
DrPhish··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
2xEPYC Genoa w/768GB of DDR5-4800 and an A5000 24GB card. I built it in January 2024 for about $6k and have thoroughly enjoyed running every new model as it gets released. Some of the best money I’ve ever spent.
DrPhish··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Its also easy to do 120b on CPU if you have the resources. I had 120b running on my home LLM CPU inference box in just as long as it took to download the GGUFs, git pull and rebuild llama-server. I had it running at 40t/s with zero effort and 50t/s with a brief tweaking. Its just too bad that even the 120b isn't really worth running compared to the other models that are out there.

It really is amazing what ggerganov and the llama.cpp team have done to democratize LLMs for individuals that can't afford a massive GPU farm worth more than the average annual salary.

DrPhish··on Qwen3-235B-A22B-Thinking-2507
Thanks Daniel. I know you upload them, but I was hoping for some solid numbers on your dynamic q8 vs a naive quant. There doesn't seem to be anything on either of those links to show improvement at those quant levels.

My gut feeling is that there's not enough benefit to outweigh the risk of putting a middleman in the chain of custody from the original model to my nvme.

However, I can't know for sure without more testing than I have the time or inclination for, which is why I was hoping there had been some analysis you could point me to.

DrPhish··on Qwen3-235B-A22B-Thinking-2507
I generally download the safetensors and make my own GGUFs, usually at Q8_0. Is there any measurable benefit to your dynamic quants at that quant level? I looked at your dynamic quant 2.0 page, but all the charts and graphs appear to cut off at Q4.
DrPhish··on Las Vegas is embracing a simple climate solution: More trees
Trees are pure carbon. I have heard a number of weak “yeah, but…” arguments that try to diminish the fact, but a central, common sense thesis remains.

If we are truly worried about climate change and are unable to curb our consumption, then we should plant as many trees as we can and aggressively shift as much of our long-lived infrastructure to using wood products as possible.

Grow it, use it, maintain it.

DrPhish··on Snake on a Globe
Here’s a demoscene prod from 2013 that executes a similar idea in 4 kilobytes https://m.pouet.net/prod.php?which=61035
DrPhish··on Mermaid: Generation of diagrams like flowcharts or sequence diagrams from text
I can second this. I’ve been using R1 to both straight up generate mermaid as well as making custom mermaid syntax generators for dynamic diagramming
DrPhish··on OpenAI Audio Models
In my opinion GPT-SoVITS is the best if you can put in the effort. I'm still using v2 since the output is so good. Its also the best multilingual one in my testing on Japanese inputs.
DrPhish··on Nping – ping, but with a graph or table view
Smokeping is an amazing and underrated resource as a network health metric and diagnostics tool.

If you have a network monitoring or asset system you can export IP addresses from, you should use a small glue script to automatically build a smokeping configuration. I've got one for our LAN and one for the WAN at each of our sites so I can track down issues at either level.

The LAN connection charts are a great daily sanity check, and the WAN connections (I have every-to-every for each site so any and all inter-site issues can be seen) can help keep your ISP honest with the service they're delivering.

DrPhish··on Show HN: Transform your codebase into a single Markdown doc for feeding into AI
That's very nice and compact. I do the same with a short bash script, but wrap each file in triple-backticks and attempt to put the correct language label on each eg:

Filename: demo.py

```python

   ...python code here...
```
DrPhish··on Building a personal, private AI computer on a budget
The build guide index page is newer, but to be fair, the mikubox rentry is from Oct 6, 2023.

If that isn't "ancient" in terms of AI workstation build guides, then I don't know what is.

DrPhish··on Building a personal, private AI computer on a budget
This is just a limited recreation of the ancient mikubox from https://rentry.org/lmg-build-guides

Its funny to see people independently "discover" these builds that are a year plus old.

Everyone is sleeping on these guides, but I guess the stink of 4chan scares people away?

DrPhish··on US bill proposes jail time for people who download DeepSeek
"Being in possession of a contraband Chinese Artificial Intelligence" is honestly one of the most cyberpunk things I can imagine. I hadn't felt like I lived in the future until now, honestly.
DrPhish··on How to run DeepSeek R1 locally
“you can get a dual EPYC server with 768GB RAM - CPU inference only at around 6-8 tokens/sec.”

This is what I run at home. I built it just over a year ago and have run every single model that has been released.

DrPhish··on Ask HN: Has Anyone Successfully Run DeepSeek-R1 on an Nvidia Jetson Nano?
The Nano only has 4GB VRAM and DS-R1 is 671B FP8 parameters (equivalent to 671GB model size).

You need something with about 800GB to run the full model with context. You'd still need 400GB to even run a half-sized Q4 quant of R1, so there is no reasonable way that it would work.

DrPhish··on DeepSeek-R1
Making your own ggufs is trivial: https://rentry.org/tldrhowtoquant/edit

It's a bit harder when they've provided the safetensors in FP8 like for the DS3 series, but these smaller distilled models appear to be BF16, so the normal convert/quant pipeline should work fine.

DrPhish··on Show HN: News Minimalist – News ranked by significance
Thanks for the reply! This is perhaps not so much a Hacker News type question since this place is very VC focused, but have you considered publishing any papers on your system? I think it would make a fascinating and valuable bit of research.

Or, even farther off the deep-end: have you considered open-sourcing any old versions of your prompts or pipeline? Say one year after they are superseded in your production system?

Page 1 of 10Next →