HNHacker News
TopNewBestAskShowJobs

apitman

13,687 karma · joined December 11, 2014

Software architect on the iobio team at the University of Utah Eccles Institute of Human Genetics. I'm passionate about the application of computer science to solving health problems. Hans Rosling is my hero.

Personal site:

apitman.com

Projects:

IndieBits.io - A community for data ownership, self-hosting, and decentralization.

LastLogin.net - A free, privacy-focused login provider

TakingNames.io - Domain names for self-hosters

boringproxy.io - Simple, e2ee tunneling proxy

droplock.apitman.com - Simple secure secret sharing

submissionscomments
apitman··on Qwen 3.8 27B
Yeah make sure you're using MTP and potentially tensor parallelism.
apitman··on Unsloth Qwen3.8-27B GGUF files
This is specifically about unsloth having day-1 GGUF's available.
apitman··on Qwen 3.8 27B
I used GPT-5.6 Sol high to optimize it, and it claimed it was getting 50. I'm seeing ~40 on my goto smoketest: "Make me a vector add in CUDA".

Funny side note. It successfully one shot the program, but it wasn't able to run it because there literally wasn't enough VRAM left to allocate CUDA memory. Watching it try to debug that was fascinating. I'm pretty sure it would have killed the llama-server (and thus itself) if it hadn't been running in a separate container.

apitman··on Qwen 3.8 27B
Running it on 2x3060 now. Works pretty well but VRAM is tight. 4bit quants. 1x128k context, 8bit KV, MTP on.
apitman··on DeepSeek peak/off-peak pricing update
The Luna price dropped the day before a massive update to Flash (0731 update). GP may be referring to that version.
apitman··on DeepSeek API Pricing Update
Look at OpenRouter. They have lots of providers
apitman··on DeepSeek V4 Pro 0813
But why? Luna Max is almost the same intelligence as Terra xhigh and way way cheaper. And Terra max is almost the same as Sol high. I just don't really see a place for Terra but slower.
apitman··on DeepSeek V4 Pro 0813
Wait people use terra?
apitman··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
Are there any projects that track how much usage of each model translates to how much percentage drop in weekly/5hr windows?
apitman··on Grok 4.6
https://artificialanalysis.ai/models/grok-4-6
apitman··on Docker Sandboxes – Disposable, isolated sandboxes for AI agents
This is what I do. Nice benefit is it lets me passthrough my GPU and share it between multiple containers. Incus is awesome.
apitman··on Docker Sandboxes – Disposable, isolated sandboxes for AI agents
Fun fact: QEMU runs natively on windows and supports acceleration with WHP. It works surprisingly well.
apitman··on DeepSeek V4 Flash 0731
> OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4

At DeepSeek's absurdly low rates or market rates?

apitman··on DeepSeek V4 Flash 0731
I'm getting like 25 tok/s on 2x RTX Pro 6000. This is with llama.cpp, but I had GPT tune it for me. I was under the impression vLLM was at most ~2x faster, and usually for highly parallel loads. Any tips on where I should look first for an obvious blunder?

I'm guessing tensor parallelism or similar?

apitman··on DeepSeek V4 Flash 0731
These are very interesting results, and honestly hard to believe, even as a big 0731 fan.

If I'm reading the chart correctly, a couple observations:

* deepseek-v4-flash-0731 max is better than kimi-k3 max

* glm-5.2 is dumber than a box of rocks (this must be on low reasoning or something, right?)

This is way more extreme than other results I'm seeing, like those from Artificial Analysis.

apitman··on DeepSeek V4 Flash 0731
I've found it to be pretty good so far.
apitman··on DeepSeek V4 Flash 0731
DeepSeek has far cheaper cache pricing. That's the difference.
apitman··on Qwen3.8 Max now ranked as the best overall model by agentic index
These numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?
apitman··on Qwen3.8 Max now ranked as the best overall model by agentic index
Welp. That didn't last long
apitman··on Qwen3.8 Max now ranked as the best overall model by agentic index
As low as it is, switching between providers on OpenRouter is still lower.

That said, it's a fair point. For me, it boils down to things covered here: https://earendil.com/posts/session-portability/

Things like obscured reasoning traces.

apitman··on Qwen3.8 Max now ranked as the best overall model by agentic index
Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe "agentic index" is still the metric to look at for coding performance. I could be wrong about that though.
apitman··on Qwen3.8 Max now ranked as the best overall model by agentic index
For one thing, providers of open models can't arbitrarily increase their prices without facing competition.
apitman··on Devtools must be open source
If your coding agent workflow isn't sandboxed, backed up, and rollback-able, it's fundamentally broken. And yes I realize most of us aren't working this way today. But I think we'll have the tooling pretty soon. There's currently a cambrian explosion of solutions in this space.
apitman··on Rust All Hands 2026 Retrospective
Oh are you saying the human remains the bottleneck?
apitman··on Rust All Hands 2026 Retrospective
I don't doubt this is true for many projects currently (though it's not for a bevy project I'm working on).

Have you tried Cerebras, Groq, Taalas, et al? It was a paradigm shift for me.

apitman··on Rust All Hands 2026 Retrospective
I could see Rust becoming the language for coding agents, because it has such solid guardrails built in. I could also see it slipping into obscurity as LLMs get faster and compile times become a more and more obvious bottleneck on iteration speeds.
apitman··on DeepSeek-V4-Flash Update
I grew up on Dune lore and the first movie is one of my favorites of all time. Highly recommended for any Dune fan. I didn't care for the second one.
apitman··on The session you cannot take with you
Can you give some more details on the technical differences between completions and responses that people actually care about? I found this information surprisingly hard to drum up.
apitman··on The session you cannot take with you
I'm sure you were referring to the desktop app but codex CLI is open source
apitman··on Show HN: BitBang – Reach machines behind NAT from a browser, no account
Would love to have this on https://github.com/anderspitman/awesome-tunneling as soon as it hits 100 stars.
← PreviousPage 3 of 34Next →