HNHacker News
TopNewBestAskShowJobs

pich

19 karma · joined June 21, 2012

CTO at Archdesk (construction ERP). Earlier: fintech exit, and a long time ago the kind of curiosity that got me labelled a hacker.

I write about the cost side of AI — intelligence per joule, verification cost, agent autonomy that survives contact with evidence: https://piszczek.pl

Two terms I coined and keep refining: "Joule Wars" (the race shifting from capability to useful intelligence per joule) and "Proof-Adjusted Autonomy" (raw autonomy discounted by whether the work is proven, independently validated and delivered in time).

Kraków, Poland.

submissionscomments
pich··on Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
A 3090 has 936 GB/s of bandwidth vs 432 GB/s on the 70W RTX PRO 4000 SFF, so 70 t/s there is not surprising. The interesting constraint here was fitting a workload-tuned 5.01 BPW quant + 256K + MTP into 24 GB while working with less than half the memory bandwidth
pich··on Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
Its not quite that simple here. The iMatrix-guided hybrid uses different quantization levels per tensor/layer, so there isnt one honest Q4/Q5/NVFP4 label I can put in the title
pich··on Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
5.01 BPW custom hybrid: bulk NVFP4, selected Q5_K/Q6_K tensors from an iMatrix, Q6_K embeddings and Q8_0 lm_head. The iMatrix was built from 5,472 messages across 296 real Hermes sessions
pich··on Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
vLLM is probably the key difference there… its scheduler is built around batching/concurrency, while this setup is heavily optimized llama.cpp for single-stream latency