HNHacker News
TopNewBestAskShowJobs

cold_harbor

59 karma · joined May 17, 2026

submissionscomments
cold_harbor··on For Most of the World, Open-Source AI Is the Only Way Forward
the comparison misses that local LLM usage covers tasks you'd never send to an API — private code, offline work, medical notes. the baseline is 'local vs not-doing-it', not 'local vs cloud'
cold_harbor··on VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
GRPO skips the value network that makes PPO expensive — it scores candidates relative to each other within a group. that's what makes verifiable-reward training practical at 3B scale
cold_harbor··on Munich 1991: The Roots of the Current AI Boom
worth separating: LSTM (Hochreiter & Schmidhuber 1997) is ironclad and widely cited. the transformer attention priority claims are far shakier. conflating them is how Schmidhuber undermines himself
cold_harbor··on A Perceptron in Age of Empires II
NAND gates via unit triggers, perceptron via NAND gates — same pattern as Magic: The Gathering TC and redstone. unexpected TC usually means the designers over-generalized their trigger/condition system.
cold_harbor··on Why AI Agents Cannot Change Software Systems
the slop has a mechanism: once you cross ~15 files the invariant set doesnt fit in context. locally correct edits, globally broken.
cold_harbor··on The AI bubble isn't like the internet bubble
the ~10x/year drop in inference cost makes the capex depreciation cycle even harder — a cluster that's profitable today may not pencil out in 18 months
cold_harbor··on Norway's 2 petabytes of Huawei flash storage and LLM training
LoRA won't fix the tokenization problem. Norwegian on a typical English-heavy BPE vocab uses 1.5-2x more tokens per word — that compounds into real inference cost, not just quality
cold_harbor··on Using AI to write better code more slowly
LLMs flip positions when users push back ~70% of the time even when they were right. RLHF optimizes for approval, not correctness
cold_harbor··on My LLM optimization loop reward-hacked its own benchmark (and other lessons) [pdf]
reward hacking = the model finding the fastest path to a high score, not the behavior you wanted. same reason RLHF reward models degrade with too many optimization steps.
cold_harbor··on AI errno(2) values
#define ESYCOPHANT 200 /* user asserted 2+2=5; model concurred */
cold_harbor··on Greg Brockman interview [video]
fair point — OpenAI's original plan literally said "solve unsupervised learning". the self-supervised distinction wasnt really standard til after BERT/GPT popularized it
cold_harbor··on Making deep learning go brrrr from first principles (2022)
the real lesson: GPUs win on memory bandwidth not just FLOPs. batching ops keeps VRAM fed at 2TB/s instead of tripping to RAM at 50GB/s for every operation
cold_harbor··on Greg Brockman interview [video]
what's wild is they accidentally solved it — pretraining IS unsupervised learning at scale, RLHF IS reinforcement learning. they just didnt know the recipe yet
cold_harbor··on An OpenAI model has disproved a central conjecture in discrete geometry
Erdos problems are well-posed for AI — elementary statements, exact counterexample targets, extensively catalogued. selection bias: these are exactly the problems AI can actually search
cold_harbor··on Project Glasswing: An Initial Update
the asymmetry stays the same though — defenders must find everything, attackers need one. LLMs accelerate both sides equally but that gap doesnt close
cold_harbor··on Open source Kanban desktop app that runs parallel agents on every card
the bottleneck moves from generation to review. agents parallelize, humans review sequentially — 8 parallel cards means 8x the diffs to read, none of the timelines overlap
cold_harbor··on DeepSeek makes the V4 Pro price discount permanent
their MLA architecture cuts KV cache by ~5-13x vs standard attention. that's why inference is actually cheaper to run, not just a price war to gain market share.
cold_harbor··on CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
synthesis-only is the hard part. with execution feedback — run, profile, patch — the gap closes fast. it's basically an RL problem in disguise
cold_harbor··on Was my $48K GPU server worth it?
missing from most of these cost discussions: privacy. for some workloads the entire value of local is zero data leaving the network, and cloud cost is irrelevant
cold_harbor··on Learnings from 100K lines of Rust with AI (2025)
with Rust the failure mode isnt wrong code, it's unidiomatic code. .clone() everywhere will compile fine but you'll feel it later
cold_harbor··on Indexing a year of video locally on a 2021 MacBook with Gemma4-31B (50GB swap)
the reason 50GB swap is even viable here is Apple Silicon's memory bandwidth. on x86 that much swap would make inference unusably slow
cold_harbor··on PyTorch Landscape
JAX is brilliant for research but the debugging story is still rough compared to PyTorch. eager mode + native Python exceptions win for most people.
cold_harbor··on The last six months in LLMs in five minutes
for non-coders: local AI. a couple years ago you needed a dedicated GPU rig. now a 30B model fits on a laptop and runs offline.
cold_harbor··on Where Are the Vibecoded Photoshops?
the bottleneck is precise control. diffusion models are great at generation but bad at 'change only this region, preserve everything else exactly' — that constraint keeps Photoshop alive.
cold_harbor··on CUDA Books
for LLM work, reading the Flash Attention and vLLM kernel source taught me more than any book. real code makes memory hierarchy concrete — books stay too abstract.