The Attention Rebuild: How 2026's Open Models Made 1M Context Fit on Machinesvettedconsumer.com·2 pts·ermantrout·0
The year Claude users sued over limits is the year the limits mostly went upokaneland.com·3 pts·ermantrout·0
Speculative Decoding: The Free Speed Toggle Your Local LLM Is Probably Not Usingvettedconsumer.com·4 pts·ermantrout·0
The open-source Claude Design alternative has 77k GitHub stars and 1M+ installsokaneland.com·7 pts·ermantrout·3
What a ChatGPT citation is worth, and which answer engine a small site can winokaneland.com·1 pts·ermantrout·0
Kimi K3: The Largest Open Model Ever, and Why Almost No One Can Run It Locallyvettedconsumer.com·2 pts·ermantrout·0
Wrap an LLM, charge $20, lose money: the pricing trap the model owners hit firstokaneland.com·2 pts·ermantrout·0
AI takes two-thirds of venture money, and your odds are still one in sixokaneland.com·2 pts·ermantrout·0
Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can'tvettedconsumer.com·74 pts·ermantrout·58
AI saves about 3% of your hours, and almost none of it reaches the moneyokaneland.com·77 pts·ermantrout·94
Qwen-AgentWorld-35B-A3B: a local 'world model' you can run at home Open Modelsvettedconsumer.com·2 pts·ermantrout·0
AEO and GEO: one real study, a pile of mythology, and a traffic cliffokaneland.com·4 pts·ermantrout·1
AI coding: faster MVP, slower review, and the security bill nobody mentionsokaneland.com·2 pts·ermantrout·0
GLM-5.2: The Most Powerful Open Model yet and the Brutal Reality of Running Itvettedconsumer.com·43 pts·ermantrout·28
Prompt processing vs. generation: two phases, opposite bottlenecksvettedconsumer.com·2 pts·ermantrout·0
Show HN: Quant Picker – which GGUF file fits your model and machinevettedconsumer.com·20 pts·ermantrout·0
Mixture-of-Experts (Moe), Explained: Why "Active Parameters" Decide What Runsvettedconsumer.com·2 pts·ermantrout·0
GGUF vs. GPTQ vs. AWQ: The Plain-English Guide to LLM Quantizationvettedconsumer.com·2 pts·ermantrout·0