HNHacker News
TopNewBestAskShowJobs

kkm

2,208 karma · joined August 2, 2016

submissionscomments

How to serve trillions of tokens for trillion-parameter coding agents

modal.com·1 pts·kkm·
0

Unlocking parallel test-time scaling for long-horizon agents

blog.doubleword.ai·3 pts·kkm·
0

Claude.md is good for taste and project context.It's a weak place for invariants

tesseracted-labs-blog.vercel.app·3 pts·kkm·
0

Moving coding-agent guardrails from prompts to hooks

tesseracted-labs-blog.vercel.app·4 pts·kkm·
0

Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads

j9s.io·3 pts·kkm·
1

Guide to the Kimi DeltaNet Family of linear attention

blog.doubleword.ai·3 pts·kkm·
0

Forensic Analysis of Container Snapshot Chains for Post-Event Reconstruction [pdf]

radostin.io·2 pts·kkm·
0

The Agent swarm that designs itself

peterbhabra.com·1 pts·kkm·
0

Don't Build a Router. Train the Small Model to Know When to Defer

distillabs.ai·2 pts·kkm·
1

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai·306 pts·kkm·
250

InfiniBand, RoCE, and All That

fergusfinn.com·5 pts·kkm·
0

2678x Faster Matrix Multiplication with a GPU

0mean1sigma.com·2 pts·kkm·
0

UCCL-EP: DeepEP-style expert parallelism on any NIC, no GPU-initiated comms

fergusfinn.com·9 pts·kkm·
0

Hacking Google with A.I. For $500k

brutecat.com·1 pts·kkm·
0

How to setup a local coding agent on macOS

ikyle.me·507 pts·kkm·
127

Anatomy of a high-performance EP kernel

fergusfinn.com·16 pts·kkm·
1

No Token Left Behind: Demystifying Token-in-Token-Out in Miles

lmsys.org·2 pts·kkm·
0

MoE expert co-activations: Reordering inputs yields easy throughput gains

blog.doubleword.ai·2 pts·kkm·
0

The Economics of Speculative Decoding

fergusfinn.com·30 pts·kkm·
6

Speculative KV coding: losslessly compressing KV cache by up to ~4×

fergusfinn.com·155 pts·kkm·
48

70x faster cold(ish) starts for SGLang

fergusfinn.com·1 pts·kkm·
0

Bringing Up DeepSeek-V4-Flash on AMD MI300X

fergusfinn.com·120 pts·kkm·
25

Brave AI privacy:LLMs on NEAR AI Nvidia-Backed Trusted Execution Environments

brave.com·1 pts·kkm·
0

How fast can an LLM go?

fergusfinn.com·2 pts·kkm·
0

FHE can be leveraged for LLMs such as ChatGPT in a privacy-preserving manner

huggingface.co·4 pts·kkm·
0

Harnessing the Power of Large Language Models for Insightful Review Analysis

techblog.holidaycheck.com·1 pts·kkm·
0

A Privacy-First approach to use AI for understanding our Customers Better

techblog.holidaycheck.com·1 pts·kkm·
0

How to make LLMs go fast

vgel.me·2 pts·kkm·
0

Leveraging Large Language Models for Sentiment Classification in Hotel Reviews

techblog.holidaycheck.com·1 pts·kkm·
0

Managers Should Think More Like Hackers

hbr.org·3 pts·kkm·
0
Page 1 of 6Next →