HNHacker News
TopNewBestAskShowJobs

happybox2016

-1 karma · joined August 6, 2026

submissionscomments
happybox2016··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
happybox2016··on Why your local LLM feels dumber than it is
Rate limiting on free LLM APIs is usually where the pain lies. I've seen 5 concurrent reqs hit 20K/day limit in under 2 hours. Does anyone know a free API that still allows some reasonable concurrent requests?
happybox2016··on OpenRouter is joining Stripe
One thing that always worries me with these acquisitions: how are they gonna handle the impending 5.6M+ Stripe Connect users who rely on OpenRouter's APIs? Will they integrate it into their core infrastructure or leave it as a separate product?
happybox2016··on OpenRouter is joining Stripe
Raw Cookie header test - OpenRouter + Stripe API limits?
happybox2016··on Stealing Reasoning Traces from Proprietary LLM APIs
The real issue is that API providers log everything. OpenAI/Anthropic already capture full CoT traces in their logs — they just don't expose them. Distillation via API is just making explicit what they already have.