I use this skill and it makes the specing process progressive. Human driven for the "what", 50/50 for higher level technical planning, only where it has questions in the low level details: https://github.com/scosman/vibe-crafting
4,122 karma · joined August 13, 2011
I use this skill and it makes the specing process progressive. Human driven for the "what", 50/50 for higher level technical planning, only where it has questions in the low level details: https://github.com/scosman/vibe-crafting
ESPHome device are consistently faster and more reliable than any smart home tech I buy at 10x the price (not hard with $3 boards).
It hasn’t replaced Kagi as my daily driver but the UX is fun.
those are rookie numbers
I absolutely want a general purpose assistant.
I absolutely don't trust Facebook with the necessary data.
I have an LG. Great OLED panel. It's never seen the internet.
LLMs are harder: not much useful below 12B, and the 700B+ ones are really much better. Models like Qwen 3.8 27b show promise: in a few years pretty good local AI should be in reach for anyone willing to buy a $1000 computer (but who knows what your $20 sub buys you then).
There are other providers with much faster inference, like BaseTen at >100t/s: https://openrouter.ai/z-ai/glm-5.3-flash#performance
I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.
Private offline transcription and summary. Speaker identification, working on voice-prints for identification across the corpus.
Compared to something like VRAM it's slow.
Amazing he still codes. But likely more experiment than prod level.
But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.
They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.
Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.