Disclaimer: I'm the author of the app.
2,790 karma · joined January 20, 2009
Disclaimer: I'm the author of the app.
Here's a good overview of the problem: https://arxiv.org/abs/2206.02658v3
Strategically speaking, I think this only makes sense as an anti-NVIDIA play.
https://developer.apple.com/documentation/uikit/uiapplicatio...
if it looks like a hack, walks like a hack, and quacks like a hack...
> Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
IMO, asymptotic advantages are marginal. At least for coding, we got a glimpse into how much of an advantage it gives (or doesn't) when the Claude Code codebase leaked[1] ~4 months ago. :)
I suppose the road to technical hell is paved with marketers and grifters. :)
[1]: https://ollama.com/library/deepseek-r1:1.5b [2]: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-...
We've had so many advancements in LLM samplers for improved text generation (off the top of my head: min-P, adaptive-P, XTC, DRY, p-less, Top-H, Top-n-Sigma, and so many more) but hosted LLM APIs only provide three basic knobs: temperature, top-k and top-p which are old as the mountains in LLM years at this point.
One thing that doesn't help the local LLM case is that all the popular VC backed local LLM wrappers also only support the same three ancient knobs because I suppose they're more preoccupied with their next fundraise than with keeping up with the advances in tech.
Does it have to be? There are plenty of coding tasks, where it's good enough.
The #1 post on HN right now[1] is full of people jubilating about how they can run Qwen 3.8 27B on their > 5 year old GPUs. If that isn't democratization of AI, I don't know what is.
I'm sure he's smart enough to instantaneously realize this too, but as the famous Upton Sinclair quote goes, he won't mention it even if he does.
[1]: https://xcancel.com/magikarp_tokens/status/20878591737488549...