HNHacker News
TopNewBestAskShowJobs

briansun

4 karma · joined August 11, 2023

Founder & CEO, Xmind, NativeMind, Quin.
submissionscomments
briansun··on Ask HN: Why aren't local LLMs used as widely as we expected?
Thanks for the view from a very privacy‑sensitive environment — agreed that hosted SOTA still leads on broad capability.

Could you share a quick split: which tasks truly require hosted SOTA than open‑weight? I think gpt-oss is quite good for a lot of things.

SMBs can’t get enterprise contracts with OpenAI/Anthropic, so local/open‑weight may be their only viable path — or wait for a hybrid plan.

briansun··on Ask HN: Why aren't local LLMs used as widely as we expected?
Wouldn't it be cool to have a local AI agent? It could access search engines and browse any website through a headless browser.
briansun··on Ask HN: Why aren't local LLMs used as widely as we expected?
Thanks — I agree with your three big pain points: quality vs hosted SOTA, token speed, and economics/utilization.

Have you run into cases where on‑device still makes sense?

1. Data that is contractually/regulatorily prohibited from being sent to any third‑party processor (no exceptions).

2. Very large datasets where throughput can be low (overnights acceptable) but the cost is high for cloud models.

3. Inputs behind a password-wall that hosted assistants/chatgpt/claude can’t reach and can't do agentic things with them.

briansun··on Ask HN: Why aren't local LLMs used as widely as we expected?
Well put. Management overhead + unclear capacity planning kills many pilots.
briansun··on Ask HN: Why aren't local LLMs used as widely as we expected?
Totally fair. On a normal laptop you also need headroom to do your actual job, and KV cache + context length can eat that quickly.
briansun··on Ask HN: Are you running local LLMs? What are your key use cases?
Thanks for raising the privacy angle. Do you have a source or plan details for the 30‑day retention and the lack of deletion options (non‑enterprise)? It would help to know account tier and where that policy is documented.

Beyond policy, how are you actually using local LLMs—what tasks do you run locally vs. in the cloud—and which scenarios feel most privacy‑sensitive to you (e.g., proprietary code, contracts, health notes)?

briansun··on Ask HN: Are you running local LLMs? What are your key use cases?
Gemma3n as a daily driver sounds nice—4b or 8b? and rough tokens/sec on your laptop? And have you A/B‑tested code generation quality across local models (e.g., Gemma3n vs others)?
briansun··on Ask HN: Are you running local LLMs? What are your key use cases?
Thanks-this is genuinely encouraging; I'd assumed AI help was strongest on front-end work(web apps/SwiftUI), so this is my first concrete example of an LLM catching memory‑unsafe C/C++-could you share your toolchain (CLI/IDE integration) and model details (name/quant/runtime), and what "awesome" 100% on‑device Windows/Mac apps you most want to see?
briansun··on Ask HN: Are you running local LLMs? What are your key use cases?
Super useful config dump—thanks. Do you have wall‑clock numbers for prefill/gen tokens/sec and power draw on the 24GB card for those three setups? Also curious where quality starts to degrade vs. context length in your tests.
briansun··on Ask HN: Are you running local LLMs? What are your key use cases?
Haha, a cute pet dragon. Two knobs that helped me tame VRAM: KV‑cache quant/eviction and sliding‑window attention (if your runtime supports them). What model/runtime and context are you running when it tips over? Are you using Ollama?