Show HN: Rebuilding my journaling app to run on-device
tryweave.app
tryweave.app
Stack: Llama 3.2 3B (Q4_K_M) via llama.rn, Whisper small.en (q5_1) via whisper.rn, MiniLM sentence embeddings for semantic memory, everything in a local SQLite file that never syncs. Metal GPU offload on iOS, CPU-only on Android.
I ran the same reflection prompts through Qwen 2.5 3B and Gemma 2 2B first. Llama 3.2 3B was least-bad at "notice a pattern without sounding like a fortune cookie." Details in the essay.
Numbers from device tests (means over 20 sequential entries):
- iPhone 17 (A19 Pro, Metal): save-tap → complete reflection = 11.6s, first token at 7.6s - OnePlus 9R (Snapdragon 870, CPU): 59s complete, 46s first token - 100% pipeline reliability on both, 100% JSON extraction success - Everything runs in airplane mode after the one-time model download
Two things I'd love pushback on:
1. Prompt engineering for a 3B is genuinely different from a 70B — more explicit, less inference. Anyone else worked through that gap? What actually helps beyond few-shot examples?
2. The Android CPU-only situation — is there a llama.cpp Vulkan build path I'm missing, or is it really "wait for llama.rn to add it"?
The essay writeup has the honest limitations (reflection quality gap on a 3B, iPhone 12's 4GB RAM at the memory edge, thermals on Android during a session). Happy to answer anything.