Now I have persistent chats and a persistent environment that I can talk to from any device, including my phone. Even have X11 + CUA + Chromium inside the VM so the agent can use a real browser for sites that require it. DSH bwrap is the first layer of containment, VM is the second layer. I can swap between API models and local inference with a drop-down in the UI.
I realise this is way more setup than Meta's customers would tolerate, but the HN crowd could slap something together quite quickly. Using a coding agent with a nice chat UI as your general chat client is surprisingly smooth: just create an empty workspace.
I think Muse is the right shape in a lot of ways but I come unstuck at the point where my personal details and credentials are inside the VM.
I host all of my custom applications on a machine in my home with Tailscale, k8, infrastructure as code
It’s wired, always on, has access to my gpu running on another machine or falls back to cloud inference
Self hosted GitHub action runners on each different platform for build/deploy of various projects: tauri v2, apis, web uis, native macOS apps, Unity 3d
None of this existed a month ago and it’s been a blast pretending to run my own platform
The other advantage long running sessions have is that very low tok/s are perfectly acceptable. In fact, hardly noticeable.
(aside, of course, from Google and Apple both having incentives to prevent this from happening, as both depend on selling centralized services to you and advertisers, and the local model breaks both)