193 karma · joined July 9, 2020
Website: https://prologic.dev/ Twtxt: https://twtxt.net/~prologic Blog: https://prologic.blog/ Git: https://git.mills.io/prologic
It isn't quite the question/answer I was hoping for, but anyway. Thanks!
[1]: https://www.nvidia.com/en-au/products/workstations/dgx-spark...
This is something I had been working on for a while now (few years, even before things like Claude/Codex were considered "mainstream").
It's basically a Linx from Scratch (LFS), with a twist. Every part of the system (except the Linux kernel itself of course) is written entirely in Go. I've done similar things in the past with what GoNix is derived from, called uLinux, and some shell-script based LFS (but widely different then the book).
Anyway, GoNix is pretty neat (IMO), it runs https://gonix.dev itself! And it boots in roughly ~20ms in Vultr Cloud where I run the site from. I'm also using GoNix as the base OS for an appliance I'm building with OTA and A/B updates. So this isn't just a toy, it's quite serious really.
Target market/use-case is appliances and use-cases where you need a system to be quite small and quite fast to boot-up, e.g: Ephemeral CI/CD runners.
Lately though, I've switched tact a bit and have resurrected some ~20yr old research called Behaviour Tree Engineering or Behaviour Trees, so I'm now writing specifications called BT specs.
> There for you, every step of life. > > A galaxy full of paths — we'll find yours. Explore, find your fit, and chart the life you want.
This tells me nothing.
The Skill itself basically is a 6-step process and a simple contract with some guardrails:
(1) identify the alert, (2) size the problem, (3) find the signature in metrics and logs, (4) test the leading hypothesis against a second source until it is confirmed or dead, (5) write the finding + plan, (6) AskUserQuestion for the decision, (7) implement via the sibling skill and verify, (8) update memory if the cause was new.
There'a 3 scripts to automate some of the tasks so that they are deterministic and 3 data sources as references as well as MCP servers it has access to for metric, logs and alerting sources.
That's it. It not only diagnoses the problem every time, correctly, but also figures out what went wrong, and why, and proposes how to fix the underlying causes.
I then have a follow up skill called /incident-postmortem that writes up a full incident report and post-mortem for the next time to learn from and feed back into the system -- Which is basically Claude's own memory over time.
This is quite slever. I also really like the concept of an "Error Budget", inspired by SRE and SLO(s) no doubt :)