135 karma · joined July 17, 2026
It's been a while since I read up on them and things might have changed, but I have a lot respect for what they're doing. Full embracing right to repair and customer privacy when the entire industry is moving at lightning speed in the opposite direction.
I actually built my own tool that maintains a work graph (DAG-like) with task leases. It allows me to copy and paste pre-written prompts into Claude Code, Codex, or OpenCode and all the agents self-coordinate through MCP calls.
I built this after trying hermes and Openclaw but not liking the lack of human-in-the-loop judgement. So I'm wondering if I should keep refining my tool, or evaluate something like pi?
Unticking the WebGPU checkbox and reloading didn't do anything.
That's typical executive lack of awareness of the skills and roles of the little people charged with fulfilling their grand vision.
Most days I feel like I must be the dumbest person on HN. I don't understand what 80%+ of submissions are about. But if it sounds interesting, I'll dig into it a bit and learn a few things along the way.
I don't know why this has to be reinvented?
I wholly agree with your comment, but is it legally "your code"? Copyright is implicit at the moment of human creation. But there isn't yet settled law on AI-assisted creation.
So it might be a problem for projects to accept contributions where it's not clear who actually owns that work.
My wife intentionally overpaid on a credit card statement balance in order to gain some credit limit headroom in the current month. Using made up numbers: The balance said we owed $1,000, but she paid $2,000 to make room for a big purchase that month. My application rejected that overpayment as a data integrity error because it would have pushed the credit card balance to -$1,000.
This is a valid state though and Reg Z actually specifies the rules around that case. A bank has to refund a positive balance upon request or automatically after X number of days (I forget the amount).
So now, anytime my agents touch any code associated with credit instruments, they have to run a Reg Z audit to ensure that the data model reflects how banks actually operate in the real world.
Seems like a weird thing to post on a status page. Shouldn't this have happened automatically and therefore precluded the need to inform users of it?
I'm on my third iteration, after constantly log jamming previous versions. Most steps aren't "run pytest", they're gates that guard against an LLM's bias to continuously add more and more complexity to a system.
I'm aware of the irony of having a complex system to mitigate complexity, but the key difference is these rules bound complexity growth. If you're legitimately interested in the details, let me know. I'm too tired to write up much more but would be willing to drop in a LLM-authored summary of the details.
I've been designing software for a long time, but I'm nowhere near as good as most career SDEs I know (my career path has been SDE-adjacent). So it's not like I sat down and independently told various LLMs how to build out all these guardrails. I make high level architecture decisions and nudge them in the right direction ("use RabbitMQ", "trunk-based branching, not gitflow", etc).
A lot of this stuff evolved piecemeal and organically. But at no point was a churning out garbage and I never had a runaway agent completely derail the project (or my budget). But, thinking about it more, I guess there are some things I might have done differently than most people:
- I started with documentation: user interview --> user stories --> functional spec --> frozen design contract. These were all done before I wrote any code.
- I specified the tech stack and the architecture in broad strokes, rather than let the LLMs make that decision. I went with boring choices because that's what I know best: Flask/Jinja, Alpine.js, Postgres.
- I've constantly gone back and refactored accumulated tech debt and have added hard CI gates for things like cyclomatic complexity, ensuring that docs don't drift from the underlying code, and an "apparatus ledger" that keeps track of all the rules and constraints that keep getting added.
- I make sure that each session proves that it's tests can fail before shipping a PR, so it's not writing meaningless tests.
Maybe I'm underestimating how impactful all those things add up to shape the behavior of the LLM agents? Because individually, I wouldn't expect them to have saved my from nearly all the AI pitfalls I read about.
In my defense though, it does way more than a spreadsheet. Stuff like OCRing screenshots of bank transactions with a specialized, locally-hosted LLM to avoid data harvesters like Plaid. This became an entirely separate subsystem with verification, automated model benchmarking, prompt provenance, etc.
You'd be surprised at how quickly edge cases start to pile up when an accounting system makes contact with the real world. (If you buy something on a credit card and then return it after your statement closes but before your payment is due, do you still owe a minimum payment based on that purchase? Well... depends on your bank. Capital One and Chase: yes, US Bank: no.)
> Just how much functionality are you getting out of that?
I'm still dogfooding it. It's a pretty opinionated app that has things a month-end closing ceremony, reconciliation processes, envelope-based budgeting cycles. So unfortunately my feedback cycle is largely locked to the calendar. But my wife absolutely loves it so far.
My agents have never created code that is straight-up garbage and I have never flushed a week's worth of tokens down the toilet. I just can't identify with all the constant complaints about AI-assisted coding.
And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline. I'm doing 24x7 mutation testing on a dedicated box against the accounting engine and temporal systems. I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks.
Most of my complaints about everything are nits, like the overly verbose and dense way LLMs communicate with me. Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).
This is the inherent problem with tech start-ups. Once they solve the problem they were founded to solve, then what? It's feature complete, but investors insist that the line must go up. Can you imagine if `sed` or `vim` was brought to market by a venture-backed company?
IMHO, capitalism and FOSS are fundamentally incompatible.
There's also an implicit judgment in your question. Is well-designed, well-tested AI-generated code worse than poorly written, but hand-typed code?
They're struggling to maintain a single nine these days.
Unit and integration tests serve different purposes and one isn't necessarily better than the other.
Their hosted, API-based service is something like a third of the cost of this model.