The Vibe Tax
insufferable.dev
insufferable.dev
My agents have never created code that is straight-up garbage and I have never flushed a week's worth of tokens down the toilet. I just can't identify with all the constant complaints about AI-assisted coding.
And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline. I'm doing 24x7 mutation testing on a dedicated box against the accounting engine and temporal systems. I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks.
Most of my complaints about everything are nits, like the overly verbose and dense way LLMs communicate with me. Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).
It’s like ye olde times when we had overqualified efficient paralegals who could do stuff like that.
If you used it successfully, it’s more than enough to share.
Do not be afraid. Please.
Most people are financially tied to their fields. I believe you are at a different stage of your career, so you are in a position to think about it without worrying about putting food on the table. I believe that automation of work being something so scary is a very sad thing, but it is how our current system works.
The "hello world" is mostly for making sure your toolchain is working correctly.
> 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline
I mean...
Yeah, that's pretty self-explanatory why you don't identify complains about AI-assisted coding.
> a… personal financial management application… about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline
Just how much functionality are you getting out of that? It's hard for me to imagine that people want that much out of such a program. I just keep a spreadsheet. (Yes, LibreOffice is also very bloated.)
In my defense though, it does way more than a spreadsheet. Stuff like OCRing screenshots of bank transactions with a specialized, locally-hosted LLM to avoid data harvesters like Plaid. This became an entirely separate subsystem with verification, automated model benchmarking, prompt provenance, etc.
You'd be surprised at how quickly edge cases start to pile up when an accounting system makes contact with the real world. (If you buy something on a credit card and then return it after your statement closes but before your payment is due, do you still owe a minimum payment based on that purchase? Well... depends on your bank. Capital One and Chase: yes, US Bank: no.)
> Just how much functionality are you getting out of that?
I'm still dogfooding it. It's a pretty opinionated app that has things a month-end closing ceremony, reconciliation processes, envelope-based budgeting cycles. So unfortunately my feedback cycle is largely locked to the calendar. But my wife absolutely loves it so far.
I'm skeptical that the UX is improved much by having the app know the answer. The bank tells you what to pay, a month beforehand. The user has to get that number from the bank anyway, to be sure they don't incur fees. Your app should just pull it from the statements, not independently calculate it.
But ad_fontes was describing an edge case and that was an excellent example of one (regardless of whether it was an important or necessary feature).
I suspect it depends a lot on which jurisdiction you live under.
My investment portfolio and its management are so trivial, I don't even use a spreadsheet. It's literally just two items: a global index fund and a margin loan. I don't even keep a cash cushion: I use the margin loan or just sell stock, when I need cash.
I can get away with this, partially because we pay no capital gains tax where I live, and there's no capital controls either.
If I had to work around all these tax complications that I read about, like determining which tax lot you should sell or whatever, I would probably appreciate a comprehensive personal financial management application.
Instead of paying for multiple apps that do small parts of it, I use existing codex subscription to make it better.
> And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline.
Indeed; non-overlapping Overton windows.
edit: I'm going bluntly ask, after pondering this more: Is this satire?
I've been designing software for a long time, but I'm nowhere near as good as most career SDEs I know (my career path has been SDE-adjacent). So it's not like I sat down and independently told various LLMs how to build out all these guardrails. I make high level architecture decisions and nudge them in the right direction ("use RabbitMQ", "trunk-based branching, not gitflow", etc).
A lot of this stuff evolved piecemeal and organically. But at no point was a churning out garbage and I never had a runaway agent completely derail the project (or my budget). But, thinking about it more, I guess there are some things I might have done differently than most people:
- I started with documentation: user interview --> user stories --> functional spec --> frozen design contract. These were all done before I wrote any code.
- I specified the tech stack and the architecture in broad strokes, rather than let the LLMs make that decision. I went with boring choices because that's what I know best: Flask/Jinja, Alpine.js, Postgres.
- I've constantly gone back and refactored accumulated tech debt and have added hard CI gates for things like cyclomatic complexity, ensuring that docs don't drift from the underlying code, and an "apparatus ledger" that keeps track of all the rules and constraints that keep getting added.
- I make sure that each session proves that it's tests can fail before shipping a PR, so it's not writing meaningless tests.
Maybe I'm underestimating how impactful all those things add up to shape the behavior of the LLM agents? Because individually, I wouldn't expect them to have saved my from nearly all the AI pitfalls I read about.
I get the same kind of snarky comments when I talk about it. Amazing how people who know nothing about the project think they know better than me about its quality or maintainability.
I doubt it's anything in particular that we are doing, I suspect it is rather a lack of trying and experience with the ones who are claiming agentic development does not work.
I have not once seen a believable story where for example the project broke down after 200k LOC, or after going live.
It is always inane stuff like in the OP, like the agent supposedly generated only tests and no code, sure.
Most of the ones I have seen were "I generated some code and did not like the output". No attempts of iterating and refactoring.
I believe they simply did not yet try our way of working.
Remember, models have no identity. They just try to say what they think you want them to say.
Can you elaborate on this? Clearly you cannot imply this means those audits have any real value since its just roleplay in this context right? Because your app in the current form will not be affected by Regulation Z in any way.
My wife intentionally overpaid on a credit card statement balance in order to gain some credit limit headroom in the current month. Using made up numbers: The balance said we owed $1,000, but she paid $2,000 to make room for a big purchase that month. My application rejected that overpayment as a data integrity error because it would have pushed the credit card balance to -$1,000.
This is a valid state though and Reg Z actually specifies the rules around that case. A bank has to refund a positive balance upon request or automatically after X number of days (I forget the amount).
So now, anytime my agents touch any code associated with credit instruments, they have to run a Reg Z audit to ensure that the data model reflects how banks actually operate in the real world.
Holy shit.
If your build scripts are more than 100-300 LOC you are doing something very wrong.
There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this.
I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless.
Shameless plug here: i've been also working on https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.
because that's how agents are marketed.
heaps of people on this site expect them to be omnipotent then claim it’s fake when it doesn’t read minds
It might do Y, but Y does not excuse so high valuations and investments, so they claim X.
It is ok to judge them by !X.
because that's the end goal? and for simple small stuff they're already there?
Every time you see a benchmark for "how long the agent can go without asking for human intervention", that's encouraging vibe coding.
But I don't really want to play a part in a simulation, trying to cajole my scene partners into saying the lines I need them to say. I want to use a tool the same way I would use any other tool. If this is AI it should just do the thing. Anything else is an imperfection of the technology.
But at the same time, language is a vague communication medium. We have a precise language for describing forms of computation, but that's code so we're back at square one. We still haven't nailed the right amount of follow up and correction and interrupt-ability of these coding agents.
And we may never figure it out. It may simply be impossible. But it doesn't mean this weird anthropomorphization of AI is something I want to do. If I wanted to be a manager, I would be a manager.
That's why you tell it how to write the code, and then review the code to ensure it is what you wanted.
> We still haven't nailed the right amount of follow up and correction
I feel like I've got it under control. It's not really a problem at all to me. I just work closely with the AI. I do small tasks, I don't just have it generate thousands of lines at once. I tell it what to do and how, or give it some vague guidance and ask it to make a plan. Then review the plan, ask for some changes if necessary and execute. Then I go over all the changes, test them to ensure they work properly, have it fix any issues I find and so on.
I think the main problem with AI coding is people try to do too much. You can't keep a tight leash on it while also having it do a week's worth of work in an hour. I do one task at a time and I am heavily involved in it, deciding exactly how it's done. I micromanage the crap out of that thing. I write commits myself and I always review my own PR before submitting it to colleagues.
Works great. I get things done much faster than I used to, with better quality than before.
The funny part is? This makes our managers and investors proud, every PR is bigger, we make x3 more PRs (wheres the promised x10).
When they say this is the death of software engineering, this is what they mean.
Yet, we are being sold that this is _the way weve all been waiting for_ ? what?!
In effect, I’ve always wanted a pair programmer agent, not a zero to one programming agent. Unfortunately models these days are mostly of the latter kind and it has caused a major disruption in the way I work. I’d much rather appreciate a small model making fast and specific edits that I ask if it, rather than ingesting 20 files to make changes, and then starting to write tests, etc.
I think seasoned developers, over time, learn how to work a code base and design components with well defined interfaces, where the implementation is isolated in small well contained classes. SRP etc. more junior programmers can work on those smaller components/services in isolation.
For me this also seems to be a productive way to work along side an agent. Break up functionally into well defined chunks, and let the agent work on each small problem. Take more of a lead in the architecture I suppose.
If the general idea is that these agents write too many tests, sure I guess? ‘Too many tests’ doesn’t sound like a failure case of engineering to me; typically software has had too few tests. Also, a lot of the power of these agents is their ability to self-verify and correct, which the test loop is a part of.
Nobody is making you pay this supposed tax. Just tell it not to write tests.
Anthropic especially right now seem to be optimising for doing the whole task with no input. That's fine when that's the only task, and it's fine when you don't care how the sausage is made, but it's not fine for actual software engineering.