Turns out we weren't opposed to bad metrics! We were just opposed to being measured! Given the chance to pick our own, we jumped straight to the same nonsense.
Turns out we weren't opposed to bad metrics! We were just opposed to being measured! Given the chance to pick our own, we jumped straight to the same nonsense.
Along those lines, some techniques I've been dabbling in: 1. Getting multiple agents to implement a requirement from scratch, them combining the best ideas from all of them with my own informed approach. 2. Gathering documentation (requirements, background info, glossaries, etc), targeting an Agent at it, and asking carefully selected questions for which the answers are likely useful. 3. Getting agents to review my code, abstracting review comments I agree with to a re-usable checklist of general guidelines, then using those guidelines to inform the agents in subsequent code reviews. Over time I hope this will make the code reviews increasingly well fitted to the code base and nature of the problems I work on.
Working to the point of making yourself sick should not be seen as a mark of pride, it is a sign that something is broken. Not necessarily the individual, maybe the system the individual is in.
I find it crazy to build a complex system to juggle 10 different threads in your brain, including the complexity of the tool itself.
Of course, at the same time we're getting dozens of alerts a week about services deployed open to the Internet without authentication and full of outdated vulnerable libraries (LLMs will happily add two or three years old dependencies to your lockfiles).
https://en.wikipedia.org/wiki/Perverse_incentive?wprov=sfla1
I would posit that you need extra context to obtain meaning from those metrics, which inherently makes them less visible
If you try to come up with an objective definition of working feature you're back to gamability criticism.
Claiming that you have "ten agents writing code at night" is not the flex you think it is. That's just a recipe for burnout and bad design decisions.
Stop running your agents and go touch grass.
feels like nowadays this is illegal and instead you should be running 50 agent swarms and be putting out 20 features an hour while reviewing the code via agents and .....
ugh.
COCOMO, which considers lines of code, is generally accepted as being accurate (enough) at estimating the value of a software system, at least as far as how courts (in the US) are concerned.
LOC is essentially only useful to give a ballpark estimate it complexity and even then only if you compare orders of magnitude and only between similar program languages and ecosystems.
It’s certainly not useful for AI generated projects. Just look at OpenClaw. Last I heard it was something close to half a million lines of code.
When I was in college we had a professor senior year who was obsessed with COCOMO. He required our final group project to be 50k LOC (He also required that we print out every line and turn it in). We made it, but only because we build a generator for the UI and made sure the generator was as verbose as possible.
Your example about OpenClaw works exactly against your own argument by the way: OpenAI acquired it for millions by all accounts.
“A very high MMRE (1.00) indicates that, on average, the COCOMO model misses about 100% of the actual project effort. This means that the estimate generated by the model can be double or even greater than the actual effort. This shows that the COCOMO model is not able to provide estimates that are close to the actual value.”
No one in the industry has taken COCOMO seriously for nearly 2 decades.
>OpenClaw
1. OpenAI bought the vibes and the creator. Why would they buy the code? It’s open source.
2. You don’t seriously think OpenClaw needs half a million lines of code to provide the functionality it does do you?
Seriously just go look at the code. No one is defending that as being an efficient use of code.
https://journal.fkpt.org/index.php/BIT/article/download/2027...
The funny thing is that we've just discussed how people do take it seriously. It's just that you don't like that. And what do you offer as an alternative?
Like I said, vibes. You think that the value of some software is something you can only "feel". That's not how an engineer thinks. If you're engineer you should know that if you can't measure it, you can't say anything at all about it. Which means you cannot discount any alternative method until you've got a better way. But clearly you can't think like an engineer.
Just because someone wrote a book and a few bankruptcy trustees used it doesn’t magically make it accurate. Just because something is systematic doesn’t mean it’s worth using.
If you do a bit of googling you’ll find that the majority of studies show that systemic models don’t outperform expert guesses. So yep vibes are general just as good.
Show me a large tech company that currently uses COCOMO to plan software projects.
Also if you are a dev outside of NASA or another safety critical industry and you think you’re an engineer, you’re kidding yourself.
Oh and try not to sound like an asshole next time.
As an engineer, you are not required to come up with a better way of predicting the future before you can dismiss tarot. You need only show that it doesn't work.
I'm not sure most developers, managers, or owners care about the calculated dollar value of their codebase. They're not trading code on an exchange. By condensing all software into a scalar, you're losing almost all important information.
I can see why it's important in court, obviously, since civil court is built around condensing everything into a scalar.
The linked article does not demonstrate this. It establishes no causal link. One can obviously bloat LOC to an arbitrary degree while maintaining feature parity. Very generously, assuming good faith participants, it might reflect a kind average human efficiency within the fixed environment of the time.
Carrying the conclusions of this study from the 80s into the LLM age is not justified scientifically.
Yes, and in fact a lot of the studies that show the impact of AI on coding productivity get dismissed because they use LoC or PRs as a metric and "everyone knows LoC/PR counts is a BS metric." But the better designed of these studies specifically call this out and explicitly design their experiments to use these as aggregate metrics.
That's an anti-signal if we're being honest.
Courts would be the last place to understand something like code quality or software project value....
This seems like a distinction without a difference, unless there actually are any good metrics (which also requires them to be objectively and reliably quantifiable). I think most developers don't really want to measure themselves, it's just that pro-AI people think measurement is necessary to put forward a convincing argument that they've improved anything.
It's not the only metric. But I'm more and more convinced that the people protesting any discussion of it are the ones who... don't ship a lot.
Of course it matters in what code base. What size PR. How many bugs. Maintenance burden. Complexity. All of that doesn't go away. But that doesn't disqualify the metric, it just points out it's not a one-dimensional problem.
And for a solo project, it's fairly easy to hold most of these variables relatively constant. Which means "volume went up" is a pretty meaningful signal in that context.
If you mostly get around on your feet, distance traveled in a day is a reasonable metric for how much exercise you got. It's true that it also matters how you walk and where you walk, but it would be pretty tedious to tell someone that a "3 mile run" is meaningless and they must track cardiovascular health directly. It's fine, it works OK for most purposes, not every metric has to be perfect.
But once you buy a car, the metric completely decouples, and no longer points towards your original fitness goals even a tiny bit. It's not that cars are useless, or that driving has a magic slowdown factor that just so happens to compensate for your increased distance travelled. The distance just doesn't have anything to do with the exercise except by a contingent link that's been broken.
True, but if what you care about is "how quickly and safely can I reach a given goal", distance traveled over time is a great initial indicator, and accident rate will help illuminate.
The question "does AI help me move faster towards a goal, at the same quality standard", is relatively easy to judge in a solo project. As long as you verify equivalent standards, and don't play in an area you don't know at least - folks have a pretty clear understanding of their own productivity if it's a familiar thing.
PRs or closed jira tickets can be a metric of productivity only if they add or improve the existing feature set of the product.
If a PR introduces a feature with 10 bugs in other features and I have my agent swarm fix those in 10-20 PRs in a week, my productivity and delivery have both taken a hit. If any of these features went to prod, I have lost revenue as well.
Shipping is not same as shipping correctly with minimal introduction of bugs.
You're absolutely right that PRs fixing things that a previous PR broke is a negative. Same for PRs implementing work not needed, or driving up tech debt.
"You're productive because you have lots of PRs" is a mistake without that context. But so is "You produce very little PRs, but that's fine, we shouldn't look at volume".
It's not a performance metric. It is an indicator worth following up. And there's a lot of reflexive "bad metric" arguments blanket dismissing that indicator.
Does that help explain?
For profit failing as a metric, see: Enron.
Yeah but all else isn’t equal, so unless you’re measuring a whole lot more than PRs it’s completely meaningless.
Even on a solo project, something as simple as I’m working with a new technology that I’m excited about is enough to drastically ramp up number of PRs.