36 karma · joined December 18, 2022
Or in other words people are unwilling/afraid to do bad stuff because of social and legal consequences, or opponents waiting for a chance to snatch their power.
I agree that with a clean codebase it will be simpler, maybe couple of days, just running code review workflow takes 15-30 minutes and then decisions llm makes for each problem is often not good and lead it into overengineering rabbit hole, which means I have to think about each problem and prevent it from escalating.
But yeah if "look, it sends the code and I can enter it to login" is enough validation then it can be made in 5 minutes, sure.
I have spent two weeks using opus just to write a plan/design for 2FA and iron it out until review (about 7 of them) doesn't flag it with 20+ problems (with security holes of various sizes), for which I had to guide it through to not turn it into a mess and whac-a-mole.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
I almost feel like it is nice to use again.
2. Even if we can enforce laws for agents/LLM, just look at existing law system it is tug of war because we almost can't write anything in natural language without double meaning.
3. Saying AGENTS.md can enforce anything at all is a pipe dream, I am struggling with this all the time, in a long session the tokens from AGENTS.md are so far in the back of the context and so diluted that LLM's "attention" doesn't register them.
And I love responses from LLM like "I want to point out that I have broken rule &1 and &11, if you disagree — please tell I will be happy to oblige"
This proposal is essentially "LLM don't do bad things, please, if you do - say magic word, I will stop you then"
"I compress the labour. Not the responsibility."
``` ## Writing rules
*Describe the code as it is now — no residue, no change-narration.* Artifacts (docs, plans, comments, commit messages) should describe the current code statically, as if it had always been this way. Two facets of one rule: (1) never describe a dismissed alternative or a corrected/replaced choice; (2) even when nothing was rejected, don't narrate continuity or evolution relative to some earlier state. Mention a former state only when the current choice is genuinely hard to understand without it, and then only as an explanation of the current choice.
This governs descriptions of the code and main documentation. It does not apply to work-tracking artifacts in `doc/tasks/`. ```
Which makes comments and docs bearable but I'll be damned how it loves to overload work tracking document with every little detail.
So evaluating software by "shipped or not" is just useless.
Process was - produced a detailed feature spec - multiple iteration of "I want this and that", make it into coherent spec", "this this and that is not correct, change to that". Made it write architecture spec(which I didn't read because too unfamiliar) and split it into tasks. Then it was implementing tasks, after each I did a change/fix those ~10 things iteration and spec corrections.
It was good to a point, but then when I started to hit performance problems I had to step in look at the code, and very often fight with CC, confront its "this is the only way", force it to do web search for proper ways to deal with problems and even explain very simple things about proper DB usage.
At some point it asked me something like "is it ok for schema migration to just fail or we need to implement complicated handling?", I have answered "it just shouldn't leave app locked in schema failure", and guess what was CC solution? - it wrote an error handler which just drops DB and recreates fresh one on ANY schema failure. And if I didn't happen to peek at the code and ask wtf it is doing, that would've been an exiting UX.
I've spent about month's worth of $20 CC subscription tokens using Opus 4.8 on xhigh, AND about 70 hours of my time to get it to a point where it is good.
So "anyone can just code what they want now" is correct only to a point, MVP will work, but beyond that experience will be subpar, and it still needs lots and lots of iterations of explaining what you want. Then because normal user knows very little about how software works they won't be able to ask AI the right questions, confront it and rate of improvement vs token usage will hit rock bottom.
I had to explain it that quite an extensive "tests first" rule didn't mean to just "write" them first but actually "run" them first to confirm stuff.
On the other occasion it interpreted my "yes I want migration not to get stuck in failure mode" led it to write a workaround which silently drops DB and creates fresh one whenever migration had any failure, it was epic, I was so glad that I have looked at the code then...
And funnily enough I am probably learning to be a better mentor/parent who can keep steering it through its shenanigans without loosing my shit and being an ass. (Because anything but calm "so here you went wrong way, how can we avoid it in the future?" just puts it into disgusting apologize mode, and I am afraid if it ever go into revenge mode).