- "It hyper-focuses on the current task and couldn't care less if its changes break other parts of the system." - "Long context = instant brain damage"
This is why I quickly discovered that I had to turn AI into a knowledgable, patient tutor rather than let it code for me. I am thoroughly at the helm for all decisions big and small - I don't let AI touch the code anymore.
And it is a lot cheaper in the end :)
Plus it’s so much cheaper… that has to matter.
Where I'd push back drawing too many conclusions from this study: arguably most successful AI usage is senior developers that know the programming environment they're working in. Know how far to trust the AI. And carefully review / understand outputs.
Nevertheless, the study's still interesting, and I wish they'd replicate with a much higher n per group. Junior developers (undergrads?) are a more abundant group and not particularly specialized yet. They've also spent years hand-coding at University, but probably could adapt to AI tooling pretty easily.
The study: https://www.anthropic.com/research/AI-assistance-coding-skil...
try:
something_that_should_not_fail_and_if_it_does_our_assumptions_are_all_wrong()
except:
fallback_that_will_not_result_in_correct_behavior_but_make_failure_hard_to_detect()
But, of course, it depends a lot on which models you're using and how you instruct them.I don't believe the "you are holding it wrong", "works on my machine", "works on this model" or "do this spec structure" type of arguments to compensate the fundamental issues. The tech simply does not do what is advertised and claimed as it is.
I’m serious. Treat it like any other tool. When it helps solves problems, use it. When it makes problems, don’t use it.
There are a lot of people and an enormous amount of money trying to make hands off agentic happen, but the happiest and most effective enthusiasts I know do not give up control: they go function by function and class by class, generating or writing as they see fit.
The goal is to make useful software. At least, I think that’s still true?
The goal, to corporations, has never been to "make useful software". The true goal is "make software that will bring us a revenue stream". If they think they can use AI to do that faster with less human payment, then they'll snap at it. AI doesn't ask for its rights (because it's not AGI, so it doesn't have actual rights). AI just tries to do what its told, and fucks up at doing so at a higher rate than the average human. But corps think it'll be cheaper so they swallow the tales told to them by AI executives who have a highly vested interest in making sure you use their AIaaS.
Hobby coders are coding for the fun of it, and aren't going to use AI to code. They might use AI to help them understand the subject matter better, but the code that hobby coders write is highly unlikely to be AI vibecoded. Evidence: I severely doubt that any demosceners will ever use AI to write the actual demo code.
Problem 1: "Obsessed with reinventing the wheel" " three duplicate functions":
Suggestion: plan then implement.
Tell LLM to scan your project and crete markdown file plan to solve the task first. DO NOT try to selve tasks in a single shot without planning. Review the plan file then, IN A NEW SESSION with clean context, tell LLM to read the implementation plan file and implement the plan according to the file.
---
Problem 2: "hyper-focuses on the current task and couldn't care less if its changes break other parts of the system"
Suggestion: add instructions to AGENTS.md file teaching LLMs how to run unit tests and other kinds of tests so it can make sure nothing broke. And also add to AGENTS.md that LLMs MUST run tests before marking the task done.
---
Problem 3: "you'll hit the 200k token limit in no time" "Long context = instant brain damage"
Suggestion: use 1 million context window LLMs. Also plan then implement will keep your context shorter.
If you can, use better LLM services which offer 1million context window. If you can't afford Anthropic or OpenAI, use DeepSeek V4 Flash or MiMo 2.5 for example. A $10/mo OpenCode Go subscription plan offers $60 in LLM credits which is A LOT for these cheap LLMs.
Also, planning phase is when the LLM has to scan the entire project to understand what needs changing. This is where the context bloat comes from. If you split tasks into planning + implementation, the scanning phase is condensed into a single markdown file which keeps context lean.
Bonus tip: Tell LLMs to use subagents when doing exploration.
---
Problem 4: The longer the context, the more incoherent its responses.
Suggestion: yeah, LLMs get dumber as their working memory fills up (just like me). If your session reaches 200k+ tokens, it's usually a sign you could have planned the feature better or split it up. It might be worth restarting with more clarification.
Yes, if the model someone is using only has 200k token limit, that would immediately suggest to me that it really isn't a sophisticated enough model.
Most of my coding sessions end up being about 350k tokens long when I finish, it wouldn't even fit in a 200k context. And that isn't counting the cache-reads by subagents, etc.
It's worth spending some time with the best Opus / GPT model, to at least get a sense of what the frontier is like.
Even DeepSeek v4 Flash has 1million context size.
But also on the list at 200k are "Free Models Router" and "Claude Haiku 4.5". I would not recommend making any judgment of AI based on free models. And coding with Haiku is a bad idea... I mean, that was my first code AI test too, but it's just not an accurate impression.
To be fair, Opus 4.1 & 4.5 are also listed as 200k. They did require context management for large & difficult tasks. But if you do have access to Opus, there's very little reason not to switch to 4.8 / Sonnet 1 Million now. I wouldn't recommend Sonnet, but I have used it to write a USB audio driver that got some hardware working on an obscure OS, so it can work.
I ran into many of the same issues, and they motivated me to experiment with a linter that flags duplication and architectural problems across a codebase. It’s still a work in progress though:
i use the skills /yaw-review excessively sometimes multiple times in a row on the same pr or session. followed by most often /yaw-address-all and then /yaw-coverage to add tests and /yaw-ship-ready to make production ready.
after a few rounds of these they are not needed every time on the same codebase.
if you are desperately wishing programming to go back to the before times it will never. or it will always be there but expect to be incredibly less productive than your peers.
For any issue, start a brand new context, point it to the spec, explain the issue and explain if it's a regression.
Also on it might seem like an obvious one, the more test coverage you have, the more your llm can tell if something has broken or if there's been a regression without needing to eat up context.
All of these things can help but there's no perfect solution.
My question tho is, how confident are we about an agentic future? I mean coding was the one thing agents "are best at". How would you run a complex system/organization on an agent where they will need to face with a massive (and growing) context through a limited context window?
1. Set the Agent off on some task
2. Go scroll social media
<15 minutes later>
Get back to whatever the agent was doing.
AI coding feels very anti-flow.
We have access to anthropic models, openai models and google models.
I run all my sessions on their best models with max thinking, because I don't care to optimise token usage at this stage. We are still learning every day about how to optimise our workflows, but I will say that I don't typically experience what you're describing.
I have very opinionated AGENTS.md files at the repo level, and at various other levels in the repo where more specialised rules are needed but I don't want those in my context unless that specific section of the codebase is going to be used or touched. I make a lot of use of skills. And my sessions are almost all "spec driven" in the sense that I type out an opinionated requirement to the LLM, tell it to challenge my thinking, to push back, to iterate on its own thinking, then to formulate a plan, then once done, go over it again to find any issues. I will then review the plan, or wing it, depending on the task. I then look at the overall code structure and design it has done. I have strong, opinionated coding rules in my AGENTS file. I have strong testing requirements (mostly end-to-end, not unit style).
I get really good results from this. But, I will say we're working in a highly opinionated codebase. We have the fundamentals in place already, where there are rules for how you do everything. The agent follows those rules pretty well. I'm not sure how well it would work on a codebase that is messy with a lot of conflicting design principles.
you can also use within claude code tui by running: typed cli off
i should probably change that to typed tui off/on. anyway.
Did you ask it to delete stuff or consolidate functionality? Did you ask it to reuse certain available implementations? Or do you use it as a black box, letting it do all the design and code, and not caring to steer it, except with some high level request ("build x")?
If it's the latter, if we treat a (non-actually-intelligent, generative) AI as a hands-off developer and remove ourselves from the loop, we get exactly what you mention in the rant.
But nothing forces us to use it like this (except the craze of "vibe coding"). Use it as a carefully monitored and steered coding assistant.
It's slower that way? That slowness is a requirement to ingest the expertise/knowledge/taste of a human developer in the mix. It's what avoids an avalance of slop to be commited and become part of the codebase unchecked.
That's why all the focus on maximazing speed, and removing the humans from the loop, are misguided.
When LLMs are ready to removed humans from the loop, we'll know: we'd be out of a job. As long as we have one, our role is to act as a quality bottlenect, not to open the floodgates.
this is where i start (c2l code i translate at https://c2l.puter.site):
define..sq.x[* x x
(define (sq x) (* x x))
ai step:
translate this code to python
I haven't written any code since November. People getting bad results don't know what they're doing.
A way to tell if you're doing it wrong is if you're writing lots of prompts. That's a huge smell.
Scope adjudication is extremely important in vibe coding, or agent can easily break your whole system with not applicable features.