Building Personal Software with Claude
blog.nelhage.com
blog.nelhage.com
It also hits its 'maximum limit allowed at this time' quite often due to it's inability to understand its own environment and just modify an existing software artifact rather than regenerating the entire thing. I know it can do it, it knows it can do it because it will tell you why it failed but it just wont do it for some reason. Of course this comes after you manage to convince it to quit removing sections of code so its context window is already polluted from past failures.
I spent a solid three days trying to get it to successfully do a somewhat simple thing (something that would've taken me an afternoon if I dedicated myself to the task) just to see if it could manage to do it and it failed horribly, multiple times, even after it was able to say why it was failing, no matter how many tries I gave it. And it's not like it is some impossible task either, it was able to successfully (might I even say impressively) accomplish the exact same task with no problems at a smaller scale.
It's also pretty horrible at debugging, even if you tell it exactly the problem, but that's just me testing the limits of what it can do.
So, like TFA states, it is pretty amazing at smaller tasks but utterly and completely fails at anything larger than a couple hundred lines of code.
Seems more like throwing good money after bad as they appear to make it so you have to get an API key to do anything useful and then there's the possibility of spending significantly more than the $20 month 'pro' plan. Not something I'm really all that interested in considering this has nothing more than entertainment value and, from a video on the youtubes that broke the costs down, they have the highest per token rate.
I have enough yaks to shave to justify $20/month but anything over that...
A lot of the problems you describe are due to the need to provide adequate context, project details, and rules in your prompt.
Claude will generate files it doesn’t think already exist, and in a SaaS project, you have many files.
Claude will struggle with syntax and proper usage of library/module/package versions beyond its training data. You have to provide your project with knowledge to work around this.
Lastly you will hit usage limits while working on projects because it’s a fixed cost offering. You can track and generate a project status to pass on to the next agent. This works “ok” and when I hit my sonnet limit I use haiku to bug and type fixes.
Bottom line, an out of the box chatbot is a great playground to flush out your techniques, but most software projects have complexities which must be managed in a separate system designed to manage all the details and break down the project into hundreds or thousands of individual tasks.
I want this magic wand too, but you have to build that yourself (or buy it when it becomes available). It’s been a fascinating learning process.
The task was/is to take a grammar for APL from some long forgotten paper and turn it into a lemon parser. Easy, peasy, well within its wheelhouse and it had spectacular initial results with the help of DeepSeek-R1 analyzing its work.
"Oh, good job, robot," me types, "let's work on a lexer. Hmm... you seem to have clipped out some important rules at some point, we need to add those back." Then, boom, Claude is completely worthless.
I want Claude to succeed. It was doing so well then it hit a self-reinforcing wall of failure that it just can't get over even though it can analyze its behavior and say exactly why it keeps failing.
I mean, exactly zero people think the world needs an APL interpreter written by the robots but the point of the project is to see how far they can get without having a human write a single line of code. I know they have limitations and have no problem helping them work around them.
But, alas, this project is shelved until the next big hype cycle.
The thing is it starts to fail when you are just prompting it without any example code.
And to fully prompt it to be exactly what you wanted to write is basically writing the code yourself.
So for novel code I have had far less success.
But any kind of translation or matching an API shape, it works pretty much easily.
Also Deepseek R1 can do this for much much cheaper and at virtually the same level of quality.
At home I use it with open router and ollama/vllm/llama.cpp server