18 karma · joined March 15, 2025
Email: Max @ ^^^
> The key is to view the AI as a partner you can coach – progress over perfection on the first try
This is not how to use AI. You cannot scale the ladder of abstraction if you are babysitting a task at one rung.
If you feel that it’s not possible yet, that may be a sign that your test environment is immature. If it is possible to write acceptance tests for your project, then trying to manually coach the AI is just a cost optimization, you are simply reducing the tokens it takes the AI to get the answer. Whether that’s worth your time depends on the problem, but in general if you are manually coaching your AI you should stop and either:
1. Work on your pipeline for prompt generation. If you write down any relevant project context in a few docs, an AI will happily generate your prompts for you, including examples and nice formatting etc. Getting better at this will actually improve
2. Set up an end-to-end test command (unit/integration tests are fine too add later but less important than e2e)
These processes are how people use headless agents like CheepCode[0] to move faster. Generate prompts with AI and put them in a task management app like Linear, then CheepCode works on the ticket and makes a PR. No more watching a robot work, check the results at the end and only read the thoughts if you need to debug your prompt.
[0] the one I built - https://cheepcode.com
Now with headless agents (like CheepCode[0], the one I built) that connect directly to the same task management apps that we do as human programmers, you can get “good enough” PRs out of a single Linear ticket with no need to touch an IDE. For copy changes and other easy-to-verify tweaks this saves developers a lot of overhead checking out branches, making PRs, etc so they can stay focused on the more interesting/valuable work. At $1/task a “good enough” result is well worth it compared to the cost of human time.
In fact, I built an entirely headless coding agent for that reason: you put tasks in, you get PRs out, and you get journals of each run for debugging but it discourages micro-management so you stay in planning/documenting/architecting.
The result you described is coming soon. CheepCode[0] agents already produce working code in a satisfying percentage of cases, and I am at most 3 months away from it producing end-to-end apps and complex changes that are at least human-quality. It would take way less if I got funded to work on it full time.
Given that I'm this close as a solo founder with no employees, you can imagine what's cooking inside large companies.
[0] My product, cloud-based headless coding agents that connect directly to Linear, accept tickets, and submit GitHub PRs
It works by connecting directly to Linear and dispatching assigned tasks to agents that submit PRs in GitHub when finished. My agents work in a fully-integrated Linux development environment, including internet access. This means that they can browse the web, install dependencies, and creatively work around environment issues to make sure they run and test the code they ship.
It's really gratifying to see people asking all over the internet, "Where can I just create tickets and get pull requests?" because that's exactly the workflow I built CheepCode to support. As an engineer for almost 15 years, I knew what I personally wanted, and it really makes me happy to see that what I built will work for so many others too.
As a bootstrapped solo founder, it's challenging to juggle product/growth/development/strategy all at once, but also incredibly rewarding. I wouldn't necessarily say no to funding ;) but in the meantime, it's quite a thrill!
It often is, if you pick the right tasks (and more tasks fall into that bucket every few weeks).
You can get a simple but fully-working app out of a single prompt, though quality varies widely unless you’re very specific.
Once you have a codebase, agent output quality comes down to architecture and tests.
If you have a scalable architecture with well-separated concerns, a solid integration test harness with examples, and good documentation (features, stack, procedures, design constraints), then getting the exact change you want is a matter of how well you can articulate what you want.
One more asterisk, the development environment has to support the agent: like a human, agents work well with compiler feedback, and better with testing tools and documentation/internet access (yes my agents have these).
I use CheepCode to work on itself, but I am still building up the test library and preview environments to de-risk merging non-trivial PRs that I haven’t pulled down and run locally. I also use it to work on other apps that I'm building, and since those are far more self-contained / easier to test, I get much better results there.
If you want to put less effort into describing what you want, have a chat with an AI to generate tickets. Then paste those tickets into Linear and let CheepCode agents rip through them. I’ve got tooling in the works that will make that much easier, but I can only be in so many places at once as a bootstrapped founder :-)
[0] My headless coding agents product, similar to “assign to copilot” but works from your task board (Linear, Jira, etc) on multiple tasks in parallel. So far simple/routine features are already quite successful. In general the better the tests, the better the resulting code (and yes, it can and does write its own tests).
This paradigm feels like the obvious next step for agents. It more closely models human interaction (to the degree that this is desirable) and unlocks a lot of optimizations + powerful functionality.
It is going to be an exciting rest of the year!
*I have a few more safety/scalability changes to make but expecting public launch in a few weeks!