9 karma · joined September 5, 2022
Tried the usual approaches - longer context windows, better prompts, chat history - but none of it worked reliably. The fundamental issue is that LLM agents are stateless by nature.
So I built a state machine that persists everything to DynamoDB: - Which development phase we're in (requirements → frontend → backend → etc) - Granular todos within each phase - What's been completed vs what's pending - Sandbox state (E2B sandboxes can die/restart) - S3 code sync status
Now when something goes wrong (and it always does), the agent just resumes from the exact todo it was working on. No context loss, no duplicate work.
This post shares honest lessons on agent reliability, context engineering, and avoiding “buzzword” distraction. Posting for feedback and real debate—are we missing something by not building Layer 2, or is this the right call?
I started this project out of frustration. When I tried to clone other projects using Claude Code and customize them a bit—simple Next.js, ECS, CDK, and Express server setups—it took several hours just to get everything working. I realized that while vibe coding is great, it's still time-consuming to build a production-ready, functioning product.
Building justcopy.ai - lets you clone, customize and ship any website. Built 7 AI agents to handle the dev workflow automatically.
Kicked them off to test something. Went to grab coffee.
Came back to a $100 spike on my OpenRouter bill. First thought: "holy shit we have users!"
We did not have users.
Added logging. The agent was still running. Making calls. Spending money. Just... going. Completely autonomous in the worst possible way. Final damage: $200.
The fix was embarrassingly simple: - Check for interrupts before every API call - Add hard budget limits per session - Set timeouts on literally everything - Log everything so you're not flying blind
Basically: autonomous ≠ unsupervised. These things will happily burn your money until you tell them to stop.
Has this happened to anyone else? What safety mechanisms are you using?