i am looking forward to when we can put this behind us, it is still a major issue
535 karma · joined January 14, 2015
i am looking forward to when we can put this behind us, it is still a major issue
https://github.com/waynenilsen/crumbler
This uses recursive task decomposition but is single thread by design. Honestly fast enough for me and makes it easier to reason about
such as agent native shopping platforms whereby a human will bring you something from walmart or what not could spring up and disrupt your instacart of the world
this of course is just one simple example, when it works better for the clawdbot or whatever comes next what are the users going to choose they'll say 'get me some apples from walmart using instacartforbots' because they know the agent success rate will be higher
still not great in my opinion sadly i think this will likely get worse and perhaps never get better
commit messages as marketing materials and 'kpi's is an anti-pattern i would have preferred to avoid
forms, PRG, semantic HTML and no js needed
The amount of manual QA I am currently subjected to is simultaneously infuriating and hilarious. The foundation models are up to the task but we need new abstractions and layers to correctly fix it. This will all go the way of the dodo in 12 months but it'll be useful in the meantime.
agent-browser helped a lot over playwright but doesn't completely close the gap.
https://david.rothlis.net/declarative-schema-migration-for-s...
hopefully someone is independently archiving all documents
my understanding is that some are being removed
$ phone call bill
ok call_id=3f2a
$ phone status 3f2a
dialing
$ phone status 3f2a
answered
bill: hello
$ phone say 3f2a "hey, quick question"
ok
Indeed I believe this is probably the next thing to solve but even here I don't think it is out of reach. What we aught to be able to do is disconnect and make asynchronous the goals of the project with where we are. This, in normal software building, is encapsulated by the roadmap. I am building roadmapping prompts now and broadening the scope of the software development lifecycle even further to the encapsulate the roadmap as well which was previously out of scope for the experiment I am running now.
The prompts I am using now give the agent autonomy over 'make the next prd that makes sense' however I think it is a straightforward extension to add 'in the context of the @roadmap/ ' or similar with probably decent results.
Have you tried something similar?
Even without a roadmap the agent continues to do useful work over 24 hours in. You can see the commits and PRDs they really are quite sensible and I pulled and tested and everything really is working quite well. Frankly, I am shocked it is working at all. I have had to step in once or twice you definitely need to keep an eye on the logs every once in a while. Getting the loop booted up in a reliable way was the hardest part to be honest and even that was not terribly difficult.
does that make sense?
the ticket burndown is a very nice feature because whenever you want to add a ticket it'll just pick it up and do its best
I am still hoping that openSCAD or something similar can grab hold of the community. openSCAD needs some kind of npm as well as imports for mcmaster-carr etc but I think it could work.
> Windsurf reports Opus 4.1 delivers a one standard deviation improvement over Opus 4 on their junior developer benchmark, showing roughly the same performance leap as the jump from Sonnet 3.7 to Sonnet 4.