3,397 karma · joined February 20, 2007
RedwoodJS & RedwoodSDK - https://rwsdk.com Agent CI: https://agent-ci.dev Machinen: https://machinen.dev Kindling: Coming soon.
Previously built Snaplet: https://github.com/supabase-community/seed
That's false. The LLM will only answer competently if it was trained on that data; and if it has enough data to make the correct connections between your question and the "correct" answer.
In the case of this article they're specifically saying the LLM has limited training.
It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.
There are a lot of interesting ideas in it, mostly none that I have applied to my own workflows. Curious if anyone has used it?
I have tried to use act many times, and many times I've failed.
P.S. pause on failure is also helpful for humans, but I'm trying to be realistic about where the future of programming is going...
The jobs runs via containers.
No, it's not like "act," because it uses the standard Github runner, the difference is that the control plane is an emulation of api.github.com, because of this we can do all kinds of nice things:
Caching in ~0 ms. Pause on failure, so you can let your AI agent fix it and retry without pushing.
I built this because I treat CI as the last line of defense. Agents also need validation. They should use CI, and they shouldn't bother you unless everything is green!
GH Actions is usually in the top-5 expenses for dev-teams. Add agents to that mix? It'll easily double. It's the wrong tool for the right job: Slow boot, slow cache, retrieving logs is token expensive for agents, the list goes on...
So I built a tool with one amazing feature: live-reload for failures. Agent-CI is a local CI runner.
I tweaked the control pane and mounts to provide 0ms caching, insanely fast boots. When a step fails it pauses, provides the agent with the failure, and waits for the agent to fix and retry just that step.
It uses the standard GH Actions image (via Docker), but emulates the control pane via a local HTTP server. You don't have to change any of your existing GH workflows.