2,464 karma · joined July 15, 2014
The feature set is pretty simple:
- Agents that can write their own tools.
- Agents that can write their own skills.
- Agents that can chat via standard chat apps.
- Agents that can install and use cli software.
- Agents that can have a bit of state on disk.
Perhaps more advanced llms + specifications + better tests.
If you can make good tests the AI shouldn't be able to cheat them. It will either churn forever or pass them.
https://github.com/magnet-linux/magnet-linux
Not really ready for prime time, but I think I have some interesting ideas there at least.
My point is about making it so that you have to actively risk money to push the truth needle in the wrong direction.
I feel like this is the sort of thing a prediction market might be able sort out.
I kind of consider them the same thing. Openpilot can drive really well on highways for hours on end when nothing interesting is happening. Claude code can do straight forward refactors, write boilerplate, do scaffolding, do automated git bisects with no input from me.
Neither one is a substitute for the 'driver'. Claude code is like the level 2 self driving of programming.
This has always been good practice anyway.
If it really can't be done then you aren't really shipping a library that is meant to be depended on.