Raven brings Claude Code, Codex, and its own Research, Code, Design, and Oncall agents into a shared task graph. The idea is to let different agents handle the parts of a project they are suited to, with shared memory across subagents and context carried across sessions.
The RSI work extends to the harness itself: prompts, policies, strategy code, and playbooks. Raven's specialist harnesses and orchestration layer can be improved independently. Candidate changes are evaluated before adoption. This concerns Raven's own components; it doesn't rewrite Claude Code or Codex internals.
We've used Raven for long-running research and experimentation workflows and for building a Godot game. The repository includes examples and outputs, along with installation instructions. We're also exploring how to develop and refine specialist agents for particular domains.
Raven is pre-alpha and Apache-2.0 licensed. The self-improvement work is experimental; the Curator currently ships in the repository rather than the installed package.
Code and examples: https://github.com/EverMind-AI/Raven
Where do you find coordination between agents breaks down today? We'd also be interested in what evidence you'd want before trusting an agent-generated change to its own harness.