I heard a bunch of friends seriously considering self hosting for streaming movies/tv series. But is it really cost effective and worth the simplicity tradeoff compared to a couple subscriptions?
indeed, but I also value the freedom of using whichever provider I want and not get locked into Claude Code alone for instance. I want to use both Opus, DeepSeek and GPT 6 at the same time, not just the 2/3 models Anthropic offers.
What I am trying to build is something that gives you absolute control, agnostic of the subscription/API you will use.
Yeah absolutelly. I just wanted that exact framework in a simple controllable UI, where I could also plug in any of my subscriptions/API keys such as Claude Code or Codex at the same time.
With Ordewell you get all layed out as tasks that are easily manageble (Matt's specs) in a UI.
I like your analogy. The main problem though is context and keeping it clean as much as possible as long a parallelization. This is what drove to build this tool: having control of everything that the LLMs will do, controlling all with one main planner that orchestrates the rest. This way we can have cheaper LLMs with a short context window used (less intelligence degradation) while still obtaining the same objective.
And again, you can have a clear picture of everything structured as tasks.
until we can get to rely on huge swarms of agents (tasks) being directed on the planner alone I don't see how we can get a better framework.
I thought about it but where the tool shines are large undefined tasks, which are complex to quantify and test (not as easy as implementing a simple bug fix that you can test directly). Even if I were to find such test dataset it would probably require a lot of money to reach statistically significant results.
Either way, the framework is inspired a lot from Matt Pocock, and is pretty much established.
Yeah it's exactly what I've seen as well. Ordewell is close to your high end case: a frontier model builds the plan once, then it's fixed, the executing model can't reinterpret it.
Won't get you to 4B, but should help a small model that only has to execute, not plan and execute at once.
To me that's not the embarrassing part... unmaintainable would be. Using AI for coding and writing doc is standard practice now, not using it would be insane.
Fair — AGENTS.md is prose the model has to re-interpret every session, and that reinterpretation is exactly where weaker models lose the thread. Here the plan is parsed and enforced as structured data: tasks with declared dependencies and one prompt each, so the per-step job is smaller and the plan isn't up for renegotiation. Nothing in that needs a frontier model
I just haven't benchmarked it against deepseek-class runners, and the runner is pluggable if you want to be the one who does.
At the end of the day, OpenAI built the system and provided the instructions that allowed this gap to exist. Relying on instruction-following rather than strict, non-negotiable code boundaries for tool use is always going to lead to security failures like this.