There used to be this thesis in software of [Cathedral vs Bazaar](https://en.wikipedia.org/wiki/The_Cathedral_and_the_Bazaar), the modern version of it is you either 1) build your own cathedral, and you bring the user to your house. It is a more controlled environment, deployment is easier, but also the upside is more limited and also shows the model can't perform out-of-distribution. OpenAI has taken this approach for all of its agentic offering, whether ChatGPT Agent or Codex.
2) the alternative is Bazaar, where you bring the agent to the user, and let it interact with 1000 different apps/things/variables in their environment. It is 100x more difficult to pull this off, and you need better model that are more adaptable. But payoff is higher. The issues that you raised (env setup/config/etc) are temporary and fixable.
cathedral = sandbox env in the provider's cloud, so [codex](https://chatgpt.com/codex) uses this model. Their codex-cli product is the Bazaar model, where you run in your computer, in your own environment.
Claude Code, on the other hand, doesn't have the cloud-based sandboxing product, you have to run in on your computer, so the bazaar model. You can also run in in a way that anthropic never envisioned (e.g. give it control to your house). Curser also follows the same model, albeit they have been trying to get into the cathedral model by using the background agent (as someone also pointed out below). Presumably not to lose the market share to codex/jules/etc.
Can deploy as a github action right now.
Tag it in any new issue, pr, etc.
Future history will highlight Claude Code as the first true form agent. These other analogies are not intuitive enough for the evolution of an os-native agent into eventual ai robotics.
-----
> The software essay contrasts two different free software development models:
> The cathedral model, in which source code is available with each software release, but code developed between releases is restricted to an exclusive group of software developers. GNU Emacs and GCC were presented as examples.
> The bazaar model, in which the code is developed over the Internet in view of the public. Raymond credits Linus Torvalds, leader of the Linux kernel project, as the inventor of this process. Raymond also provides anecdotal accounts of his own implementation of this model for the Fetchmail project
-----
Source: Wikipedia
If you're a software developer and especially if you're doing open source, CATB is still worth a read today. It's free on the author's website: http://www.catb.org/~esr/writings/cathedral-bazaar/cathedral...
From the introduction:
>No quiet, reverent cathedral-building here—rather, the Linux community seemed to resemble a great babbling bazaar of differing agendas and approaches (aptly symbolized by the Linux archive sites, who'd take submissions from anyone) out of which a coherent and stable system could seemingly emerge only by a succession of miracles.
> The fact that this bazaar style seemed to work, and work well, came as a distinct shock. As I learned my way around, I worked hard not just at individual projects, but also at trying to understand why the Linux world not only didn't fly apart in confusion but seemed to go from strength to strength at a speed barely imaginable to cathedral-builders.
It then goes on to analyze why this worked at all, and if the successful bazaar-style model can be replicated (it can).
Both situations you've described are Cathedrals in the CATB sense: all dev costs are centralized and communities are impoverished by repeating the same dev work over and over and over and over.
Sometimes I just realize that CC going nuts and stop it before it goes too far (and consume too much). With this async setup, you may come after a couple of hours and see utter madness(and millions of tokens burned).
A tight feedback loop is best for me. The opposite of these async models. At least for now.
But pushing this existing process - which was designed for limited participation of scarce people - onto a use-case of managing a potentially huge reservoir of agent suggestions is going to get brittle quickly. Basically more suggestions require a more streamlined and scriptable review workflow.
Which is why I think working in the command line with your agents - similar to Claude and Aider - is going to be where human maintainers can most leverage the deep scalability of async and parallel agents.
> is way better than having to set up git worktrees or any other type of sandbox yourself
I've built up a helper library that does this for you for either aider or claude here: https://github.com/sutt/agro. And for FOSS purposes, I want to prevent MS, OpenAI, etc from controlling the means of production for software where you need to use their infra for sandboxing your dev environment.
And I've been writing about how to use CLI tricks to review the outputs on some case studies as well: https://github.com/sutt/agro/blob/master/docs/case-studies/i...
It's interesting that most people seem to prefer local code, I love that it allows me to code from my mobile phone while on the road.
Tapping the prompts in is the easy part, but async model is different to work with, I feel more like a manager, not a co-developer.
No special environment instructions required.
One nice perk on the ChatGPT Team and Enterprise plans is that Codex environments can be shared, so my work setting this up saved my coworkers a bunch of time. I pretty much just showed how it worked to my buddy and he got going instantly
Ideally, a combination of both I feel like would be a productive setup. I prefer the UI of Codex where I can hand-off boring stuff while I work on other things, because the machines they run Codex on is just too damn slow, compiling Rust takes forever and it needs to continuously refetch/recompile dependencies instead of leveraging caching, compared to my local machine.
If I could have a UI + tools + state locally while the LLM inference is the only remote point, the whole workflow would end up so much faster.