But the community not.
Anthropic's community, I assume, is much bigger. How hard it is for them to offer something close enough for their users?
Not gonna lie, that’s exactly the potential scenario I am personally excited for. Not due to any particular love for Anthropic, but because I expect this type of a tight competition to be very good for trying a lot of fresh new things and the subsequent discovery process of new ideas and what works.
Stories like this reinforce my bias
I think Anthropic's docs are better. Best to keep sampling from the buffet than to pick a main course yet, imo.
There's also a ton of real experiences being conveyed on social that never make it to docs. I've gotten as much value and insights from those as any documentation site.
Early adopters are some of the least sticky users. As soon as something new arrives with claims of better features, better security, or better architecture then the next new thing will become the popular topic.
For your sake, I’m not saying they’re wrong. I’m just pointing out something I’ve noticed.
My personal approach is to start from minimal, push that as far as I can, keep the human in the loop so I know how good it actually is, then find ways that have a better overall ROI. I'm performing similar loops to clawd/ralph manually, with my attention paid.
1. Stable models
2. Stable pre- and post- context management.
As long as they keep mothballing old models and their interderminant-indeterminancy changes, whatever you try to build on them today will be rugpulled tomorrow.
This is all before even enshittification can happen.
The practical workaround most teams land on is treating the model as a swappable component behind a thick abstraction layer. Pin to a specific model version, run evals on every new release, and only upgrade when your test suite passes. But that's expensive engineering overhead that shouldn't be necessary.
What's missing is something like semantic versioning for model behavior. If a provider could guarantee "this model will produce outputs within X similarity threshold of the previous version for your use case," you could actually build with confidence. Instead we get "we improved the model" and your carefully tuned prompts break in ways you discover from user complaints three days later.