It's really tiring to have to tweak everything with each model release and then watch those changes mess up cheaper models in the process.
I am on board with not putting stuff like "write clean code" into an agent file, or using plugins for tools that are now built into the harness. I don't see enough evidence to support models being significantly better at figuring out intent, or getting the assumption correct. I've always gotten better results (as ever) with very constrained instructions, vs "fix the install".
I’m sure this just intended to steer you to vendor lock-in. Remember, these are the same people who are so insecure/petty about their product that they won’t make it recognize the .agents/AGENTS.md standard.
And a lot of people latched on to it as a form of self-soothing.
"I may not write much code anymore, but I can still be an expert prompt engineer!"
I wrote this last year, it's still true:
--- start quote ---
https://dmitriid.com/prompting-llms-is-not-engineering
In reality these are just shamanic rituals with outcomes based on faith, fear, or excitement. Engineering it is not.
--- end quote ---
There are some prompts useful for the user like brainstorming [1] but on the whole it's nothing but lucky charms
[1] brainstorming from superpowers: https://github.com/obra/superpowers
You can see this visually in older image generation models. Stable Diffusion 1.5 would produce wildly different images based on slight variations in prompt and seed, but the latest image gen models are nearly seed indifferent and can tolerate a decent amount of prompt tweaking while staying "consistent"
Yes, they keep surprising me that they still keep doing all of this: https://news.ycombinator.com/item?id=48962703 with no improvement despite all the marketing assurances that "hey you don't need to read code anymore"
I have some wrappers but the only one I really use is a variant of that pinned to haiku for quick questions.
Anthropic had something about API-only in the description a couple months back, but they've been waking all those back for months since the fable access rollback fallout
Finally he also "recommends" that you do not look at the code, or even understand it.
His "recommendations" are designed to get you to spend even more tokens and get you hooked on the Opus / Fable slot machine in order to extract as much money as possible from your wallets.
News at 10.
Now, the codebase is managed by agents and, while the company has the source, it's as if they bought it from a vendor and pay the vendor for changes. The vendor of their own system is the AI company selling them tokens.
He can be informative to listen to, as long as you keep that in mind.
If you were, please share the knowledge with us. Where can I learn more?
Tried out Opus 5 and it's been a super annoying experience out of the box.