2,635 karma · joined March 27, 2009
Agreed, so why frame it as avoiding unpleasant people? My experience with Anthropic's latest models is that no amount of instruction can overcome their latent verbosity (nor can AGENTS.md overcome OpenAI models' reticence). I tried A/B testing several models and Opus was the clear winner, so I'm not likely to retire it just because it generates unpleasant prose. I do curate model-specific notes to mitigate some of the worst offenses.
I realized that all clues could be represented by Boolean formulas, allowing a solver to check whether my deductions and hypothetical scenarios were consistent without revealing the solution.
Of course the same solver can solve for individual variables, but that takes too much of the fun out of the game, so I left that feature out.
One could argue the opposite conclusion, that technology helps break monopolies, but either view depends on reductionist historical readings. The truth is somewhere in between.
> Bash(DATABASE_URL=$(grep -E '^DATABASE_URL=' .env 2>/dev/null | head -1) echo "ok")
even though I have in CLAUDE.md:
> For database queries, use tidewave first.
I then prompted:
> use tidewave as per CLAUDE.md. also diagnose why you failed to heed that
> ● Diagnosis first: I defaulted to shell habits (env grep → psql) instead of pausing to recall the CLAUDE.md rule that tidewave is the first-line DB tool. The trigger was "look at this record" — I should have read that as "run a SQL query" and reached for tidewave immediately.
If Opus 4.7 doesn't follow simple CLAUDE.md instructions, I'm not sure what benefits other markdown files could bring. I don't trust Opus's own explanation, but it could point to the fact that the system prompt for bash is much longer than CLAUDE.md with tidewave.
While LLM judging could be helpful, I think the tool-call assertions (https://github.com/darkrishabh/agent-skills-eval#what-you-ge...) may be the most useful thing in agent-skills-eval given that it's the only objective measure of compliance.
1. https://github.com/rails/rails/blob/fa8f0812160665bff083a089...
I do like the basic concept and directory structure, but those are easy enough to adopt without all the cruft.
1. https://github.com/Fission-AI/OpenSpec/blob/main/src/core/te...