Opus 4.7 is horrible at writing
Similar experiences?
Similar experiences?
Work experience and actual outliers are better signals. Anyone can read a book and get a 100.
It goes to show that there's a very large and vocal user base using it for writing, and yet it's not part of the benchmark for Anthropic.
Anyway, try Sonnet 4.5 while it's still available?
This is something it spit out just now (trimmed a 9 line comment though):
let keepSize = 0;
let overBudget = false;
await this.items.orderBy('[priority+dateUpdated+size]')
.reverse()
.eachPrimaryKey((primaryKey, cursor) => {
if (overBudget) {
evictKeys.push(primaryKey as string);
return;
}
const key = cursor.key as [number, number, number];
const itemSize = key[2];
const contribution = itemSize > 0 ? itemSize : 0;
if (keepSize + contribution > maxSize) {
overBudget = true;
evictKeys.push(primaryKey as string);
return;
}
keepSize += contribution;
});
Come on now... what? For a start that entire thing with its boolean flag, two branches, and two early returns could be replaced with: let totalSize = 0;
await this.items.orderBy('[priority+dateUpdated+size]')
.reverse()
.eachPrimaryKey((primaryKey, cursor) => {
const key = cursor.key as [number, number, number];
const itemSize = key[2];
const contribution = itemSize > 0 ? itemSize : 0;
totalSize += contribution;
if (totalSize > maxSize) {
evictKeys.push(primaryKey as string);
}
});
I'm back to 4.6 for now. Seems to require a lot less manual cleanup.I guess they broke continuity with a 0.1 in model version change in some ways.
- economics: i'd wager a bet that Opus 4.7 is just distilled Mythos Preview - performance: surgery like this would explain the spiky performance and weird issues
just spitballing tho
It is not only the model that affects the end results. Good technical specification, architecture documents, rules, lessons learned, release notes, proper and descriptive prompting are also important.
Regardless of which one. They're too verbose. They repeat information. They lack cohesion. Overly agreeable. The flaws are part of the tool.
Meaning: You managed your ways around the system prompt and usage intention - Congrats! Now it doesn't work any more - Bummer!
Have you tried opus 4.7 in comparison to 4.6 with a general purpose / writing system prompt in the app? Thats where this would make more sense.
i asked it to look at writing research (especially from NN/g) and come up with some alternatives for a heading that is roughly supposed to convey:
"this app lets you create custom shortcuts for all mac apps, even sophisticated ones, mouse wheel ones, ..."
it came up with the following headlines:
1. ONE APP, MANY MAC APPS
2. ONE INSTALL, MANY APPS ONE APP.
3. ONE PACK PER MAC APP.
4. ALL YOUR APP SHORTCUTS IN ONE PLACE
5. [app name] IS ONE APP. PACKS COVER THE REST.
whereas GPT-5.4 came up with 1. The Shortcuts Your Apps Are Missing
2. What Your Apps Still Don't Let You
3. Do The Parts You Still Do by Hand
4. When Built-In Shortcuts Run Out
5. Where Your Apps Stop Short
now both of these aren't amazing, but please tell me how in the world "ONE APP MANY MAC APPS" makes sense as a headline for fucking anything lolthat's not something even GPT-3.5 would come up with.
"one install, many apps" ........huh???