2,648 karma · joined March 21, 2012
https://twitter.com/eugeniyoz
https://www.upwork.com/users/~01d95397aacaef6e88
http://careers.stackoverflow.com/oz
https://www.linkedin.com/in/newmanoz/
Thank you for doing this, I love your benchmark the most!
Some songs I added to my playlist :)
No, you can not: https://www.pasteboard.co/dNXUdT-h8Giy.png
If you want to say that "admin can" - it doesn't matter, I'm not going to ping admin every day to check how it goes. I'm not going to ask admin about every session to check how cost efficient a model was.
You’ve already done great work here. That said, the feedback seems to come from someone who spent considerable time analyzing your work. Even if only a few of the suggestions are ultimately valuable, that’s still a meaningful contribution and worth considering.
> Now: Let Claude use judgement
No, it should follow my rules exactly. I don't care what code examples it was trained on - it will either write code the way I want, or I'll use another model.
If you ask it to be fair and non-biased and provide pros and cons and give possible alternatives - it will. The catch - you might understand the explanation if you don't know the domain good enough.
Overall - a very, VERY good article, thank you!
UI is UI. It is naive to expect that you build some UI but users will "just magically" find out that they should use it as a terminal in the first place.
That's it. For some things you need MCP, for some things you need SKILLs - these things coexist.
and others. There are free to use tools also.
Gave the same prompt to GPT 5.4 (high) and Opus 4.6 (high).
GPT 5.4 implemented the feature, refactored the code (was not asked to), removed comments that were not added in that session, made the code less readable, and introduced a bug. "Undo All".
Opus 4.6 correctly recognized that the feature is already implemented in the current code (yeah, lol) and proposed implementing tests and updating the docs.
Opus 4.6 is still the best coding agent.
So yeah, GPT 5.4 (high) didn't even check if the feature was already implemented.
Tried other tasks, tried "medium" reasoning - disappointment.