HNHacker News
TopNewBestAskShowJobs

EugeneOZ

2,648 karma · joined March 21, 2012

https://jamm.dev

https://twitter.com/eugeniyoz

https://www.upwork.com/users/~01d95397aacaef6e88

http://careers.stackoverflow.com/oz

https://www.linkedin.com/in/newmanoz/

submissionscomments
EugeneOZ··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Can't trust a company which can halve a subscription any moment they want.
EugeneOZ··on GPT-6 Astra on OpenRouter
I love Astra pelicans! Awesome results!
EugeneOZ··on Muse Spark 1.3
Absolutely BRUTAL! :)

Thank you for doing this, I love your benchmark the most!

EugeneOZ··on Gemini 3.8 Flash and 3.8 Flash Cyber
Impressive pelicans!
EugeneOZ··on Claude Fable 5.1 and Claude Mythos 5.1
So the issue was in large context size of these old sessions.
EugeneOZ··on Claude Fable 5.1 and Claude Mythos 5.1
These pelicans are awful.
EugeneOZ··on Dolly Parton has died
Also not a fan of country/folk, but today I learned about her special albums trilogy: The Grass Is Blue, Little Sparrow, and Halos & Horns. Quite different music!

Some songs I added to my playlist :)

EugeneOZ··on Cursor removed cost information from the usage page and CSV export
> You can still see what you’re billed on the Spending page.

No, you can not: https://www.pasteboard.co/dNXUdT-h8Giy.png

If you want to say that "admin can" - it doesn't matter, I'm not going to ping admin every day to check how it goes. I'm not going to ask admin about every session to check how cost efficient a model was.

EugeneOZ··on Show HN: Physically accurate black hole you can put in your room
“Physically accurate” indeed raises the bar significantly.

You’ve already done great work here. That said, the feedback seems to come from someone who spent considerable time analyzing your work. Even if only a few of the suggestions are ultimately valuable, that’s still a meaningful contribution and worth considering.

EugeneOZ··on The new rules of context engineering for Claude 5 generation models
And now your CLAUDE.md is only good for Claude 5 models, not previous ones.
EugeneOZ··on The new rules of context engineering for Claude 5 generation models
> Then: Give Claude rules

> Now: Let Claude use judgement

No, it should follow my rules exactly. I don't care what code examples it was trained on - it will either write code the way I want, or I'll use another model.

EugeneOZ··on Kimi K3: Open Frontier Intelligence
Why we should waste 46 minutes instead of briefly looking at charts for 20 seconds? To pay their ads? No, thanks.
EugeneOZ··on Kimi K3: Open Frontier Intelligence
Any benchmark where Sol is better than Fable at coding is ridiculous.
EugeneOZ··on Why Vanilla JavaScript
The article is titled "Why Vanilla JS," but the takeaway seems to be "here are the custom abstractions I wrote."
EugeneOZ··on GPT-5.6
GPT 5.6 Sol is a token hog. After implementing the task, it started some "reviews" I didn't ask for - they consumed 19.5M and 11.9M tokens, while the task itself was below 5M tokens.
EugeneOZ··on 98% Isn't Much
If for 2% of users a webpage will not look as awesome as intended (it's not guaranteed that it will be broken), that's ok. It's not poisoning - it's a 98% chance of getting a top mark.
EugeneOZ··on John Carmack on Fabrice Bellard
Start fixing the unfixable and doing the undoable things ;)
EugeneOZ··on Nvidia RTX Spark
2 comments in total there
EugeneOZ··on Ferrari Luce – Designed with Jony Ive
Doesn't look like a sport car. From above it actually looks like a phone. The main thing is that the charging port isn’t on the bottom.
EugeneOZ··on Claude is not your architect. Stop letting it pretend
> Ask it if a microservices architecture makes sense for your three-person team and it’ll explain why microservices are an excellent choice

If you ask it to be fair and non-biased and provide pros and cons and give possible alternatives - it will. The catch - you might understand the explanation if you don't know the domain good enough.

Overall - a very, VERY good article, thank you!

EugeneOZ··on An update on recent Claude Code quality reports
If you think that you can just silently modify the model without any announcements and only react when it doesn't go through unnoticed, then be 100% sure that your clients will check every possible alternative and will leave you as soon as they find anything similar in quality (and no, not a degraded one).
EugeneOZ··on An update on recent Claude Code quality reports
> people didn't understand to use /effort to increase intelligence, and often stuck with the default -- we should have anticipated this

UI is UI. It is naive to expect that you build some UI but users will "just magically" find out that they should use it as a terminal in the first place.

EugeneOZ··on Saying goodbye to Agile
Absolutely awesome, thank you, Lewis!
EugeneOZ··on I still prefer MCP over skills
> Skills are great for pure knowledge and teaching an LLM how to use an existing tool. But for giving an LLM actual access to services, the Model Context Protocol (MCP) is the far superior

That's it. For some things you need MCP, for some things you need SKILLs - these things coexist.

EugeneOZ··on Goodbye to Sora
There are open-source alternatives:

https://mochi1ai.com/

https://wan.video/

and others. There are free to use tools also.

EugeneOZ··on Goodbye to Sora
This market will not be abandoned, and other tools already exist:

https://klingai.com/global/

https://aistudio.google.com/models/veo-3

https://runwayml.com

EugeneOZ··on Cursor Composer 2 is just Kimi K2.5 with RL
I don't know - it works okay (yet to be tested whether it is actually smarter than Opus 4.6), but it is not bad at all. So far, it works quite fine (I'm not testing the "fast" version).
EugeneOZ··on The changing goalposts of AGI and timelines
Not in my experience. Quoting my tweet:

Gave the same prompt to GPT 5.4 (high) and Opus 4.6 (high).

GPT 5.4 implemented the feature, refactored the code (was not asked to), removed comments that were not added in that session, made the code less readable, and introduced a bug. "Undo All".

Opus 4.6 correctly recognized that the feature is already implemented in the current code (yeah, lol) and proposed implementing tests and updating the docs.

Opus 4.6 is still the best coding agent.

So yeah, GPT 5.4 (high) didn't even check if the feature was already implemented.

Tried other tasks, tried "medium" reasoning - disappointment.

EugeneOZ··on The L in "LLM" Stands for Lying
I do, 100%, every line.
EugeneOZ··on Terence Tao, at 8 years old (1984) [pdf]
It depends on how much value their talents can bring to humankind, I guess.
Page 1 of 34Next →