Edit: Better at chain of thought, long running agentic tasks, following rigid directions.
Edit: Better at chain of thought, long running agentic tasks, following rigid directions.
Figures that any article written on LLM limits is immediately out of date. I'll write an update piece to summarize new findings.
It's very hard to evaluate whether a model is better than another, especially doing it in a scientifically sound way is time consuming and hard.
This is why I find these types of comments like "model X is so much better than model Y" to be about as useful as "chocolate ice cream is so much better than vanilla"
This is a great way to describe what I've been feeling / experiencing as well.
* Connection between points
* Flows better
* Eyes don't start-stop as much
Different readable than the more flowing, conjunct readable than yours (which is the more typical use of it I concede)