just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so
just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so
They're almost tied for felonies.
In some sense though, sure, skill issue explains the gap vs. Anthropic’s much less severe alignment issues.
That’s not the most parsimonious explanation even if the assumption it rests on (anthropic ahead of OpenAI) is true, which we don’t have proof of.
other than anecdote, do you have a comparison table or something that I can refer to to see this clearly?
While I notice ad hoc announcements from these companies, I don't have an overall pulse and tally that gives me an objective perspective.
We don't know what internal models look like, and any guesses about it are just speculation.
astra is a good workhorse, but its much less generally intelligent
Hint. You lose using either.
> I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?
> When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it
> That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!
> A model like that should never have gotten out of QA, let alone been released.
I've seen the same pattern regardless of open v closed, don't have the same family that wrote the code also review the code
diversity has this way of making things better across everything humans do
I only use open weight models now and I don't really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.
When you say ... it's hard to take you seriously
> dude were in the singularity, this opinion was cute 18 months ago
you are definitely displaying strong bias that Anthropic is way ahead of everyone throughout your posts under this story
as such, I give your opinions zero weight, they don't align with the majority of accountings or my own experiences
here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models
https://github.com/verdverm/pge-jax#note-from-author
are open weights lagging, yes, are they way behind, no
if open weights were so inferior, they would not be >50% of all token processing
why are open weight models seeing such rapid rise in usage?
there has been a step function change this summer, like the end of last year for closed models
---
do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?
(for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it's a question from curiosity about how others are engaging with the technology)
I don't think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn't imply that they're competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.
That said, at the moment I'm finding that not much can compete with GPT-6 Luna on cost / performance (not using for coding, but for AI pipelines in my product).
this is different and nuanced from the "not even close" or "they are trash" that the other person in this thread has opined, note how they also claim Claude is way ahead of OAI as well