Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these days.
Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these days.
Claude 3.7 Sonnet was consistently on top of OpenRouter in actual usage despite not gaming benchmarks.
I personally couldn't care less about them, especially when we've seen many times that the public's perception is absolutely not tied to the benchmarks (Llama 4, the recent OpenAI model that flopped, etc.).
Not all benchmarks are well-designed.
so effectively you can only guarantee a single use stays private
> By default, we will not use your inputs or outputs from our commercial products to train our models.
> If you explicitly report feedback or bugs to us (for example via our feedback mechanisms as noted below), or otherwise explicitly opt in to our model training, then we may use the materials provided to train our models.
https://privacy.anthropic.com/en/articles/7996868-is-my-data...
Don't forget the previous scandals with Amazon and Apple both having to pay millions in settlements for eavesdropping with their assistants in the past.
Privacy with a system that phones an external server should not be expected, regardless of whatever public policy they proclaim.
Hence why GP said:
> so effectively you can only guarantee a single use stays private
AI companies are grasping at straws by selling us minor improvements to stale technology so they can pump up whatever valuation they have left.
What we've seen from Veo 3 is impressive, and the technology is indisputably advancing. But at the same time we're flooded with inflated announcements from companies that create their own benchmarks or optimize their models specifically to look good on benchmarks. Yet when faced with real world tasks the same models still produce garbage, they need continuous hand-holding to be useful, and they often simply waste my time. At least, this has been my experience with Sonnet 3.5, 3.7, Gemini, o1, o3, and all of the SOTA models I've tried so far. So there's this dissonance between marketing and reality that's making it really difficult to trust what any of these companies say anymore.
Meanwhile, little thought is put into the harmful effects of these tools, and any alleged focus on "safety" is as fake as the hallucinations that plague them.
So, yes, I'm jaded by the state of the tech industry and where it's taking us, and I wish this bubble would burst already.