It’s crazy that people feel confident making judgments like these when the model’s been out for only a few hours.
It's a pretty easy spot if you're already using an older model daily.
Its even crazier that people are sitting here trying to calculate intelligence per dollar from metrics. At least first impressions have more basis in real performance.
It overthinks quite a bit above medium effort, try using that.