And Phi-3 is something else, even from relatively limited time playing with it, so that’s useful signal for anyone who hadn’t looked at it yet. Wildly cool stuff.
It seems weird to not mention Anthropic or Mistral or FAIR pretty much at all: they’re all pretty clearly on more modern architectures at least as concerns capability per weight and Instruct-style stuff. I’m part of a now nontrivial group who regards Opus as basically shattering GPT-4-{0125, 1106}-Preview (which is basically the same as 4o for pure language modalities) on basically everything I care about, and LLaMA3 is just about there as well, maybe not quite Opus, comparable if you ignore trivially gamed metrics like MMLU.
And I have no idea why we’re talking about GPT-5 when there’s little if any verifiable evidence it even exists as a training run tracking to completion. Maybe it is, maybe not, but let’s get a look at it rather than just assume that it’s going to lap the labs that are currently pushing the pace now?