Ran some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all.
Major analytical errors in their response to multiple of my technical questions.
Major analytical errors in their response to multiple of my technical questions.
Otherwise you're doomed to "sample size of one" level of relevance.