The claims of certain models outperforming GPT-3.5-Turbo and approaching GPT-4 fail to hold up to their benchmark results in real-world scenarios, potentially due to data contamination in assessments, based on my testing.
As noted in the linked survey paper, some models may outperform 3.5-Turbo in specific, narrow areas, depending on the model. Yet, we still lack a general model that definitively exceeds 3.5-Turbo in all respects.
I'm concerned that while we're still striving to reach 3.5-Turbo's performance level, OpenAI may unveil a new next-generation model, further widening the performance gap! Back in the summer, I had higher hopes that we would have surpassed the 3.5 threshold by now.
The performance gap has been surprisingly large. It is especially noticeable in areas requiring consistent structured output or tool use from the LLM. This is where open models particularly falter.