Qwen 3.8 and Claude Opus 5 show why raw benchmark scores don't predict the billventurebeat.com·6 pts·ashurandi·1