How do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness?
It's all vibes, and the numbers contradict the vibes. There's 30 years of literature trying to explain the "productivity paradox" where we can't see any excess productivity driven by computer technology. Lots of FOMO, no hard data. For an entire generation. And people come here every day and say stuff like you just said and they really seem to think that "this time is different".