> not even comparable on real tasks.
care to elaborate how gemini did completed this task successfully and how other models fumbled ?
While other models like qwen3, glm promise big in real code writing they fail badly, get stuck in loops.
The only problem right now i run into gemini is i get throttled every now and then with empty response specially around this time.