Claude 4 Opus and Sonnet seem much better for me. The models needed alignment and feedback but worked fairly well. I know Copilot uses Claude but for whatever reason I don't get nearly the same quality as using Claude Code.
Claude is expensive, $10 to implement a feature, $2 to add some unit tests to my small personal project. I imagine large apps or apps without clear division of modules/code will burn through tokens.
It definitely works as an accelerator but I don't think it's going to replace humans yet, and I think that's still a very strong position for AI to be in.