GPT5, Sonnet 4, and Gemini Pro 2.5 are all 1x. Opus is 10x, for comparison.
https://docs.github.com/en/copilot/reference/ai-models/suppo...
Also worth keeping in mind that Copilot has reduced context windows even for the premium models, which has a very real impact on agentic performance.
I use Copilot because work is paying for it and it can be made usable, but requires being really deliberate about managing context to keep things on the rails. It's nice that it gives you access to a pretty decent selection of models, though.
At home, I'm mostly using the $100 Claude plan. It's definitely not cheap, but I've found it has a pretty decent balance for my casual experiments with agentic coding.
Another option to seriously consider is setting up an account with OpenRouter and just tossing some cash into your bucket on occasion. OpenRouter lets you arbitrarily make API requests to pretty much any model you want. I've been occasionally tossing $10 or so into mine and I'll use it when I've hit my usage limits with Claude or if I want to see how another model will attack a particular task.
FWIW, I use Roo code for all of this, so it's pretty easy for me to switch between models/providers as I need to.
Unlike some other workarounds this is a fully supported workflow and does not break Copilot terms of service with reasonable personal usage. (As far as I understand at least. Copilot has full visibility into which tools are using it to make chat requests so it isn't disguising or impersonating Copilot itself. When first setting it up there's a native VS Code approval prompt to allow tool access to Copilot and the LM API is publicly documented).
But anything unlimited in the LLM space feels like it's on borrowed time, especially with 3rd party tool support, so I wouldn't be surprised if they impose stricter quotas for the LM API in the future or remove the unlimited limit entirely).
For any problem with a lot of reasoning or problem solving I use "architect" mode first with Gemini 2.5 Pro or Claude Sonnet 3.7/4 to break it into discrete subtasks that 4.1 can follow pretty successfully. This approach is very cost effective as Gemini can do a lot of high level planning quickly and cheaply.
I'm sure a lot of the experience depends on how 4.1 is being used, I've fine tuned my custom Roo code configuration to work around its limits without a lot of sacrifices, I'm sure using it out of the box with Copilot is asking a lot more from a weaker model on its own.
The latter is a commonly recommended strategy in general for any large task even with more powerful models to keep context manageable and allow recovering easily if on step 9/10 the LLM loses it and starts mangling all the previous work it did. That way you don't have to start all over from the last good checkpoint or commit.