Until it doesn't...
Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter.
And I am feeling a lot better about it now that I've finally got it working.
Deepseek V4 Flash 0731 is surprisingly capable and cheap. [0]
Checkout pi coding agent. You can create as many different sub-agents as you wish, to specialize and understand and tackle or pass off any problem you like. It's refreshing, really. I feel like a coder in control again.
Huh, what's happened? I'm on the 20x plan and haven't noticed any fiasco, what went down exactly?
But since you don't know when resets are coming, it becomes kind of frustrating trying to game it out. One week they're resetting like crazy so you're trying to burn tokens as fast as possible. The next week they're not resetting at all so you have to adjust your workflow to be more conservative.
This whole thing sounds crazy to me, why are you so focused on making sure you hit 0% usage left when it's supposed to reset? Why can't you just use what you have and if it resets, it resets, and if it doesn't, it doesn't?
I've calculated I can spend about ~12% of usage every day, more or less, this is my "budget". If it resets, then the "12% per day" gets reset for that day, but that's it. Sounds crazy to me that I'd "invent work out of nowhere" just to spend more usage, why on earth would I adopt such a workflow? Sounds like you're burning tokens just to burn tokens???
Sure, in a video game where you have one attribute and you can minmax a strategy just focusing on that, but that's not how real-life works.
You can't just stack pending work on top of each other, expect yourself to be able to stay equally on top of everything and have the same results as if you didn't. With this comes the consideration about the tradeoffs of "produce mediocre but large body of works" vs "produce high quality but small body of works".
Who's to say what's more "rational" or not, it's not a straight-forward calculation which culminates in "Must consume all available usage to maximize resource usage" like some robot, as we are not.
It's quite literally inventing work out of nowhere as you wouldn't put the agent to do that and forcing yourself to be conscious about that work until it completes, unless you actually had the usage available. Asked another way, wouldn't your workflow clearly change if you had unlimited usage available? You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Not really. The AI will output the same sort of code at a certain skill level regardless of speed, it's not a human so the above is a false dichotomy. Also, it's sort of strange that you're saying it's inventing work out of nothing, you've never heard of a backlog? In companies that can grow very long and the more usage means the more that can be tackled.
> You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Yes? But that's probably after the backlog is complete unless they are truly low hanging or high priority fruit. So not sure how that reasons with your point.
It'll output the same code given the same prompts yes, but you don't just accept whatever it puts out, it requires iterations before it's actually ready to be committed as none of the agents write perfect code on their first try. So, it's not a "false dichotomy", I'm just looking at larger things than "LLM does inference"
I don't get the point of this. We all seem to agree that these companies have almost no moat, if one stops being a good deal, you can switch to another. That doesn't invalidate the existence of a deal that is currently good.
My point was that chasing deals like this is just kicking the can down the road. You're going to have to reckon with harsh price increases sooner or later.
So I have resolved to avoid that future-dreading and fixed it, basically.
Seems weird to me to not take advantage of the great deal the frontier labs are currently giving for subscription pricing when there's such an easy fallback in the worst case.
the rest of the comment is about moving from expensive switching costs of what harness you are using to having the difference be closer to changing some config in openrouter
Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
https://artificialanalysis.ai/models/grok-4-6#token-use
So depending on how you want to define "token efficiency", Grok is either tied with OpenAI, or in the lead.
[1] Though I grant that 4.6 appears to be wordier, on the order of Terra max.
I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
I haven't tried so it's pure speculation based on benchmarks, but I'd assume Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5.
For building a full stack custom CRM and media pipeline tool with video conversion, transcription, and indexing. Supabase, AWS, Meili, NextJS, GCS - lots of surfaces and planes.
4.8 basically couldn't do it, I abandoned the project as the fallback was, "current business processes".
With F5 it's been 4 weeks and almost ready for production release.
It’s not as quite as smart as opus 4.8 but it’s close and x4 the cheaper.
https://github.com/features/copilot/plans
https://github.blog/changelog/2026-08-06-kimi-k3-is-now-avai...
Have they converted entirely to transparent API rates + base allocation now? One of the reasons I left was that if I was going to be billed at API rates anyway, I'd just rather use the APIs. The value proposition still sucks for individuals now, when the other major providers are bundling at below-API rates.
On the OpenCode Go workspace there is a huge box:
Providers
Control which providers are used for routing.
Enable models hosted in China <toggle-button>
If you turn this off, DeepSeek Flash & Pro both stop working: Error: Provider request failed with HTTP 403: The latest version of this model is only available hosted in China and requires explicit opt in: https://opencode.ai/workspace/wrk_msh_is_a_liar_stop_bullshitting_please_be_real_example_url/go
Kimi K3, MiMo V2.5 Pro, MiniMax M3, Qwen3.8 Max all work (remarkably!) without the box ticked! That was a strong surprise for me.My request stands: I would like to have up-front information on what models only have Chinese providers. This information is not available except by buying a plan and experimenting, currently. And I had to try each one to find out, which I wish could be avoided. I want to end this thread where I started it (now that I have done the work to cut through the din and noise and misinformation), with my original request: please OpenCode Go make the information about which providers have Chinese-only/non-China hosting readily available on your website.
Apparently 5x usage when using Kimi Code too.
If you willing to share to no zdr, meta is waaaaaay cheaper vs Kimi.
With recent offerings from spacex and meta , I hardly imagine why would you pay money to any Chinese vendor it’s not as cheap and it’s not as intelligent neither .
Maybe deepseek is an exception , but it’s only good for narrow use cases that probably goes into modal.com and other gpu + fine tune me easy vendors , not vanilla dumb but cheap model .