I would test this, might be cheaper per task even costing more per token, probably faster too
But yes, I intend to support several models, to handle the overload situation (automatic failover). And switch to a cheap paid plan, though it seems like that'll mostly improve rate limits, which barely matters for my usage.
Faster is always good, though. I do care about latency.