The rate limits I've encountered with free api keys has been way lower than the limits advertised.
I agree. I found it unusable for anything but casual usage due to the rate limiting. I wonder if I am just missing something?
I think it's the small TPM limits. I'll be way under the 10-30 requests per minute while using Cline, but it appears that the input tokens count towards the rate limit so I'll find myself limited to one message a minute if I let the conversation go on for too long, ironically due to Gemini's long context window. AFAIK Cline doesn't currently offer an option to limit the context explosion to lower than model capacity.
I'm pretty sure that's a google maps' level of free where once in control they will massively bill it
There is no reason to expect the other entrants in the market to drop out and give them monopoly power. The paid tier is also among the cheapest. People say it’s because they built their own their inference hardware and are genuinely able to serve it cheaper.