I have been using the “free” Ling Flash model on openrouter for a bit over a month now. My work is an aside project at home building a Rust binding for an open source Zig code base library. The result is nothing impressive but also not a total failure: I have a feature parity binding to use in Rust vs. Python/java/typescript.
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.