Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.
I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.
Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.