At 300 tokens per call, you'd get:
One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.
Would probably make sense to self-host Clef, if you have the capability/resources. If not...
At 300 tokens per call, you'd get:
One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.
Would probably make sense to self-host Clef, if you have the capability/resources. If not...
For privacy perhaps but on pricing you're unlikely to come out ahead versus datacentres with scale and industrial power pricing
Since it benches better than Jev, Jev is probably smaller and easier to host. They could also be losing a lot of money.
It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input