Fyi, I haven't tested this yet.
Fyi, I haven't tested this yet.
When it comes to the high volume market of business automation, it seems that ultra-low cost, rather than expensive frontier intelligence, is exactly what you want, and low latency is also nice to have for customer-facing applications like customer service chatbots.
When I hear “architecture” I am thinking number of parameters and latency.
When I hear “accuracy” I think training recipe, data, and (later) number of parameters.
So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.
Hence my note about the risk of cost being the only moat.
But it is likely more than just a fine tune + novel training. At the very least the LM head is swapped out for a classifier one and then or also idk, bidirectional attention for the encoding pass I'm out of my depth at this point and will stop guessing. The training is probably where they have the biggest moat though, not that it's necessarily huge.
I have a project that fits jev as advertised almost comically well and I've been playing with it, and the various hacks and open versions. Jev doesn't necessarily perform better overall but it is quite different. It's sensitive to prompt phrasing in ways the others aren't, it's easy to generate questions where all the other models cluster in confidence but jev is an outlier. Not necessarily more correct, but it does feel like it's getting its answers in a different way.
I'm guessing just as much as anyone else but I've been spending a ton of time on this the last couple weeks, it landed right when I was most ready to dig into it.