We've stumbled into general differentiable models..
We've stumbled into general differentiable models..
After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.
To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.
The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.
Probably because doing it wrong (using an llm in place of a classifier) is more profitable? (For the people selling inference.)
There have been tons of applications for this. People were using earlier LLMs like BERT for classifiers long before LLMs became viable chatbots.
People didn’t know they wanted classifiers until OpenAI gave them a taste.
I think the hype with Jev is just that, while structured generation is great, LLM judges tend to kind of suck for precise classification. And the more powerful the base model, the more accurate they can get, but they get increasingly expensive/impossible to finetune. It was specifically the latency/price point Jev offered vs. the general accuracy it claimed that generated all the excitement. Plus the promise of cheap calibration (tuning).
"Jev exists because LLMs exist" is kind of a truism, as Jev apparently is literally a Transformer model.
In early 2016, Clarifai's core product was basically just an image classifer exposed via an API. And at that point, they'd raised $40 million--the same amount as TypeSafe.ai/Jev--and they were experiencing viral growth among developers + signing contracts with a bunch of flashy logos. The demand was so high that Amazon launched Rekognition and Google launched their similar APIs to compete.
The AI hype cycle and the number of people thinking about using AI certainly puts more wind at Jev's back, but even in an alternative universe where we don't have contemporary LLMs, Jev's core claims would be remarkable and there would be a big market for it.
LLMs still are better than Jev at the task, just across the board slower.
Anyone who had a reason to try this already tried it (ads/recommendations) - back in 2023/2024 during the first fine tuning wave and it was accurately determined that it was not worth the effort, the results were more bogus than just using CoT, so frankly parallelism meant nothing if bogus * parallel = bogus.
So thrown into the dumpster and nobody really cared to revisit because it was already tried.
Pretty much sometime between then and now it somehow became the state where the tradeoff makes sense now.
Maybe it's just a case of the idea just now crossing the threshold into working just good enough to be worth it
I have 80,000 voice recordings to classify. Very few are in English, and the classes are not in English. Many of the classes are project names or other proper nouns. How could a system not trained on data specific to the problem possibly be expected to work?
and you might find the scores it gives back about its confidence are useful for escalating to a more expensive classifier.
but also, there's no requirement to use it as a zero-shot classifier, you can provide it as many examples as you like, and prioritise giving it examples it had previously gotten wrong. and with input caching it might be economical.
I'd be surprised if they or others don't start offering a fine tuning API for models like this, like openai does for some models (or used to, I haven't checked in a long time).
As are, honestly, most projects I work on.
There are already Jev-shaped open weight local models coming out that can be fine-tuned on specific data. We'll see which of them the community settles on.
There's probably other areas where people are using LLMs where a more tailored ML solution might work better.
I adhere to the idea that this is software's "Tower of Babel" moment where everyone just fundamentally ships things in completely diverging architectures, because creating a ground up architecture is no longer something that needs to be avoided for an economically viable business mode that in the past two decades would have otherwise incentivized people into industry standards. In a world where "taste" is the focus, single ingredients in the recipe aren't enough.
This particular "innovative" concept already has a rich, open research background. What Jev appears to have done is scale that up a bit and isolate good training data, which results in a great product but not really something impossible to imitate. The only major difference currently is that the open source decision models need to be fine-tuned as they're not trained off of the entire internet yet.
My thoughts too. It sounds like a minor feature being framed as a whole new business.
Then again, Dropbox and Docker are too.