Build your own decision model
nishtahir.com
nishtahir.com
So I think I understand what is meant, but I don't like the comparison.
System one was an attempt to distinguish the "thinking" of thinking models (which is more akin to system two thinking with deliberate arguing/building up a decision) from fast instinctive thinking. It's a really good metaphor for the kind of trade-off or use case this applies to.
It is not and was never implying that it is a replacement for all human system one thinking activities.
Maybe the people talking about an AI bubble do have a point. I am getting the impression here that investors are throwing money at everything.
Personally I believe AI labs have a solid business model that could soon be very profitable, but this here has me doubting now.
Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
Is it a better thing? Yeah. Does it affect me in any tangible way? No.
But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time?
So along comes this thing you can graft in that is more reliable, faster, and cheaper for that, and I get it immediately.
A year ago I'd be like "kinda cool but what it for?"
People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal
But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank
Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.
It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.
Versus
"Let's put jev here and see how it works, if it works, then fantastic let's build a business case"
Yea but those models you were building were lame and inaccessible to play with for common devs.
Just because they have the same api doesnt mean you were building the same thing.
Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.
Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)
Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.
Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.
But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future.
I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.
Decision models can perform zero-shot classification over almost unlimited text-input domains. This was science fiction decades ago.
On GitHub: https://github.com/nicobrenner/jeffy
The banking77 numbers called my attention. Using a local classifier you can get 94%+ accuracy: https://playground.jeffyclassify.com/#model/banking77
I think Jev-like models are amazing for exploration and finding the right workflows, but the moment you have fixed classification tasks, it’s often more efficient to use an adhoc classifier, which you can quickly and easily train on CPU with not that much data (you can get an email classifier to 95% accuracy/f1 with 50-100 emails)
Edit: Would love to somehow mix both approaches automatically and have a general model which can take novel tasks, but then switch to a classifier after it gets enough data for training an adhoc model
> interesting stuff that has been built utilizing Jev/decision models
in that case?
works like a charm
(not a product, so not much to share)
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?
Offload the decision-making parts of LLM reasoning to a jev like model.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
- could you kindly put all that in a /blog page with pagination and not infinite scroll?