Text classifiers are an underrated application of LLMs
blog.kasperjunge.com
blog.kasperjunge.com
I cannot recommend guidance enough. You can use shockingly small Llama models for some tasks with guidance while only actually generating a handful of tokens.
You should highly consider some form of guidance/logit bias for classification especially if you have a known set of classes. This will ensure you get it in the format that you want, with the correct classes that you want.
Keep in mind LLMs perform much better with COT. So you make it explain what the text/image is, then explain the possible classifications, then list its final decision. Again guidance can ensure it follows the correct format to do this.
LLM's still massively benefit from finetuning, especially if you want too classify it in a particular format. Notebook tags vs SFW/NFSW vs important subjects, etc. Existing alignment can sometimes mess with some of these classifications too which finetuning helps smooth out.
lmql might be a decent alternative.
Any form of logit bias should work though.
Of course, GPT-4 is insanely expensive to use at scale, and still isn't a perfect classifier. So the next step is to take the outputs you get from GPT-4 and use them to fine-tune a smaller model that's really fast and good at your specific problem. In my experience, even without using any human annotations or online learning, a model fine-tuned just on GPT-4 outputs can actually outperform GPT-4 as a classifier! This seems really counterintuitive at first, but my guess is what's happening is that the training process is a kind of regularization, so the weird mistakes GPT-4 occasionally makes are overwhelmed in the training data by all the times when GPT-4 gets it right.
As a disclaimer, we're building open source tooling to ease the transition from prompt to cheaper fine-tuned model at my company OpenPipe.
Between open source modeling tools being incredible, transfer learning allowing dirt cheap fine-tuning and now mega-models being able to instantly give you a "mostly right" data set, the cost of creating ML features has dropped to almost nothing.
Products that took quarters/years and required big budgets for labeling, ML specialist, GPUs etc just a few years ago can now be done in an hour or so for free (if you are scrappy). I imagine this is going to lead to a ton of great ML features that weren't worth funding in the past but are very valuable in aggregate. Similar to the mid-2000s when the cost/ease of web development came down enough that there was a lot more experimentation and fun to be had.
Being able to teach an AI assistant to look for specific (but not too specific) things with just a prompt would be incredibly helpful.
I have written some stories myself, with the help of GPT, i will try to parse my stories with your method. It is very interesting.
As a side note, GPT is definitely not a toy. I use it for coding, it is great! I use it to write command line apps, which do some simple data manipulation, some more complex than others, but in the order of hundreds of lines of code. They work flawlessly, without me writing even a single line of code.
One interesting use case we see is a SaaS company using the our REST API to access a Particle with custom instructions just for integration with other systems. They will provide a CSV row and the GPT-4 model will classify and map the columns into their key columns. In effect, they are able to integrate with almost any system in their vertical with an out-of-the-box integration. Albeit, it is more expensive, but it is great for the initial trial phase and the costs can be passed to the customer. https://www.particlesy.com
[0] https://matt-rickard.com/categorization-and-classification-w...
For some evidence, say a sentence, then say a sentence with the words scrambled. The performance nose dives, relative to audio levels.
I used to love those posts.
They achieve this via a fairly complex set of regular expressions, which must represent a significant time investment to research and maintain.
The same effect can be achieved with a properly prompted GPT-4 completion request, which took about five minutes to write.
You can prompt gpt4 and get something that looks plausible for a few test cases with very little effort, but can you get any guarantees that it will behave reasonably for most inputs? And if you can, will those guarantees last as the model is updated underneath you?
How do you update your prompt to take the new data point into account? Or do you just add it as an example inside the prompt and let it grow?