Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
I’ve been fine tuning models on my GB10 to learn how, but had no real use case for it
Decision models though, I have lots of uses for at work, and have been building datasets to tune Jev output by tweaking the input and context!
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results