I've been looking forward to someone providing a detailed guide on how to "fine tune it with your custom data" for ages!
Like with "prompt engineering", a lot of people are just hiding how much of the heavy lifting is from base models and a fluke of the merge. The past few "secret" set leaks were low/no delta diffs to common releases.
I said it a year ago, but if we want to wowed, make this a job for MLIS holders and references librarians. Without thorough, thoughtful curation, these things are just toys in the wrong hands.
This also helps distributes traffic as a side effect.
I guess the problem is how the conversation would flow. If the user changes topics from say art to quantum physics then asks a question about quantum physics and art then I'm not sure what the algorithm should do.
I'm not sure it's "distributing" traffic so much as amplifying it.
Load is divided across 2 models. Load balancing is a feature for free and division is across subjects. Of course this is assuming each model owns it's own set of gpus.
What you're suggesting is just simply intent classification and using a specific model per intent. That's what everyone did _before_ LLMs.
same here, it doesn't adhere to explicit instructions, maybe one or two simple instructions are ok but not more complex ones