The presenter, somewhat overeagerly, "Why not ask our new AI?" and went on to type: "Are you an independent model or do you use OpenAI?"
To chat bot answered in flourish language that sure it was using ChatGPT as a backend. Which it was not and which was kind of the whole point of the presentation.
so never demo LLMs. got it.
https://www.tomshardware.com/news/google-gemini-ai-video-sta...
The most common reaction I get is "wow, I didn't expect that to work so well with so little effort". For most tasks, a fine-tuned Mistral 7B will consistently outperform GPT-3.5 at a fraction of the cost, and for some use cases will even match or outperform GPT-4 (particularly for narrower tasks like classification, information extraction, summarization -- but a lot of folks have that kind of task). Some aggregate stats are in our blog: https://openpipe.ai/blog/mistral-7b-fine-tune-optimized
I've been looking forward to someone providing a detailed guide on how to "fine tune it with your custom data" for ages!
Like with "prompt engineering", a lot of people are just hiding how much of the heavy lifting is from base models and a fluke of the merge. The past few "secret" set leaks were low/no delta diffs to common releases.
I said it a year ago, but if we want to wowed, make this a job for MLIS holders and references librarians. Without thorough, thoughtful curation, these things are just toys in the wrong hands.
same here, it doesn't adhere to explicit instructions, maybe one or two simple instructions are ok but not more complex ones
This also helps distributes traffic as a side effect.
I guess the problem is how the conversation would flow. If the user changes topics from say art to quantum physics then asks a question about quantum physics and art then I'm not sure what the algorithm should do.
I'm not sure it's "distributing" traffic so much as amplifying it.
Load is divided across 2 models. Load balancing is a feature for free and division is across subjects. Of course this is assuming each model owns it's own set of gpus.
What you're suggesting is just simply intent classification and using a specific model per intent. That's what everyone did _before_ LLMs.
Experience has been great, INT8 tradeoffs are acceptable until hardware FP8(FP4 anyone?) becomes more widely & cheaply available. On-prem costs have been absorbed already for few boxes of A100s & legacy V100s running millions of such interactions.
Nvidia 40xx Tensor Cores support FP8 (but not the fancy async stuff on H100s). For some reason it's not used for models (AFAIK).
Some people report throttling issues in some cases, though.
Part of its training data was code.
Mistral 7B fine-tunes (OpenChat is my favorite) just chug through the data and get the job done.
Details: using vLLM to run the models. Using ChatGPT-4 to condense information for complex prompts (that the local models will execute).
I think, the situation will just keep on getting better with each month.
Imagine using a GPT-2 type model when everyone else is using GPT-4. Until the dust settles there's no point in investing in alt models imo, unless you're leading the research.
Hopefully os models can catch-up to gpt4 in the next six months when we fixed all the low hanging fruit outside of the model itself
Of course all of them feel like a Frankenstein compared to actual ChatGPT. They feel similar and work just as well until, sometimes, they put out complete and utter garbage or artifacts and you wonder if they skimped on fine-tuning.
Currently i use lmstudio on my m2 with 96gb ram. But i‘m looking into switchin to ollama or another oss solution.
here's a dead simple way : (1) download LM Studio, install it[0] (2) download a model from within the client when prompted (3) have a ball.
the program is fairly intuitive, it takes care of finding the relevant files, and it can even accept addendum prompts and various ways to flavor or specialize answers.
Learn the basics there, take what you learn to a more 'industrial' playground later on.
[0]: https://lmstudio.ai/
As simple as
pip install llm
# add the local plugin
llm install llm-gpt4all
# Download and run a prompt against the Orca Mini 7B model
llm -m orca-mini-3b-gguf2-q4_0 'What is the capital of France?'
Alternatively, you could use the llamafile[1] which is a tiny binary runner which gets packaged ontop of the multigigabyte models. Download the llamafile and you can launch it through your terminal or a web browser.From the llamafile page, after you download the file, you can just launch it as
./mistral-7b-instruct-v0.2.Q5_K_M.llamafile -ngl 9999 --temp 0.7 -p '[INST]Write a story about llamas[/INST]'
[0] https://llm.datasette.io/en/stable/index.html[1] https://github.com/Mozilla-Ocho/llamafile
Edit: added llm quickstart from the intro page
However, if you want to get into the weeds of how this actually works, I recommend you look up model quantization and some libraries like ggml[1] that actually do that for you.
If you want a truly general purpose front-end for LLMs, the only good solution right now is oobabooga: https://github.com/oobabooga/text-generation-webui
All other alternatives have only small fractions of the features that oobabooga supports. All other alternatives only support a fraction of the LLM backends that oobabooga supports, etc.
Follow the guide all the way until you get to "Loading our model in Oobabooga". Then ignore the rest. You can do inference in Ooba under the Notebook tab.
(You can also ignore the "enabling HTTP API" parts, but it's quite handy, it's an OpenAI-compatible API which means you can use any OpenAI-compatible web UI)
Mostly I think I need to use LLMs more effectively
I no longer use llms.