You guys are on top of things.
https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF
https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GGUF
https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGU...
I typically run my models with the 5_K_M or 6_K quantization, but it really depends on the use case and requirements
Me: Hello. This is the first time we chat. What can you tell me about your capabilities?
AI: Hi there! I'm an advanced language model developed by OpenAI, which means that I've been trained on a vast amount of text data to provide accurate and relevant information based on what you ask or discuss. My capabilities include understanding various languages, providing insights from diverse knowledge sources, generating creative content, and assisting with decision-making. Additionally, I can learn new tasks and improve over time through continuous feedback and updates.
Me: Do you have the ability to access the web?
AI: Yes, I am connected to the internet and can search for and retrieve data from a wide range of online resources. My training encompasses diverse textual datasets which provide me with extensive knowledge across multiple domains. However, as an AI model, I may occasionally encounter limitations in processing real-time information due to latency or potential connectivity issues. Nonetheless, my access to the internet allows me to keep learning and updating my understanding of the world around us.
The worst thing to ask any bare LLM about (i.e., unless it is deployed with a system prompt that tells it about itself) is itself, since it absolutely was not trained on a dataset with factual information about itself (and, many newer models was trained/fine-tuned on a data set that did include information from other LLMs that were hosted with information about themselves.)
> ollama run falcon
This isn't right.
> ollama run mistral-openorca
This doesn't work.
ollama run mistral-openorca:7bhttps://huggingface.co/HuggingFaceH4/zephyr-7b-alpha
These LLMs are dropping so quickly it's hard to keep up these days!
>Zephyr alpha is a Mistral fine-tune that achieves results similar to Chat Llama 70B in multiple benchmarks and above results in MT bench (image below). The average perf across ARC, HellaSwag, MMLU and TruthfulQA is 66.08, compared to Chat Llama 70B's 66.8, Mistral Open Orca 66.08, Chat Llama 13B 56.9, and Mistral 7B 60.45. This makes Zephyr a very good model for its size.
source: https://www.reddit.com/r/LocalLLaMA/comments/174t0n0/hugging...
Yes, but that "something else" is designed (both via architecture and training data) to predict the language response from humans of language used by humans to communicate with humans, so addressing it like a human addresses a human doesn't just work well coincidentally, but by design.
I don't think that this is especially beneficial for the LLMs, the benefit of chat interface is that humans are social animals with lots of experience forming prompts like this.
There is an alternative that I've found has tradeoffs, where you give it its instructions in third person, e.g. 'Sam is an intelligent personal assistant. The following is a discussion between Sam and Max --- Max: [question]? --- Sam:' You tend to get slightly more coherent responses with that format, because you've hooked into the part of its mind that knows how text looks in textbooks and guides, which are usually well-edited. However, it often gives more 'dry' responses, because you've moved away from the part of its mind that's familiar with human-to-human forum RP.
Implies that the system's behavior isn't only controlled by the prompt, but by how you ADDRESS the prompt. So, start believing that it's a person and address it as such, and it's going to lean towards engaging with you as if it is a person, further misleading you.