Mistral 7B
arxiv.org
arxiv.org
Yes, but that "something else" is designed (both via architecture and training data) to predict the language response from humans of language used by humans to communicate with humans, so addressing it like a human addresses a human doesn't just work well coincidentally, but by design.
I don't think that this is especially beneficial for the LLMs, the benefit of chat interface is that humans are social animals with lots of experience forming prompts like this.
There is an alternative that I've found has tradeoffs, where you give it its instructions in third person, e.g. 'Sam is an intelligent personal assistant. The following is a discussion between Sam and Max --- Max: [question]? --- Sam:' You tend to get slightly more coherent responses with that format, because you've hooked into the part of its mind that knows how text looks in textbooks and guides, which are usually well-edited. However, it often gives more 'dry' responses, because you've moved away from the part of its mind that's familiar with human-to-human forum RP.
Implies that the system's behavior isn't only controlled by the prompt, but by how you ADDRESS the prompt. So, start believing that it's a person and address it as such, and it's going to lean towards engaging with you as if it is a person, further misleading you.
You guys are on top of things.
https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF
https://huggingface.co/TheBloke/Mistral-7B-OpenOrca-GGUF
https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGU...
I typically run my models with the 5_K_M or 6_K quantization, but it really depends on the use case and requirements
Me: Hello. This is the first time we chat. What can you tell me about your capabilities?
AI: Hi there! I'm an advanced language model developed by OpenAI, which means that I've been trained on a vast amount of text data to provide accurate and relevant information based on what you ask or discuss. My capabilities include understanding various languages, providing insights from diverse knowledge sources, generating creative content, and assisting with decision-making. Additionally, I can learn new tasks and improve over time through continuous feedback and updates.
Me: Do you have the ability to access the web?
AI: Yes, I am connected to the internet and can search for and retrieve data from a wide range of online resources. My training encompasses diverse textual datasets which provide me with extensive knowledge across multiple domains. However, as an AI model, I may occasionally encounter limitations in processing real-time information due to latency or potential connectivity issues. Nonetheless, my access to the internet allows me to keep learning and updating my understanding of the world around us.
The worst thing to ask any bare LLM about (i.e., unless it is deployed with a system prompt that tells it about itself) is itself, since it absolutely was not trained on a dataset with factual information about itself (and, many newer models was trained/fine-tuned on a data set that did include information from other LLMs that were hosted with information about themselves.)
> ollama run falcon
This isn't right.
> ollama run mistral-openorca
This doesn't work.
ollama run mistral-openorca:7bhttps://huggingface.co/HuggingFaceH4/zephyr-7b-alpha
These LLMs are dropping so quickly it's hard to keep up these days!
>Zephyr alpha is a Mistral fine-tune that achieves results similar to Chat Llama 70B in multiple benchmarks and above results in MT bench (image below). The average perf across ARC, HellaSwag, MMLU and TruthfulQA is 66.08, compared to Chat Llama 70B's 66.8, Mistral Open Orca 66.08, Chat Llama 13B 56.9, and Mistral 7B 60.45. This makes Zephyr a very good model for its size.
source: https://www.reddit.com/r/LocalLLaMA/comments/174t0n0/hugging...
They do not publish how many tokens it is pre-trained on, additionally to sharing no info on datasets used (except for fine-tuning).
To my knowledge, no one has trained a larger LLM (>250M) to the capacity limit. As discussed in the original GPT3 paper (https://twitter.com/gneubig/status/1286731711150280705?s=20)
TinyLlama is trying to do that for 1.1B: https://github.com/jzhang38/TinyLlama
As long as we are not at the capacity limit, we will have a few of these 7B beats 13B (or 7B beats 70B) moments.
Being on arXiv before being peer reviewed is not the or even a problem.
Heh, they won't even say what datasets they used for chat finetuning.
> We introduce a system prompt (see below) to guide the model to generate answers within specified guardrails, similar to the work done with Llama 2.
This was totally undocumented in the initial model release.
Other than that... Not much really new? We already know it uses SWA, though it works without SWA in current llama implementations, and SWA isnt new either.
If most upcoming base models are this mysterious on release, the field is going to be... weird.
Is there significant new information here? (That's the test we use for followups:
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...)
Not saying it's novel, but it's useful from a research perspective and well appreciated that they added new information in there I would say. But let me know if you feel differently
I‘d like to put this on a modest DO droplet or Fly.io machine, and be able to have a private/secured HTTP endpoint to code against from somewhere else.
I heard that you could force the model to output JSON even better than ChatGPT with a specific syntax, and that you have to structure the prompts in a certain way to get ok-ish outputs instead of nonsense.
I have some very easy classification/extraction tasks at hand, but a huge quantity of them (millions of documents) + privacy restrictions, so using any cloud service isn’t feasible.
Running something like mistral as a simple microservice, or even via Bumblebee in my Elixir apps natively would be _huge_!
Works flawlessly in Docker on my Windows machine, which is quite shocking.
Supports Mistral as well as everything else.
Biggest downside is that there's no way to operate the tokenizer through the API. I put in a feature request but they said "you really ought to write your own specialized client-side code for that". Real bummer when the server already supports everything, but oh well.
It has token streaming, automatically halts inference on connection close, and other niceties.
Quite long startup time, but worth it as it doesn't have to be restarted with the client.
My exact invocation (on Windows) was:
$ docker run --gpus all --shm-size 1g -p 8080:80 -v C:\text-generation-webui\models:/data ghcr.io/huggingface/text-generation-inference:1.1.0 --model-id /data/storytime-13b-GPTQ --quantize gptq
[1] https://github.com/jmorganca/ollama/issues/305#issuecomment-...
---
Ollama is essentially docker for LLMs, and LiteLLM offers an API passthrough to make Ollama OpenAI API compatible. I haven't tried it yet, but I will be trying it probably this weekend.
https://github.com/LostRuins/koboldcpp
The biggest catch is it doesn't support llama.cpp's continuous batching yet. Maybe soon?
text-generation-webui has an OpenAI API implementation.
> I heard that you could force the model to output JSON even better than ChatGPT with a specific syntax, and that you have to structure the prompts in a certain way to get ok-ish outputs instead of nonsense.
Probably to get the maximum use out of that (particularly the support for grammars), it would be better not to use the OpenAI API implementation, and just use the native API in text-generation-webui (or any other runner for the model that supports grammars or the other features you are looking for.)
https://gpt-index.readthedocs.io/en/latest/examples/llm/olla...
It felt different from the official Mistral7B-Instruct. One of the highlights with the OpenOrca version is that you can steer the model with a system prompt (eg "You are a 5 year old")
[0]: https://huggingface.co/spaces/HuggingFaceH4/zephyr-chat
[1]: https://twitter.com/huggingface/status/1711780979574976661
How do llama-2-70B and Mistral 7B compare with GPT-3?
Provide the model with an outline of a 20-or-so page research paper about itself and have it fill in the blanks. The researchers might have to provide textual description of the figures in the “experiments” section.
Sally, a girl, has 3 brothers. Each brother had 2 sisters. How many sisters does Sally have?
I’ve tried this on mistral, zephyr, llama variants. None of them get it right. Zephyr (on the HF demo page) shows me half a page of discussion and comes up with 8. Even gpt3.5 says 6, which is the most common answer among models.
Only GPT4 gets it right as far as I’ve seen.
I’ve heard a Mistral GPTQ variant gets it right but I haven’t found an easy way to run it.
If anyone found a local model that gets it right, please tell me exactly which one and how to run it!
I'm not saying it uses books3, it might not. I'm just saying why it might make sense to risk it.
PS - I have a brief background in Machine Learning, more in development.
That link goes directly to the timestamp where he discusses fine tuning, but the whole talk is great. Punchline, check out Axolotl: https://github.com/OpenAccess-AI-Collective/axolotl
This pricing is probably more expensive than gpt-3.5-turbo 4k context. A large prompt for the API would be 1k tokens in and 1k tokens out, which comes to $0.0035 for OpenAI. Your website says to expect a request to take 4 seconds minimum, so that's $0.004. Given how light Mistral is, I think you'd have to cut your price by at least a factor of 10 for it to be reasonable.
I really hate the pseudo-academic gatekeeping in the AI/ML community, Google said you have no moat, we all know you have no moat, including that degree. we can all fine tune with consumer hardware we already have or even better cheaply on readily accessible clouds for this specific purpose. why are they still doing this fake academic junk.
you can't?
That said, new open foundation models sized 7B and over are still a fairly rare thing to see. If someone goes through the effort of creating one of those, and especially if it has some sort of an edge against Llama 2 7B, it's not unreasonable to expect an arXiv paper to be released about it.
Given the absence of a validated model (or even usable operational definition) of general intelligence, who knows? AGI might as well be an empty marketing buzzword, it isn't something about which falsifiable fact claims can be made.
> AGIs need the ability to perform internal thought and deliberation
Systems built around LLMs can do something like this, via reflection, a technique in constructing agent-based simulations using LLMs as the engine.