There are quantizations already. I would suggest running it on KoboldCPP, or any of the UIs recommended on this page: https://huggingface.co/TheBloke/MistralLite-7B-GGUF
Personally, I run the elx2 quantization with exllamav2 because its faster and more coherent in the same amount of vram. You can do that with Aphrodite or Text-gen-ui: https://huggingface.co/models?search=MistralLite%20exl2
<|prompter|>{prompt}</s><|assistant|>
Why is that </s> in there?And how would I prompt it with a previous history of the conversation?
<|prompter|>hello</s><|assistant|>
And got back this (truncated): hey there!</s>
introducing the first of its kind, a revolutionary AI...
Why did that </s> come back as part of the response?<|prompter|> is the human message (including the instructon/contexts from my tests) and <|assistant|> is the bot's response.
There are llm wrappers that will save chat histories and reformat them appropriately for different LLMs. For instance, see TUI and its instruction template presets: https://github.com/oobabooga/text-generation-webui/tree/main...
<|prompter|>... prompt goes here</s>
<|assistant|> ... assistant response will go here, maybe followed by </s>
Still don't understand why I got a rogue </s> in the reply when I ran a prompt through it (but it continued after that tag): https://news.ycombinator.com/item?id=38101766And is my guess about </s> at the end of the assistant thing correct? If so then a conversation might look like this:
<|prompter|>hi</s>
<|assistant|>Hi to you too. What do you want to talk about?</s>
<|prompter|>Tell me facts about snakes</s>
<|assistant|>
But the documentation on these models is so thin! Do I need that </s> at the end of the assistant things or not?Welcome to LLM research land, lol. Mostly there is no documentation, no reproducability, just random hit-and-run releases with some claimed metrics and if you're lucky a model card with the trained prompt.
> Do I need that </s> at the end of the assistant things or not?
...I dunno. I didn't add it in my prompt template, and you can see that the model itself it outputting that token in the responses. Maybe ask the trainers on the HF page.
But TheBloke doesn't (yet) do exl2 quantization. That one is almost better done yourself, as the bpw target is arbitrary, and requires profiling text that should probably fit your intended use (though many just use wikitext as a good default).