MistralLite by Amazon Web Services
huggingface.co
huggingface.co
- Rope theta of 100,000, likely from the Llama 2 Long paper which found that a large theta helped regulate attention between distant tokens[0]
- A 16k (effective 32k) context window, improving upon Mistrals 4k (effective 8k) context window
In The Llama 2 Long paper, they saw improvement in short context benchmarks as a result of long context fine tuning. I can't find any of the expected MMLU / HellaSwag / etc benchmarks yet. Benchmarks haven't been submitted to MTEB yet.
Some user anecdotally seem to be having trouble with generating quality responses [2][3][4]. I can't find any examples of users getting good results from the model outside of using exact examples from the documentation.
[0] https://arxiv.org/pdf/2309.16039.pdf
[2] https://old.reddit.com/r/LocalLLaMA/comments/17jd00g/mistral...
[3] https://old.reddit.com/r/LocalLLaMA/comments/17kzlbl/anyone_...
[4] https://old.reddit.com/r/LocalLLaMA/comments/17b0n8t/llama_2...
It still behaves more like an good unaligned LlamaV2 13B finetune.
And it loves to summarize and retrieve facts. In fact, when I give it a huge context and instruct it to storytell, deduct or whatever, sometimes it will summarize the context instead... But the summary is excellent.
I think L70b is the mark to beat, I’m hopeful mistral will take the crown with a smaller model.
But some of the 3rd party 13B finetunes/merges are also much better that Meta's chat model.
Maybe they think, why train the foundational model from scratch when you can get some company to rent your sever to do it, and then license it to you for free anyways?
To me, it feels good enough for retrieval/summarization tasks Amazon could sell, especially with grammar based sampling.
Stable Diffusion made a fuss a year ago. Everyone is hoarding A100s and H100s like they are the last GPUs ever made, and we have... Maybe 5 good downloadable base model LLMs, and ~4 good diffusion models? And only one from a really large company.
shrug
They announced BedRock and made it GA directly after a gated preview of just a couple of months.
Titan is their trained model.
November is their annual gala conference time. Expect a lot of GenAI announcements resounding through late November.
There are quantizations already. I would suggest running it on KoboldCPP, or any of the UIs recommended on this page: https://huggingface.co/TheBloke/MistralLite-7B-GGUF
Personally, I run the elx2 quantization with exllamav2 because its faster and more coherent in the same amount of vram. You can do that with Aphrodite or Text-gen-ui: https://huggingface.co/models?search=MistralLite%20exl2
But TheBloke doesn't (yet) do exl2 quantization. That one is almost better done yourself, as the bpw target is arbitrary, and requires profiling text that should probably fit your intended use (though many just use wikitext as a good default).
<|prompter|>{prompt}</s><|assistant|>
Why is that </s> in there?And how would I prompt it with a previous history of the conversation?
<|prompter|>hello</s><|assistant|>
And got back this (truncated): hey there!</s>
introducing the first of its kind, a revolutionary AI...
Why did that </s> come back as part of the response?<|prompter|> is the human message (including the instructon/contexts from my tests) and <|assistant|> is the bot's response.
There are llm wrappers that will save chat histories and reformat them appropriately for different LLMs. For instance, see TUI and its instruction template presets: https://github.com/oobabooga/text-generation-webui/tree/main...
<|prompter|>... prompt goes here</s>
<|assistant|> ... assistant response will go here, maybe followed by </s>
Still don't understand why I got a rogue </s> in the reply when I ran a prompt through it (but it continued after that tag): https://news.ycombinator.com/item?id=38101766And is my guess about </s> at the end of the assistant thing correct? If so then a conversation might look like this:
<|prompter|>hi</s>
<|assistant|>Hi to you too. What do you want to talk about?</s>
<|prompter|>Tell me facts about snakes</s>
<|assistant|>
But the documentation on these models is so thin! Do I need that </s> at the end of the assistant things or not?Welcome to LLM research land, lol. Mostly there is no documentation, no reproducability, just random hit-and-run releases with some claimed metrics and if you're lucky a model card with the trained prompt.
> Do I need that </s> at the end of the assistant things or not?
...I dunno. I didn't add it in my prompt template, and you can see that the model itself it outputting that token in the responses. Maybe ask the trainers on the HF page.