I'm just starting to get into downloading and testing models using llama.cpp and I'm curious which model you're actually using, since they seem to come in varying levels of quantization. Is this [0] the model page for the one you're using, or should I be looking somewhere else? What is the actual file name of the model you're using?
[0] https://huggingface.co/TheBloke/OpenHermes-2.5-Mistral-7B-GG...