Tested: Mixtral 8x7B vs. GPT-4 for boolean classification
jshelbyj.medium.com
jshelbyj.medium.com
However, I should say these benchmarks are not meant to be academically meaningful. My testing was strictly done on my real world use cases. I didn’t start with the intention to do these tests, but in working on my personal project I realized I could convert my integration tests to output test results.
I performed these tests using a 3090 and llama.cpp.
> Give me a list of 5 words having 5 syllables each
Mistral 8x7B and GPT 4 both choke on this frequently.