ParentFull threadrom16384·Using llama.cpp on a 128 GB server running codellama-70b-python.Q4_K_M.gguf I get 1.3 tokens per second which is too slow. With Nous-Hermes-2-Mistral-7B-DPO.Q5_K_M.gguf I get 8.3 tokens per second which is usable.View on HN