I think it's a vllm vs llama_cpp performance thing, will pay more into it.
One note I had between the two is that gemma has a much higher prefix cache hit rate in general.
One note I had between the two is that gemma has a much higher prefix cache hit rate in general.
No comments yet.