ParentFull threadsuperkuh·That's not how llama.cpp works. It's a layer split. The GPUs handle a few layers and the CPU handles the rest. The GPU layers no matter how fast they complete still have to wait on the CPU layers.View on HN