ParentFull threadsearealist·Just enable MTP on llama.cpp and you will get the same decode speeds.View on HN