Full threadsmokel·This seems to be a wrapper around llama.cpp with several tuned parameter settings. Most of the speedup is gained by enabling speculative decoding [1]. No amazing breakthroughs, just good tuning![1] https://en.wikipedia.org/wiki/Speculative_decodingView on HN