ParentFull threadlogicallee·Awesome! I tried it and after loading (which took a while as it is a large model) got 12 tokens/second and very coherent output. Great demonstration.View on HN