245 karma · joined May 14, 2021
I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for you
https://github.com/ggml-org/llama.cpp/pull/25680
Also, for Qwen, the 4 bit _XL quantization seems to have a good balance of performance to size.
archived versions here: http://archive.today/8LX8s http://archive.today/00mIw
https://duckduckgo.com/?q=https%3A%2F%2Fservury.com%2Fdatace...
It is not visible in the live webpage.
As major differences I'd highlight: local and offline, so drawings not sent anywhere trained on artist work with explicit consent
In Krita's case, they claim the AI isn't generative so it doesn't add detail.
Whereas the AI today is trained on stolen work and often on the inputs as well.
It remains to be seen how good these custom models are.