ParentFull threadam17an·Honestly you can run this on a 16GB VRAM GPU with llama.cpp. Just try it!View on HN