is there any integration with it to langchain?
+
Is there any optimization for LLM to run on RTX cards? 40XX,30XX I found out tha LLAMA.CPP is nice but I want to take advantage of my graphic cards also, and didn't found any documentations...
+
Is there any optimization for LLM to run on RTX cards? 40XX,30XX I found out tha LLAMA.CPP is nice but I want to take advantage of my graphic cards also, and didn't found any documentations...