How does this differ from anything llama.cpp offers, regarding offloading layers? The repo consistently refers to "DDR4". Is there a reason DDR5 won't work with this?
Oh, that's just the infra for the infra. Then use something like graphllm from matteo, and of course llama.cpp from greg, tailor you model selection to your hardware.