Specs say this runs an AMD RX-421BD. This is a 2015 AMD CPU with 2 bulldozer cores and a tiny IGP.
...To be blunt, you would be much better off running LLMs on your phone. Even an older phone. Or literally whatever device you are reading HN on. But if you insist, the runtime you want in MLC-LLM's Vulkan runtime.
You'll see it over and over again when you're looking for help, be careful, it's 100% a blind alley in your case. It's very likely you'll be disappointed by MLC as well, simultaneously it's your only real option. You definitely won't hit 1 tkn/sec, and honestly, id bet 0.1 tkn / sec
Sorry for reddit link.
To be blunt, there isn't much interest in support outside of apple/nvidia. There is a WIP Vulkan backend, but (last I checked) progress is slow and its not optimzed for IGPs either.
MLC-LLM is much more promising once its features get fleshed out, as it "inherits" support for many devices from its Apache TVM backend.
That's your problem. I googled and it looks like one of these all-in-one appliances like a drobo or whatever's popular these days. That's not a server. (At least, I wouldn't call it a server. It's an all-in-one appliance, or toy, depending on perspective) And yegods, that price...
Spend $500, get an actual computer, not some priced up appliance, and you'll have a much better time. Regardless of if you spend it on more CPU or more GPU. You can get a used computer off ebay for $100 and shove a $400 graphics card in it. Or maybe get a ryzen 7 7700x, I'm looking at a mobo+cpu combo with that for $500 right now.
Finally, to make sure this response does contain a answer to what you asked: ;-)
if you can run this stuff in a container on your appliance already, but it's very slow, congrats! I'd call that a win. I looked up the chip, the RX-421BD, it's of similar power as an Athelon circa 2017. I think my router might have more compute power. You _do_ have those 512 shader cores, given effort, you could try and get them to do something useful. But I wouldn't assume it's possible (well, maybe you don't mind writing your own shaders ;-)). Just because the chip has "some gpu" doesn't mean it has "the right kind of gpu you'd need to hijack for lots of matrix multiplies, without writing the assembly yourself".
Sorry this isn't more helpful, but it's the truth.
I hadn't even noticed that, I just saw "I could build that for 1/4th the cost"+"wtf, only 4 drives?"
Stuff like this prompts a dual response in me. :)
It always gives me a strong urge to educate "it's not that hard, and it's fun" to build it yourself.
AND it always makes me kick myself for not commercializing the expandable media servers I started building in the early 2000s, for me, for the dorm, for my friends, i.e. exactly the people you identify. :)
Unrelated, cool, I hadn't heard of any of those 4 programs, I'm googling now and some look useful. Thanks! Possibly saving me some time in my next project...
[1] https://github.com/ggerganov/llama.cpp [2] https://huggingface.co/TheBloke [3] https://github.com/abetlen/llama-cpp-python