Hetzner’s New AX102 Dedicated Server with AMD Ryzen 9 7950X3D
hetzner.com
hetzner.com
Hopefully these come to more than the 1 data-center! :) Hopefully we see more competitors offering great stuff like this. Wild that this easily a sub $2000 box, $700 cpu, $200 ram, couple hundred for ssd and mobo and case.
Still shocked after more than a decade, Economics 101 hasn’t kicked in and put more downward pricing pressure on the major cloud providers.
Provisioning is relatively fast and everything has an API. If you are brave enough you can just build your own server and have them co-locate it for you too, making advanced machine learning clusters possible aswell (albeit of course with more upfront investment).
I have seen startups go through their 250k cloud bonuses in 3 month because of unattended scaling, I recon on hetzner, you would have a couple of years playroom =)
But then how will you get invited to the next AWS conference to talk about your (self-inflicted) problems and then write about it on your company’s “engineering” blog?
[1] https://www.reddit.com/r/hetzner/comments/wucxs4/comment/ilf...
I'm not an ai expert yet but:
1. Is this processor suitable for that?
2. Is it feasible to customize one of the LLM models directly on that machine or retraining the model requires more power?
Park the processes that don't need it on the smaller-cores.
https://man7.org/linux/man-pages/man1/taskset.1.html
These commands exist in both Linux and Windows for a reason.
Process #0 is on HardwareThread#0 through HardwareThread#15 inclusive. These will run faster because of L3 cache (96MB IIRC).
Process #1 is on HardwareThread#16 through HardwareThread#31. These will run slower because they don't have the cache (only 32MB).
Load balancer checks queue-depth of Process#0 vs Process#1 and distributes as appropriate. Done. You probably should be doing this under NUMA-conditions anyway (which will have different memory-access speeds and different PCIe access speeds / asymmetries. NUMA#0 will talk to PCI#0 faster, and if PCI#0 is Ethernet, then NUMA#1/#2/#3 will be penalized), but its no different for asymmetric cores.
Turning off NUMA reporting doesn't make NUMA go away. Whether you can buy hardware that doesn't need this kind of optimization really depends on your needs and budget. The 7900X3D and 7950X3D don't have a very large niche, IMHO; but if you were thinking of a 7900X for a VM system, and you have some loads that would benefit with more cache and some that wouldn't, it could make sense. Depending on your needs, you might be better served with a 7950X or maybe two 7800X3D systems or ??? there's lots of options. I think if you're going with Epyc, you're needing or wanting lots of cores, lots of memory channels/capacity, and/or lots of pci-e; if you need that in a single machine, you're just going to have to deal with the fallout of NUMA that results, because it's not like you have a choice. Sure, if you can manage to split up your load into multiple small UMA machines, that'll be easier to manage, but then you do have to manage more machines and balancing at different levels. It's all tradeoffs.
Its cheaper to own/maintain 50 servers, rather than 100 servers. Running 2-processes per server is just cheaper.
You have half the motherboards, your RAM is consolidated, your SSDs / Hard Drives / Ethernet ports are consolidated. This leads to lower power and more efficiency (its easier for threads to migrate over if possible).
EDIT: AMD CPUs with nonuniform cache (such as this one), with the right motherboard, support L3 as NUMA, which means that each CCX is exposed as a NUMA domain.
But I don't have such a system, nor have I played with the BIOS settings of that chip. But it only has one memory controller, so I doubt there's a NUMA setting. Correct me if I'm wrong of course.
There are very few workloads where you will have less per-core performance on a 7950X3D than, say, an EPYC CPU. Even those relatively lower clocks are pretty high by server standards.
Do companies like google and Facebook replace 100% of their servers all at once? Or do they upgrade in stages, allowing some requests to be processed faster than others?
It's fairly likely the high cache cores would be good for PostgreSQL, while the standard cores would be good for general application servers.
I'd be trying these out if we needed new dedicated servers atm. And might do so depending on how things go over time. :)
Amazon Lightsail
Atlantic.Net
DigitalOcean
Google Compute Engine
Linode Shared
Vultr
kamatera
lunanode
upcloud
I tried to give them the benefit of the doubt, but once they asked for a photo of my drivers license, I gave up. no thanks, I will just stick with Digital Ocean.
I've been using them for a few years already and only have good things to say.
I am also unable use Hetzner, due to their overreaching security process.