The AMD Epyc build is severely bandwidth and compute constrained.
~40 tokens/s on M3 Ultra 512GB by my calculation.
The AMD Epyc build is severely bandwidth and compute constrained.
~40 tokens/s on M3 Ultra 512GB by my calculation.
1. M3 Ultra 512 2. AMD Epyc (which Gen ? AVX512 and DDR5 might make a difference in both performance and cost , Gen 4 or Gen 5 have 8 or 9 t/s https://github.com/ggml-org/llama.cpp/discussions/11733 ) 2. AMD Epyc + 4090 or 5090 running KTransformers (over 10 t/s decode ? https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...)
If the M3 can run 24/7 without overheating it's a great deal to run agents. Especially considering that it should run only using 350W... so roughly $50/mo in electricity costs.
I'd assume this thing peaks at 350W (or whatever) but idles at around 40w tops?
I don't know how to calculate tokens/s for H100s linked together. ChatGPT might help you though. :)
If this is remotely accurate though it's still at least an order of magnitude more convenient than the M3 Ultra, even after factoring in all the other costs associated with the infrastructure.