How is the $ per inference, say 4k tokens, on a Hetzner box vs an A100?
Not all tasks require low latency.
Not all tasks require low latency.
If your usecase fits inside that 32GB (no 70B models, sadly) the price to performance of a GGUF Q4KM is really attractive on this setup.