Sure, there are many ways one can be accidentally inefficient, some examples:
Batching: Your throughput (and thus cost per token) will be much worse if you don't do batching (meaning running inference for several inputs at once). To do batching, you either need to have a high sustained QPS or adjust your workload to be bursty (but this has utilization issues).
Inference engine: The regular HuggingFace, even with built in optimizations, is not competitive against inference engines like vllm, exllama, etc.
Utilization: Depending on your scale, it might be hard for you to have a nice flat utilization, or to even utilize one GPU at 100% capacity. This means you need to solve scaling up & down your machines and the issues associated with it (can you live with cold starts? etc.)
Hardware: Forget server-less GPUs like Replicate etc, their markups compared to on-demand pricing is usually >10x. On-demand A100 or H100, provided you can even get a quota for them, are also expensive. Spot instances are better. Older GPUs (A10, T4, etc.) are better. Your own GPU cluster is likely the best if you have the scale (and likely, you can resell your cluster within the next 18months without much if any depreciation).
For these reasons, I have been toying with the idea of providing a dead simple service for fine-tuning & serving open-source LLMs where the users actually own (and can download) the weights. If anyone is interested in this and would like to chat, let me know.