Depends on the application.
If you can process "offline" for an hour and then cache the results, CPU inference is fine.
GPUs are expensive.
If you can process "offline" for an hour and then cache the results, CPU inference is fine.
GPUs are expensive.
The downside is that it makes development and testing super slow without the speedup you get from having local GPU power.
> GPUs are expensive.
Depends on the GPU. I've found T4 GPUs to be cheaper than CPU compute on AWS when testing throughput per $ of spend.