Just to add some ballparks: I have a yolov5 model that I run in PyTorch on a 3080 and the inference time is roughly 20ms.
And inference with a 200estimator random forest is a few ms on a cpu.
I can imagine that in some scenarios it’s not enough though.