-masroor (author)
-masroor (author)
As an aside, I took into account the resource allocation in the parent comment. The c5.2xlarge has 8 cores, 8GB RAM [3] and does a single fp32 inference in ~17ms. If we chop that down to 4 cores and assume linear scaling we can fathom running ResNet-50 in ~35ms compared to the ~500ms achieved here. I'd recommend comparing to a known baseline rather than a "vanilla setup" to ensure you aren't missing any simple changes that may dramatically improve performance.
[1] https://github.com/IntelAI/models/blob/master/docs/general/t...
[2] https://www.intel.ai/improving-tensorflow-inference-performa...
-masroor(author)
[1] https://github.com/tensorflow/serving/tree/master/tensorflow...