Last time I took a look at TensorFlow Lite, they had a vision where you would export your model into a .tflite file (which is FlatBuffer encoded execution graph with weights) and then use it on mobile for inference like this, in pseudo-code :
model = tflite.interpreter().load_model("my_model.tflite")
model.set_input_data(my_input_buffer)
model.execute()
model.get_result(my_output_buffer)
Which is nice, since you can easily update the model by simply distributing a new file "my_model.tflite". The TFLite interpreter library would use whatever capability (SIMD instructions, DSP cores, etc) is available on the device to accelerate inference, so the application developer doesnt have to worry about writing different code for different platforms or even understand how the prediction works under the hood.Is Qnnpack a library directly competing with TFLite? Are the file formats used for the model the same between tflite and this? Does it support TensorFlow Cores created by Google for inference, and/or more generally specialized cores like Qualcomm's DSP?