Nice work with this. I was wondering, are all computations other than convolution performed on the FPGA as well - such as pooling, padding, inter-layer quantization operations (rescaling & offset additions)? If not, does the FPGA offload unsupported operations to the host before continuing?
Does the FPGA need to transfer intermediate layer IO data back and forth between the host during GEMM if the data become too large to fit on the FPGA SRAM?
Thanks