Full threadalzaeem·if using deep learning models, consider using distilled and/or quantized models to reduce the resources required for inferenceView on HN