Full threadfnetisma·What would be the difference in compute for inference on an audio<>audio model like this compared to a text<>text model?View on HN