re: Whisper v3 -- how is this possible? Whisper has a 30s context window. You have to chunk it.
Plus you could do actual batch inference instead. Or if you must carry forward the context you could still do it linearly, but the mem usage shouldn’t just explode