Just made https://feycher.com thats similar, but has realtime lip syncing as well. Let me know if you are interested and we can chat
44 karma · joined February 1, 2016
on tesla-t4-30gb-memory-8vcpu google cloud
on tiny and tiny.en
for 10 minute = 30 seconds
on medium
for 10 minute = 1m 30s
for 60 minute = 7m
on large
for 60 miutes = 13m
on NVIDIA GeForce RTX 4090
on tiny
for 10-minute = 5.5 seconds
for 60-minute = 35 seconds
on base
for 10-minute = 7 seconds
for 60-minute = 50 seconds
on small
for 10-minute = 14 seconds
for 60-minute = 1 min 35 sec
on medium
for 10-minute = 26 seconds
for 60-minute = 3 mins
on large
for 10-minute = 40 seconds
for 60-minute = 3 min 54 secInteresting, which model are you using? We use the medium model which is the sweet spot between time/performance ratio. We also chunk, We try to detect words and silences to do better chunking at word boundaries but if you do more chunking and you don't get the word boundaries right it seems like whisper loses some context and the accuracy suffers. We will soon support longer hours. We just want to make sure the wait time for transcription doesn't suffer for most users. But great demo, reach out to me if you want to collaborate