Considering the AWS API is essentially open for everyone to start a transcription service, what exactly is the difference here. If you know what you're doing you can build this in about 4 hours.
That being said, I am skeptical about the quality and would like to see some demos. Audio recordings of meetings are especially difficult to transcribe accurately.
The only real innovation here is when this is combined with language learning apps to help me practice my Chinese pronunciation, but even then I know I'll have to look to hire a tutor soon.