Automatically transcribe an interview, meeting or video
voicedocs.com
voicedocs.com
Recording conversation means saving time of an expert in exchange of additional time spent by the student, when looking things up in it. Having good notes/docs is easier for students, but more expensive for the expert, who needs to spend more time to organize the information properly.
So depending on what you're doing, there might be different tradeoffs.
Sure, searchable conversation data are better than nothing. But it is, by definition, disorganized. I worry about the future where people will stop making notes/docs just because they can record everything.
Oh wow was I surprised to see the quality. All of the cloud providers are abysmally bad at transcribing German.
I believe the reason is that in German, you can make up word combinations on the fly and use them as valid nouns. And people do that, if it's convenient or if it enables you to be more precise.
"Dampfschiffahrtsgesellschaft" = Society (Gesellschaft) for Driving (Fahrt) of Boats (Schiff) with Steam (Dampf)
In the case of pronunciation this primarily poses a problem with detecting the intended word, but in other cases "cleaning up" the output may lose contextual information (e.g. what a speaker was going to say before cutting themselves off and using a different word). This is difficult enough for a human to get right, let alone a machine.
They trip up on technical jargon but handle everyday conversations just fine, including speaker detection, punctuation, idioms, etc.
But that's also a slightly different use case, where each speaker is in their own (somewhat) quiet environment and on separate connections (and thus audio tracks).
It's much harder to do all that after the fact, like with a recorded video.
I find Trint.com, which is partially automatic, to be good for that... the AI does a first pass, and a human cleans it up afterward. YouTube has a similar assisted-auto feature for their captions, minus speaker separation.
In this context, Gesellschaft translates to Company. (GmbH=LLC)
The spelling also depends on whether you're talking about the historical Erste Donau-Dampfschiffahrts-Gesellschaft or any generic Dampfschifffahrts-Gesellschaft – note the ff vs. fff in middle; the old company name retains its pre-1996 spelling.
Donaudampfschiffahrtsgesellschaft without hyphens was as far as I can tell never officially used by the company, but used informally as part of the name of the Donaudampfschiffahrtsgesellschaftskapitänstango, a 1930s song.
SteamBoatDrivingSociety
Which actually made complete sense even in English!
Due to much more complicated grammar German is much more difficult to learn than English, but at least the spelling is easy.
I wonder why more languages never try to simplify their orthographies. Children could spend years learning useful things, instead of wasting time on spelling.
Controversial opinion here: they should have removed ß (sharffes S) completely. It is still used in some relatively rare cases.
The first is to record the sounds you hear. Look at a common stenographic "alphabet" (often called "shorthand alphabet" though that practice is essentially dead) or at the keyboard of a stenographic machine.
Then the stenographer reads the output (either hand or machine generated) and writes a text using a combination of cue (from the paper) and memory.
This is quite different from trying to do straight text-to-speech.
I won't upload recordings (with possibly sensitive information) to a third party.
You'll need to sign a DPA with them to be compliant with the GDPR tho, and they'd need to disclose where the data will be stored and processed and how they maintain control over that data if it's a third party.
Apart from Itep Pictures, all others seem to be in Turkey. Are you based in Turkey?
If yes, allow me to place zero trust on everything-Turkey, under the current government/leadership. I strongly believe that Turkey lacks the basic/fundamental freedoms and rule of law is going whichever way this regime's leader wants it to go.
I would similarly hesitate to upload such data to Iran, North Korea, Syria.
If no, where are you based?
As it was already said you should make clear which languages are supported.
And I think you should put prices in USD and/or Euros instead of TL (turkish lira), ideally Euro's for european visitors and UDS for the rest of the world. Besides the free tier, if I'm serious about the service I will be less keen to test it out before knowing the cost of it and at first I've seen the price without looking too much and thought it was pretty expensive before understanding it was expressed in TL's.
We provide the same functionnality (except for the Word export)
+ direct recording and upload from Zoom, Hangouts, etc.
+ video / audio editing & sharing by high-lighting which part of the transcript you'd like to keep.
www.spoke.app :)
In 70 languages (see language list here: https://spoke-for-sumo-lings.webflow.io/)
But, with the way insurance has been going in the US lately, you better be recording and transcribing that call. Usually, if the call line is recorded (basically all US health insurance companies do this) you can legally record the phone call without permission from the other party.
I personally have an NVIDIA Jetson AGX Xavier with AI tools for speech-to-text, person identification, and transcribing, which I use for important phone calls. I use my own AI tools and devices for privacy reasons.
They also struggle with domain specific jargon depending on what data they were trained on. While manual transcriptions will mark ambiguous utterances as such (or ask for additional information), automation can create a false sense of certainty while just "guessing" whatever it matches most closely. This is a hard problem and unlikely to be solved soon.
ML transcriptions are fast/cheap and they're fine if you mostly want to pull out some quotes or check some things in your notes. But, in general, I find they're not remotely worth my time if I'm going to publish a transcript in which case I get a human transcription. (And even that can be a bit tough with accents, technical jargon, overlapping voices, etc.)
Some tech conferences were pretty good about hiring actual people for live captioning, which was great, but with conferences mostly happening online via video streams at the moment, automated captions and transcriptions might seem like an obvious choice if you don't understand the limitations.
There’s still quite a bit of value-add possible on top of that, however. The ability to edit transcriptions is a great start, especially if you maintain timecodes against the media. Developing or curating domain-specific language models to improve accuracy is also a likely option. There also appears to be a lot of interest in using real time transcription to augment live events with content derived from the conversation.
Good luck!
Apart from your service offering an onprem option is there much else difference?
That being said, I am skeptical about the quality and would like to see some demos. Audio recordings of meetings are especially difficult to transcribe accurately.
The only real innovation here is when this is combined with language learning apps to help me practice my Chinese pronunciation, but even then I know I'll have to look to hire a tutor soon.
The privacy policy also doesn't provide all the information the GDPR generally requires you to provide, e.g. spelling out users' rights under the GDPR and what legal basis is given for collecting each specific piece of information.
I'm mostly pointing this out because it could get them sued, but I'd also expect a company based on a service like this to take privacy a bit more seriously, or at least present themselves as if they do so.
You can get sued for omitting such a page (by any bored lawyer really) because it's considered anti-competitive and a misdemeanor: https://de.wikipedia.org/wiki/Impressumspflicht#Ordnungswidr...
Here's a lengthy explainer of what should go in a privacy policy to be fully compliant (in German), note that "clear and precise" language is generally understood to mean being explicit about the legal basis (i.e. parts of the GDPR) under which the data is collected and processed: https://www.datenschutz.org/datenschutzerklaerung/
In any case, your privacy policy link on the German language version of your website gives me the policy in English, which violates the GDPR's requirements for "clear language" regardless of the actual content by not being in German: https://voicedocs.com/de/legal/privacy-policy
But to be honest, you shouldn't be asking a random person on HN, you should talk to a lawyer.