Your next conference should have real-time captioning
lkuper.github.io
lkuper.github.io
Caption all the things. There are so many benefits that it's just daft not to.
Actually it was. I discovered American TV shows on the internet (well.. had no legal way to access them) Started watching first season 24 with subtitles, they were like training wheels for ~20 eposiodes. Then I managed to ditch them and started listening (with a finger on a shortcut to jump 10 seconds back).
So, please have text transcripts if possible. They make a huge difference for non English speakers.
Contrasting that to German English speakers whom likely had high quality English classes throughout education speak with a specific accent and rhythm closer to that of the German language.
I don't know if their are technical reasons for this, but anecdotally Romanians had said they learned a lot of their English from watching dubbed TV and movies.
Someone else in this thread mentioned the show 24. It's fascinating that, indirectly, watching pirated copies of a typically Hollywood entertainment television show (which academic and intellectual communities may find "crude") can open up a wealth of knowledge and culture (English language only resources) for people whom, potentially, wouldn't have had the educational opportunities otherwise.
It makes my spine tingle when stories like this flip your understanding and perspective of things like the Hollywood Entertainment Machine. Things have value in unexpected and fascinating ways.
I knew a guy at university who had an ok understanding of thw english language with a fairly Chinease accent, but would occasionally say strange slang informal words that no one has said since Friends aired.
Hence, people speak good English with a slightly american accent. We're exposed to it pretty much all our lives.
No doubt. Not sure if they ever made a conscious choice of "we'll teach our people English better" or they just took the easier route. But watching American shows in English is really invaluable to learning American conversational English. A huge part of it is always cultural references and sayings that just never translate directly things like "Don't put the cart before the horse" saying, of even stupid movie and pop culture memes like "Hasta la vista, baby" (Terminator 2) That is stuff one has very little chance of learning in school (grammar and vocabulary always come first).
So I would say transcription is a must have and even real time transcription is a feature a large portion of attendees will benefit from, not just those who can't hear well, unless your conference space is fairly small or your sound system is truly excellent.
I guess it's also worth considering why the sound systems for these conferences are so epically bad. For example, what's with passing around microphones when we are all already carrying our own personal mic?
Though I'll be the first to admit that I'm in awe of the skills of their stenographer; I'm not nearly that fast, and rely on other meeting participants to fix up my typos (especially in code) in our collaborative editor, https://etherpad.mozilla.org/
The only thing that seems to be missing from the !!Con transcripts are timestamps. I'm not familiar with Plover so it might already be a feature, but being able to output one of the standard subtitle formats (SubRip, timed text, SSA) would make remuxing an MP4/MKV of each talk with subtitles much easier.
YouTube will sync the transcript to the audio, (removing the ambiguity of it having to guess what is being said, as you're telling it that - so now it knows what words it's listening out for) and you can download the resulting automatically timed file as an SBV or SRT type file.
It's not 100% perfect, and is something of a hack, but it usually works pretty well. :)
Better still, it's now open source - although sadly the voice recognition part of it relies on a closed source commercial product which (at first glance) might not even still be available.
Dangerous in the wrong hands, but interesting: http://subtitling.sourceforge.net/
Sometimes the acoustics of the room are bad, or there's background noise, or maybe you just missed a word and need to have it repeated. All of these problems would be solved by this system.
This kind of thing would benefit all conference goers.
What software/sdk's have you used?
There are no automated solutions that can universally do a good job of transcribing natural speech from people who aren't specifically "speaking to be recognised", if that makes sense.
Maybe one day, but not yet. It's a problem waiting to be solved, so the reward for the first who can really crack it will be substantial.
It works to some degree, and has the advantage for the subtitling companies that respeakers are easier to train and don't need to be paid as much as a proper stenographer. The disadvantage is that the output is much slower and subject to a rather greater delay. There is still nothing that beats steno - but respeaking is cheaper and people don't complain enough about the inaccuracies.
https://github.com/hausdorff/bangbangcon.github.io/blob/gh-p...
For instance, she'll separate out conversations:
>> Can you move the mic closer to your mouth? >> Yes. Is this better? Is this better? Okay. I will talk like this, then. >> You can move the mic. >> Like... >> Take it off the stand and hold it up to your face.
She can also figure out when something is an acronym (like LARP), make sure everything is capitalized correctly (Python, Ruby), separate out what's being said into paragraphs when the speaker starts talking about something new, and a ton of other things.
"Automatic speech recognition is not currently a substitute for human transcription, because computers are unable to use context or meaning to distinguish between similar words and phrases and are not able to recognize or correct errors, leading to faulty output. The best automatic speech recognition boasts that it's 80% to 90% accurate, but that means that, at best, one out of ten words will be wrong or missing, which results in a semantic accuracy rate that's often far lower than 90%, depending on which word it is."
(this is a subset of the answer to "Will speech recognition make CART obsolete?")
My favourite example of misrecognition is one of Travis Goodspeed's talks, where YouTube's VR output "Geek women are expensive, but not prohibitively so."
Voice recognition is OK but it falls a long way short of a usable level of accuracy, and even the accuracy it can muster goes right down the toilet if there is any background noise or music, or if the speaker is in any way unclear (accents, rushed or slurred speech, etc). There's a long way to go before you can just get a usable transcript of speech automatically.
Quite a lot of voice recognition engines seem to have been trained on thousands of hours of C-SPAN or Meet The Press or something, because when recognition conditions get challenging, some engines start to degenerate into outputting nonsense like "congress Muslims Kenya capitol great today Cheney".
There is no substitute for a human pair of ears and a lightning-fast means of text entry like a steno keyboard - nice to see Plover getting a mention in the source article too.
The article began with people looking into captioning as a means of meeting an accessibility need (a speaker losing a hearing aid) and ended with the discovery that there are many other benefits beyond accessibility. It was a great read and everyone won. Many planners never get to that point.
The lack of widespread real-time captioning isn't a technical or cost issue, it's an education issue.
There is an oddity in the UK where a person with a hearing impairment can get speech to text paid for, but they are the client and they get to control what happens to the text. It would be useful if text to speech paid for by the state and used in public meetings was automatically given to the meeting as a whole as well as the person as an individual.
(I just thought if you make the plug outrageously shameless it will pass for a joke, still being a plug though)
Jokes and plugs aside, it can produce a graph of audience's per-moment sentiment, that can be overlaid on the transcript/video, so you know how they felt at each moment of the talk.