Say_what: Using speech-to-text to fully check out during conference calls
github.com
github.com
I do kind of like the idea of dialing into a meeting but saying that you'll be AFK and then just reading the transcript afterward. It seems like it could be a real time saver in those cases where you're not really participating currently but want to keep up with the project.
Yep, very. Just imagine!
"So, guys, the new feature we onion maybe needs to beat the liver by the twenty and. Josh, you we'll get the baby to Reese pony ability customer. Let me know if super busy ved communication."
>Yep
The only reasonable answer to this question
http://www.dmlp.org/legal-guide/recording-phone-calls-and-co...
> Eleven states require the consent of every party to a phone call or conversation in order to make the recording lawful. These "two-party consent" laws have been adopted in California, Connecticut, Florida, Illinois, Maryland, Massachusetts, Montana, New Hampshire, Pennsylvania and Washington. (Notes: (1) Illinois' two-party consent statute was held unconstitutional in 2014; (2) Hawai'i is in general a one-party state, but requires two-party consent if the recording device is installed in a private place; (3) Massachusetts bans "secret" recordings rather than requiring explicit consent from all parties.). Although they are referred to as "two-party consent" laws, consent must be obtained from every party to a phone call or conversation if it involves more than two people. In some of these states, it might be enough if all parties to the call or conversation know that you are recording and proceed with the communication anyway, even if they do not voice explicit consent. See the State Law: Recording section of this legal guide for information on specific states' wiretapping laws.
That said, IANAL.
At the level of your personal rights, however, you have an absolute right to observe, record, copy, and display anything you are able. Personally, I feel the correct response to anyone who tells you that you cannot archive what you see and hear is laughter, followed by an explanation of their rights.
The legal question is whether text-to-speech qualifies as a "recording" when the audio itself is not stored.
If the law were as technical as you describe then VoIP calls themselves would be against the law since a person's voice is recorded, transmitted, and momentarily stored.
Nowadays I'm watching alot of tutorials on youtube. What I like to do is downloading all that stuff with youtube-dlg (https://github.com/MrS0m30n3/youtube-dl-gui) and watch them with increased speed on vlc (factor 1.5 to 2.0 works reasonably well). It always feels very satisfying to think about the time saved. Like that brain-learning interface in Matrix.
Would love to have something like this for everything. meetings, school, conferences, news, etc..
Not sure if this is not just the same thing... If it is, I would love to hear your experiences.
Edit: shame I didn't publish it :) Edit2: I had one call for which I have a colleague speaking who speaks 'the Queen's English' and same result...
why not run the audio through all the tts engines you can find and compare results. keep what has the highest matches and see if you can get the other words using phonetics engines.
for words that are specific to your company train a tts engine on those words and weight your local engine higher than other services.
you can get a pretty decent translation back, but who would be so lazy as to work that hard..... ;)
The key is using a conference call system which captures each track separately. Polycom, Tropo, and others do it. Then when you run ASR, you can feed in each channel and have a "unified" transcript.
And yes, I've worked on this problem at my last startup where we did all of that and took things further by adding search. Here's a demo: http://clarify.io/try-it-now/?id=20f5e8c4f1ef4f8c91160839b48...
> I do run the risk of losing credibility as to whether I'm actually listening in meetings now, but I'm not too concerned about that. I made this as a joke, and my coworkers know that
http://www.businessinsider.com/josh-newlan-creates-say-what-...
From the "Readme.md":
> Installation (OS X)
> 1. Sign up for, install, and run Splunk Enterprise
> ◦ This has to be enterprise; the HTTP Event Collector feature used here doesn't exist in light
Kudos, to Splunk for having the guts to "allow" Newlan to unleash "Say What?" on conference calls everywhere, and cheers to Newlan for dreaming it up.
EDIT: formatting, readability, punctuation.
My initial thought was, "who would choose Splunk for the datastore for anything?"
This now makes sense.
edit: by "anything", I mean "anything that isn't logs/security event data"
Pretty fun hack though.
"Sorry, I didn't realize my mic was on mute there."
"No problem Josh, we can hear you now just fine."
"Sorry, I didn't realize my mic was on mute there."
"... um, Josh?"
"Sorry, I didn't realize my mic was on mute there."
On fb we already have bots posting articles and people asking for a bot to deal with the deluge of articles. Bots posting for screener bots.
how so ?
If it highlights the complete waste of time that almost all conference calls are, it would be a service to mankind.
I recently wondered if there is a a video-chat app that uses voice recognition for live subtitles? Would be awesome for talking to people who don't have perfect hearing. Could this be hacked together in a similar way?
0: arxiv.org/abs/1604.08242 1: arxiv.org/abs/1512.02595
It's got to be worth $10-15/hour in wasted time and recapping.
"Oh, you thought it was due on Monday? Nope, transcript says you promised it by close of business Friday."
Kudos, sir. Kudos.
It's a funny gimmick, for entertainment and to get visibility for the author's employer's name.
Also, I wouldn't call it merely a slight convenience when multiple times a day you see the pattern repeat of there being a long pause after you ask someone a question, followed by "sorry, I was on mute."