Deepgram – Find Damning Soundbites
blog.deepgram.com
blog.deepgram.com
Coming to you 2020: president candidate X may or may not have said that he hates women, kicks babies, love hitler, and plans to nuke Wales. You don't even care any more because your airwaves have been saturated with both candidates (and their Vocaloids) saying all kinds of crazy shit in disinformation and counter-disinformation and counter-counter-information campaigns and everybody's desensitized.
http://talkobamato.me/synthesize.py?speech_key=77ff00cb8af50...
Greetings Professor Falken
Hello
A strange game.
The only winning move is not to play.
"""
[1] https://en.wikipedia.org/wiki/Utau - which incidentally was the software that created the voice for that erstwhile Nyan Cat meme
[2] http://store.vocallective.net/album/trumploid-vs-sanderpoid-...
On the other hand "audio searches reveal the presidential candidate is telling a massive lie about always having expressed profoundly opposition to Controversial Thing X, and actually repeatedly ridiculed the cause in not well-reported meetings prior to becoming the candidate" could and should be an electoral issue.
For better and (mostly) for worse, dredging up muck from student days has been around since the dawn of politics, and at least in the UK we're comfortable enough about the fact people change for even a Conservative Party chairman to be willing to volunteer they were a radical Marxist in their teens. Being able to rapidly cross-reference everything a politician has said on record on a given subject in recent years is a new and far more useful way of subjecting them to scrutiny.
This is less about the specific case of Trump's sex remarks (which was obviously a case of somebody knowing their secret video was dynamite and waiting for the right time to light the fuse) and more about the difficulty of establishing whether Trump was telling the truth about having opposed the Iraq war all along, which is where efficient search of stuff that's already in the public domain but not necessarily in the public consciousness comes in really, really useful.
http://www.businessinsider.com/bridgewater-records-conversat...
http://www.businessinsider.com/ray-dalios-bridgewater-manage...
And so far it has been used for people who knew they were being recorded, so don't talk like its recording private conversations and therefore over-exaggerating the issue.
What checks can be put in place for this?
Fraud, misrepresentation, drug use, sexual harassment, conspiracy, aiding and abetting, stolen property, just off the top of my head.
Drawn from financial records, statements, documents, phone calls, social graphs, etc.
You're coming across as rather too narrowly focused.
In this case someone is indicating that in the past they've sexually assaulted someone.
The speech is just evidence pointing towards the act, and the act is the problem.
Cardinal Richelieu wrote some six lines or so on this subject. You should look them up.
https://en.wikiquote.org/wiki/Cardinal_Richelieu
(Or perhaps not, the authenticity is, as many things are, disputed. But he carries the blame. Which is perhaps the most poetic justice of this point. Or is it injustice. Words, words, words.)
Which tells you that this problem is not new by any means.
I searched for oil, mexico, automatic weapons, isis... it found "crisis", a reference to oil without actually being on the correct time stamp, mexico got several 50% results none of which were isis. automatic weapons got "nuclear weapons". so basically 0/4 searches.
pretty sure that most networks use the closed caption data to find clips. NBC had an indexed system for all of their tapes in prototype in 98 for their whole back catalog. Not sure if it ever went online.
Voice stores distressingly cheaply in terms of space, and with the Internet of (broken) Things (that spy on you), odds of finding yourself surrounded by microphones in the most unexpected locations,[2] controlled by a wide variety of quite probably competing interests.[3] And if they cannot find what they're looking for in the surveillance tape itself, they'll simply manufacture their own evidence using your own phonemes[4] and video.[5]
________________________________
Notes:
1. https://twitter.com/pinboard/status/732985370204233728
2. http://www.inquisitr.com/3097029/government-surveillance-in-...
3. http://www.locusmag.com/Perspectives/2016/09/cory-doctorowth...
4. http://www.theatlantic.com/technology/archive/2016/09/hackin...
> In the case of Google, that processing must take place on
> Google's centralised servers
Doesn't recognition of the initialization phrase "OK, google" take place on the phone? Sending a continuous stream of audio back to google servers sounds expensive.I should have mentioned voice-activated televisions as a whole 'nother class of attack.
The title is clickbait, but the claim is rather interesting. It says that it has 80 % accuracy for transcribing an audio clip compared to 20 % for speech-to-text.
http://www.zdnet.com/article/microsofts-newest-milestone-wor...
When people talk to Siri, Cortana, or Dragon, they take unnatural care in the clarity of their speech compared to normal talk only meant for humans. Also the speaker may not have been speaking directly into a microphone, lots of background noise, etc.
All of these factors probably combine for a much lower accuracy than what Apple, Microsoft, and Google are going to be dealing with in usual cases. Also keep in mind they all have incentive to inflate their own products accuracy score. Not that the same incentive doesn't also exist for this company/product.
If you ran phone calls or tape recorder audio through a speech-to-text engine then the word accuracy rate is like 10-50% (i.e. abysmal). When you try to search for keyphrases like "frolicking kitten" the likelihood of a text match with STT is ~20%.
If you ran that same search with Deepgram then 80% of the time you'd find what you are looking for since Deepgram doesn't have to guess at what is being said, it takes the inverse approach and matches 'how it sounds' using deep learning voodoo magic™.
Cool demo, looking forward to seeing more detail about what is going on. However I would quibble with the STT WER quoted above. Maybe in noisy environments with unknown speakers (and no voice normalization) this is accurate, but the kinds of clean speech in the demo perform really well in modern recognition engines (on benchmark data, to be fair c.f. MSR 6.3% and IBM at ~6.9%).
Most word searches over speech to text work over soft matches (or ideally beam search over most likely partial phoneme/word part matches), rather than hard matches so it seems like a bit of a straw man comparison in this case.
[0] http://research.google.com/pubs/pub42543.html
[1] https://arxiv.org/abs/1510.01032
[2] https://sigport.org/sites/default/files/gloveNNLM_kaudhkhasi...
[3] https://arxiv.org/abs/1502.03044
[4] http://www-personal.umich.edu/~reedscot/files/icml2016.pdf
http://money.cnn.com/2016/10/09/media/billy-bush-nbc-today-s...
> Women have always been the primary victims of war. Women lose their husbands, their fathers, their sons in combat.
> Marriage has historic, religious and moral content that goes back to the beginning of time, and I think a marriage is as a marriage has always been, between a man and a woman