Otter.ai has saved reporters hours transcribing interviews. Caveat emptor
politico.com
politico.com
The underlying story is that there is no story: the author recorded a sensitive conversation and sent it to Otter. Otter sent him a survey in response and referenced the title of the recording. The author grew concerned that his recordings had been shared with China. Otter scrambled to respond, and eventually reassured the author that his recordings had never been shared with any foreign governments or law enforcement agencies.
The underlying point, that using transcription services puts journalists' sources at risk, is well taken. But the story unfairly tries to make it sound like Otter has a problem, when the author is really just using Otter as a vehicle for their story.
It seems like they never before considered the implications of using the service and sending interview contents off-device at all.
If this is representative of the trade, I think as tech we have a lot of education left to do for everyone's safety. And if it is, good on the author for making it transparent. I feel like they learned a lot researching the article and basically put their own inadequacies on display for public scrutiny, which many others may not do.
Consider this next time you update your company's Privacy Policy and wonder if it's understandable, to who, can be discovered at all, etc. Do your users understand what of their data you have and how you use it?
And I think creating (or mandating) that transparency is probably a precept for offering an alternative product that runs on-device and doesn't have the same trust limitations - because the value prop has to be understood to make it a business. This article may inadvertently bring about a market for "safe translation app for journalists". We need the public to start thinking in these categories if we want a product landscape with both.
Also, when you wrote "precept", did you mean "prerequisite"?
But I don't think "your data never leaves the phone" is something that only journalists and people with similar strong privacy requirements would or should value. All other things being equal, I would always prefer an app that keeps my data local, on principle. The difference being a journalist would make is that I would enforce it, and really journalists should know how to do that. They should be disabling network access, both cellular and wifi (which you can do on Android but I'm not sure you can on iOS) for any app they enter sensitive information into.
Apart from that, I'm assuming that there's an enormous cost sunk into the training of the network. I imagine that alone would make Otter hesitant to let the trained network out of the house, even if it was technically possible.
- the general quality of education we're giving journalists about privacy and security.
- whether or not Otter.ai (and other tech services in general) have a responsibility to turn away journalists that they think might be putting other people in danger by using the services.
- whether we in the tech industry are being irresponsible by just constantly assuring the public that AI-driven services don't involve interaction with human employees/contractors.
- whether we should have a set of general, open standards for how journalists conduct interviews that will help avoid problems like this.
- if you're being interviewed by a journalist about something sensitive, should you inquire into what their setup is, and should you explicitly make requests like asking them to promise not to use a 3rd-party for processing the interview.
And so on. I couldn't agree more, I don't think it's a non-story, I think it's a very important story. Otter.ai is a small part of it though, and the story is more about the context/environment that causes a story like this to be written and that causes this to be news.
----
> And if it is, good on the author for making it transparent
That's also a good point; part of the way you change journalist attitudes towards data security is you get journalists to research it and then talk to each other and very transparently say, "hey, these things we thought were safe aren't, and we are making mistakes with data right now." It's a lot more powerful and effective for a journalist to say that about themselves and about their field than for us to criticize them for bringing the issue up -- even if it feels at points in the article like the author hasn't really grasped all of the problems with what they're doing yet.
The cynic in me says that the lack of transparency about this is intentional; look at the outcry that happens when people learn about fb/meta snooping through their device, or the various apps turning on cameras and microphones when the user doesn't expect it. It's not that people don't care or have given up on privacy - it's that a lot of times developers lie to the user about what's happening, or at the least obfuscate it. Look at how torqued off people were when Apple forced apps to actually notify the user what they were doing!
It took Otter.ai three months to confirm the email was genuine - having initially said that it wasn’t.
It feels like, after the first candid response, somebody higher up went "oops, this looks bad, deny everything!", and eventually retracted it because the guy wouldn't give up.
Looks like the sort of misunderstanding that Pied Piper might find themselves into, to be charitable.
Quick Poll: Is there anyone here (on HN) who thinks that the details of Otter's privacy policy are relevant to whether or not an interested A-list national intelligence community accessed the author's data on Otter?
[/cynic]
> When I asked Otter to clarify whether it shares user data with non-U.S. government or law enforcement agencies, the answer wasn’t comforting. “We disclose Personal Information if we are legally required to do so, or if we have a good faith belief that such use is reasonably necessary to comply with a legal obligation, process or request,” Denise Mutch, an Otter service team member, told me via email on Dec. 16.
Later in the article, a different Otter representative (Mitchell Woodrow) says the opposite, but that indicates to me that they don't have a consistent internal policy for whether to turn interviews over to totalitarian regimes, which is a huge, HUGE problem.
Even Signal/WhisperSystems, who I have a great deal of respect for, would without question cave under government demands which come from the sort of people who've learn't from their failure to force Lavabit/Levison's compliance and silence. I have no doubt that successors to Levison's refusal to exploit their own previously-secure systems at the Government's request have been successfully threatened and manipulated into doing so using much more powerful threats and consequences that Levison's option of "OK then, I'll just shut my whole company down".
If a sufficiently powerful TLA decided I needed to be surveilled for a serious enough suspicion of a serious enough crime, I have no doubt that I'm "STILL GONNA BE MOSSAD’ED UPON"*
The chance of _any_ non-privacy-obsessed tech startup/company like Otter deciding to lawyer up and fight the US government to protect _my_ individual privacy as a low-paying customer is literally zero. The _best_ I can expect of them is to have a solid process in place to ensure the people demanding my data have the appropriate paperwork and jurisdictional oversight in place, more realistically I'd expect them to have a "law enforcement portal" like Facebook/eBay et al. where bored cops can log in and stalk their ex girlfriends.
* https://www.usenix.org/system/files/1401_08-12_mickens.pdf
With that said, the core idea is great. These companies are starting with transcription but it's not difficult to envisage a future where (a) supplemental information is presented depending on conversation topics, and (b) actions are taken on your behalf. Example: no more need to manually send a calendar invite, just mention it in the conversation and the "AI assistant" will schedule for you.
Well, not that shocked. I'm sure your uncle was generally being diligent. But really, you shouldn't have been able to find out what he was working on.
Definitely a grey area, but the point was more that literally anyone off the street can sign up to gig for one of these services and receive access to a bunch of potentially pretty sensitive audio.
If you have good relationships in the courthouse you can luck out with complex cases that result in appeals.
I really think on-device models like we see in Android's Live Caption tool are a major privacy boon, and they're starting to reach an acceptable level of performance in Google's case. The main pathway to better performance is loading ever-more-massive models into memory, which isn't feasible for mobile devices but could be done on people's laptops in a meeting.
https://www.nytimes.com/2019/05/07/opinion/google-sundar-pic...
[1]: https://ai.googleblog.com/2017/04/federated-learning-collabo...
I saw in a lawtech startup, when i arrived they were transfering data unencrypted over the internet. this data was everything todo with a matter, including documents. Suffice to say they were lucky i had experience building software on Secret networks and patches those holes quick smart (ie. re-wrote there entire stack which was built by 3rd-world-outsourced-devs who have zero care for your busisness continuity.
No wonder scamming is outpacing the drug buisness these days. People are dumb and the internet exposes dumbness to desperateness who are more then willing to take up the slack.
The opening stanzas describe a conversation over Signal, followed by a message from Otter.ai. But there is no initial disclosure that the journalist sent the conversation to Otter. The way the story is written, it sounds like Signal was hacked or his phone was exploited.
Later, he states:
> "Otter said that the fact that Aksu was the focus of the survey was only because I’d entered his name as the recording’s title."
So OK, you submitted a recording to Otter? Why not disclose that you used Otter intentionally:
> Apparently I wasn’t the only Otter user worried about this kind of scrutiny. “This survey has been discontinued over concerns that some customers (such as yourself) may include personal or sensitive information within the title of the conversation, and including this information within a survey may cause some concern,” Lai said.
> The next day, I received an odd note from Otter.ai, the automated transcription app that I had used to record the interview.
Archive.today link demonstrating that this isn't a stealth edit by Politico: https://archive.is/etLpm
Did you enter "Mustafa Aksu" as the title of the Otter recording? If so, they just took the text you entered and emailed it to you. Why did you write this sensationalist article?
If they transcribed it and send it to you, maybe there's something here. But expecting Otter to be secure is an entirely different issue. But the article conflates this whole thing about the title with insecurity and it's unclear.
There is certainly an article to be written about the dangers of using Otter. But this stuff about the email and the title is a distraction. It attempts to "storify" something that doesn't need it. Automated recording means sending recordings to the cloud where you don't control them. That's a terrible idea for sensitive content. Full stop. That's the article.
Disclosure: have used Otter, like it, but wary of security issues. Self hosted option sounds good.
Returning information to users can be a surprisingly delicate matter. E.g. HIV clinics and mental health services (in the UK) are careful about the information they put in appointment reminders. Sure the patient signed up for the appointment, but they don’t necessarily want a voicemail or text message to mention the type of clinic.
81 employees and no cybersecurity team from what I can tell. linkedin[.]com/company/otter-ai/people/?keywords=security
Security and privacy isn't a tool, it's a set of tools in a logical stack and used according to correct processes. This is why focused newsroom security teams are so critical, but NYT fired their rep a few years ago and WSJ had or has generic security hires double-timing in that role. It sucks that Khashoggi wasn't more of a watershed moment for this.
I wouldn't worry about privacy, I'd worry about threat actors popping otter and it's 0 security team to expose sources.
The careers page is hiring a software eng - security which means someone who can build tooling, but won't have any policy/process power beyond what a VP Eng allows them. So good luck deploying necessary changes that that impact dev/product flows.
This is the sort of stuff that gets security/privacy people upset at ex-Googler SV tech. Lot of hubris.
Is anyone aware if there's been a similar movement to create on-prem versions of popular ML-based products? Otter, Grammarly, etc.
mp4grep --model ~/apps/mp4grep-0.1.1/model/ --transcribe filename.mp3
The output is more intended for captioning so it's lots of short phrases with timestamps and no punctuation, but it'll give you a quick taste of what Vosk can do.I use LTeX [1] for VSCode, which sets up a local LanguageTool server, and its resource usage is quite significant. (I use it together with their n-gram data sets [2])
[0] https://dev.languagetool.org/
[1] https://github.com/valentjn/vscode-ltex
[2] https://dev.languagetool.org/finding-errors-using-n-gram-dat...
You may have put their lives and the lives of their friends and family at risk.
Next time Phelim Kine asks you for an interview, will you agree?
"Safe transcription app for journalists". Go do it.
Right, but in the meantime, if a government is hunting me and a journalist asks me to interview them, I want to know I'm secure. I don't want that to be reliant on whether or not a market I know nothing about exists. Definitely, it looks like there's a need for a safe transcription app for journalists. But that's not a license for them to be dangerous while they're waiting for that product to exist. Journalists have a moral responsibility to keep their sources safe in the world that exists today.
Otherwise, the takeaway seems to be that nobody with really sensitive information or in a vulnerable position should talk to journalists until after the tech industry builds an entirely new product, which probably isn't the outcome that anybody wants.
We need to hammer out some degree of data security best practices that journalists won't break even in the instances where it makes their lives more inconvenient. Ideally, we should try to hammer out best practices that make it possible for a journalist to evaluate whether or not a service is appropriate to use. Otherwise we'll just play this endless game of cat and mouse where someone builds a transcription service that's private, and then another company builds this data-leaking collaboration or markup platform and that hole gets opened... at some point we have to teach journalists that there are certain types of services that they should stay away from unless they meet certain criteria, and it sounds like a lot of journalists don't have a good grasp on how to do that kind of security evaluation.
Certainly, there's at least lack of education here about why we're using E2EE and what specific attacks it guards against if it's not ringing alarm bells with journalists when a service asks them to upload the raw interview to a remote server.
Everyone has a smartphone. I can guarantee that at any scale, there is some nut recording their entire day. I’ve worked at places where that’s a real threat that is difficult to manage.
The reason why we advise E2EE is to make sure that the recording is only unencrypted and listenable on devices you control, it's not just because we like Signal. Transferring it raw to another service invalidates the point of doing your call over something like Signal in the first place, because the purpose was to not have the recording sitting unencrypted on someone else's servers that you don't own or manage.
My impression is that if a journalist is doing E2EE and then sending the recording off afterwards, they probably don't understand why that advice was given and they probably don't have an understanding of what the risks are we're trying to protect against when we tell them to do remote interviews over an E2EE service. We want journalists to be thinking about their source's privacy on a level beyond "I use Signal because I was told to."
I get what you're saying -- E2EE is not a solution to every problem, and you're right about that, there are other security risks that are in separate categories. But what I'm getting at is not to say that it does solve every problem; just that if a journalist doesn't understand not to send recordings off someplace, to me that indicates a lack of knowledge about other security concepts.
Iirc, Dragon is $500 and human services charge anywhere from $0.25-$7 a page of text depending on the quality level and turnaround time. If you have access to tech people, you could easily and more securely leverage AWS or GCP services as well.
I think it’s appropriate to blast the author as if you’re in the business of interviewing sensitive or high risk sources, you really have no business using something like Otter or Grammarly.
Otter is the only thing approaching affordable for individuals, but I don’t really trust it given they’re playing fast and loose with user data.
There are some various other services along similar lines to Otter.ai but don't remember names off the top of my head.
Ostensibly almost all newer/flagship iphone/android devices can do offline speech to text with passable accuracy. On Android it is (I suspect intentionally) nerfed to an accessibility feature called "live transcribe." It transcribes everything star wars style on the screen, but you can't save the audio or the transcription, you have to copy/paste the output which just disappears when you close the app. It's arbitrary and infuriating that it's like this.
I'm generally a "hater" of the cloud, and I have to admit that Otter is very good. I used it speaking with someone using airpods on a shitty internet connection, with no noticeable drop in accuracy. Even spelled uncommon last names correctly.
What I don't get is why none of these free apps can take a pre-recorded audio file as the input. Even for the paid services, this seems to be more limited than using the functionality 'live.' Is there some technical reason for this, or is it a sales department decision?
Probably not, the attack vector of an apk of dubious provenance is probably that it sends all of your transcriptions to some 3rd party server. There's some non-zero probability that it does that, and it's possible you could run it in a network-less sandbox somehow to mitigate it.
But with the saas service, you know for a fact it's going to a 3rd party server, who can store it for as long as they want. And if they aren't hacked now, they could be in the future.
A 3rd party sending around a sketch apk that calls back to them is _definitely_ going to try to do something bad with it.
Otter is unlikely to do something bad with it (it'd be bad for their business). So you should really compare the odds otter gets hacked vs the odds the apk is compromised.
Shorthand
> Forward suspicious email to spoof@paypal.com
That's where the story rises above the mundane.
If support at your company got an email saying "omg! sensitive information in our recording is being referenced in this random survey we got from you!" I guarantee 2/3 companies would flub the initial response because the support person doesn't understand the technical side of the issue which is that titles are used in one of the survey templates
Journalists who are casual or reckless with this kind of data shouldn't be in the business, and companies that don't make these kinds of risks apparent to their user base, or go to lengths to disguise a clear conflict of interest as bad support should not be in that business either.
The shocking revelation should be that journalists who normally are so careful they use end-to-end client-side encryption for their communications are using third party transcription services that have full access to their sensitive interviews. That's a faux pas and critical security blunder on the journalist's part, not otter's, though one learning here is maybe otter should consider offering some sort of secure enclave service for situations like this with additional guarantees and client-controlled encryption keys.
"Not respond to that survey and delete it" is also the ending of a sentence that begins "if you would feel more comfortable, you can".
Where are you getting your information from - is there a link to this correspondence, in the article or elsewhere?
'Is also' has very different implications than 'could have been'.
I work with a geographically distributed team and it doesn't do an amazing job with English in various accents — I've noticed Google Recorder on Android is better in this regard — but it's good enough.
Nice product.
The author apparently then recorded that call and uploaded it to a cloud transcription service.
Surely any journalist reporting on human rights abuse by state actors should know that recording the interview on an internet connected device is a risk let alone uploading it to the cloud.
> Using encryption on the Internet is the equivalent of arranging an armored car to deliver credit card information from someone living in a cardboard box to someone living on a park bench.
Why the F would you even bother with Signal if you're just going to record the call to some other third party service? This is an enormous and stupid mistake.
If you're a reporter working on human rights pieces which involve nation states, you need to step up your game here or you're putting sources at risk.
Otter completely messed up and transcribed it as “I Fuck That” (not kidding) within Zoom.
And that’s pretty bad when it’s a bunch of Middle Schoolers…
If you have Microsoft 365 via your employer, you can transcribe up to 300 minutes a month. I use it a lot for recorded work calls, as my memory is awful.
I just wish they'd roll it out to the Desktop version of Word, as the online version struggles with transcripts longer than 50 pages.
https://apps.apple.com/de/app/ada-dictation-speech-to-text/i...
> Until those laws change, journalists and others who rely on transcription apps need to carefully consider the potential dangers.
What this article is revealing to me is that journalists who are dealing with sensitive information aren't well-trained enough to recognize all of the dangers in the services they use. Knowing what I know about AI-driven services, even really large ones like Alexa or Siri, it would never be acceptable to me to use a remote service to transcribe an interview that absolutely had to be private and that hadn't already had all of the sensitive info redacted. It's a lack of knowledge about how AI works and how these networks are built and maintained (real people do run into the information, and there are bugs even in giant products).
It's also a lack of knowledge about the security capabilities of the company: how is this information being stored, what encryption is being used, has the company been audited? It's really irresponsible, but I also believe the author when they said they just never thought about it before, I believe that it probably never crossed their mind that a big company might misplace data or that it might be forced to hand that data to a government.
What's frustrating is the qualifier "until those laws change". There are a lot of potential risks to using a remote transcription service even if the laws are different. For ordinary people, legislation around privacy is important. But when dealing with information of this type, you need better security fundamentals. So this feels like the author still doesn't realize just how dangerous they're being by using a service like Otter.ai for interviews where someone's life might be at risk.
----
On that subject, I also don't think calling out the journalist in this way is victim blaming; I think the victim here is the person being interviewed, and I think journalists have a moral responsibility to understand data security because the people who talk to them are trusting them to be able to keep that information safe. If you can't do that, then it's irresponsible to handle sensitive information that can harm other people. There are risks associated with handing any remote company an unencrypted interview with sensitive information that they will feed into an AI network, regardless of that company's intentions. That is not something a law can fix, you as a journalist need to be able to protect the people who trust you and you need to know the limitations of a company saying "we won't share this", you need to know what the inherent risks of the technology and process are instead of just trusting them.
Not to let Otter.ai off the hook here; if they're aware of the fact that people are sending interviews where information leaks could be dangerous, they need to do a better job of educating their clients about what the risks are. The tech industry in general needs to do a better job of education here, we share some blame. What we are seeing is the effects of Google/Amazon/Apple creating an inaccurate picture of what AI is and of how private it is: we create this narrative where people assume that Siri/Alexa are just isolated boxes that sit in a vacuum where nothing can touch the data. And it turns out that's not accurate at all, metadata (like transcript titles) gets leaked from those services, people review transcripts and use it to keep training the AI, there are bugs that leak data across accounts. And there are real-world consequences to us training the public to disregard the possibility that those leaks can happen; consequences like journalists believing us and using our technology in irresponsible ways.
I support privacy legislation, but privacy legislation will not make it OK for journalists to be irresponsible with sensitive information. There needs to be more of an understanding that for journalists who have clients who are in a position of extreme trust with them, some behavior is just inherently risky. It doesn't matter if there's legislation, if you're a journalist it's still not OK to send sensitive information over SMS, or do unencrypted phone calls to conduct sensitive interviews, or to send sensitive interviews to corporations that you don't have an extremely close relationship with. You have to learn data security, I'm sorry. To the extent we in the tech industry can help with that, it should be by making it clear that we mess up a lot and we leak a lot of data and we aren't magical wizards that can just decide to keep data safe -- all of that boasting about our privacy policies are just narratives we use to get more people to be more comfortable talking to their smart homes.
----
> In the three months since that initial exchange (and there was more to come), I’ve gone down the rabbit hole — talking to cybersecurity experts, press freedom advocates and a former government official — to try and understand what vulnerabilities and risks are present in this app that’s become a favorite among journalists for its fast, reliable and cheap automated transcription.
This doesn't require a complicated rabbit hole of research, I can tell you right now that if the person you're interviewing is being hunted by a government, if they are "a wanted man", you do not put their interview on any service that is not fully end-to-end encrypted so that even the service can not access the data, period. We have enough technology and good enough encryption tools that you don't need to take that risk anymore. That means transcription, (regardless of whether it's AI-driven or manual) needs to happen on your own hardware.
If it's not a sensitive interview, or if a leak would just be embarrassing or cost someone some money -- then sure, maybe you have different standards for that kind of information, use a remote transcription service for those interviews if you need to. But not for someone who's wanted by a government.
Then you could take the text and put it through the same process that distorted the audio and get the original text.
If you talked to your plumber and said you have a "web server" in your basement, he'd think you give funny names to spiders. The concept of "web server" is a technical shorthand for "computer attached to a telecommunication system, typically the global internet network, that is dedicated to responding to request from other computers for hypertextual media to be delivered over such network". But webserver is shorter and everybody in your sector understands what it is, so you use that. Sure, you could say "I have a computer connected to the network that does stuff", something that "would get the point across to more people", but it would be a very rough and incomplete image.
In the same way, people in literature, rhetoric, law, and medicine, use a lot of Latin words to convey specific meanings succinctly. Anybody worth their salt, in the sector, will appreciate them more than "plain English".
Except that is not the case here. "Buyer Beware" is more succinct and in the same language as the rest of the headline. It's also not a domain-specific term, so shared convention is out. The author is sprinkling on some Latin to seem smarter.
There are also proprietary ASRs which can run offline.
I wrote a small python application to evaluate some of them. Link in my profile!
> NSA whistleblower John Kiriakou and Guantanamo Bay detention camp whistleblower Joseph Hickman have both accused the same reporter accused of revealing Winner's identity, Matthew Cole, of playing a role in their exposure, which, in Kiriakou's case, led to his imprisonment.
I don't see where Kiriakou and Hickman were sources to The Intercept.
The [35] links to https://www.peterbcollins.com/2017/06/30/in-depth-interview-... which says:
> In the final 5 minutes of the interview, both men share their stories of being burned by "journalist" Matthew Cole, whose work at ABC compromised Hickman and sent Kiriakou to prison. Cole was involved in the recent episode at The Intercept, where NSA leaker Reality Winner was apprehended after Cole shared her leaked document in an effort to verify it.
The [36] links to https://www.cbc.ca/radio/asithappens/as-it-happens-tuesday-e... with (emphasis mine):
> He later spent two years in prison for disclosing the identity of a fellow CIA officer to then-freelance journalist Matthew Cole, one of four authors of Monday's Intercept report. While Cole did not publish the name, his email exchange with Kiriakou was used as evidence against him.
So while I see Cole burning several sources, I only see The Intercept burning one.
I wanted to know more about the other sources they burned, because I hadn't heard of additional cases and wondered why they didn't learn from their experience.
if you talk to a reporter, 100% assumption your full name/full transcript/full video recording is going to be out there.
"off the record" never meant anything.