Who's listening? Growing privacy concerns around transcription services (2020)
rrj.ca
rrj.ca
We’ve been working on it for over a decade and we did add machine transcription recently, but I still find a surprising number of users use and ask about the local-only “Self Transcription” mode.
One interesting side-effect of a local-only mode is that we can’t sync transcript data between computers. And this sometimes surprises people. Once we explain that this is actually a feature not a bug, since we don’t store the data on our servers to be able to do the sync, it actually seems to reassure people even more.
The privacy policy looks halfway decent, a bit chaotic with no ToC and free-floating "legal basis" & "retention policy" that are independent from the concrete processing tasks, leaving them as abstract "we know the law" blocks of little value. The policy is very dominated by all the user behavior tracking and ad-tech integrations, while the processing of recordings/transcriptions is rather short. In a way that is good, but it is almost too short, given that it is the primary function: "transcripts" and "recordings" are not even a category of data in their policy, and they are not mentioned with any sub-processor, not even by the ones running the servers (a rather curious formal error, as one would assume a focus on their core business). There is a complete lack of "we keep small samples to improve the ML/AI" which i find inteersting, but which might be true.
There is a small note in the legal texts that the user must have the consent of the recorded people before uploading recordings to the server for automatic transcriptions, but there does not seem to be one in proximity to the upload function. (or at least they are missing from the screenshots in the guides, i did not register for a trial)
I hope this is the correct way to evaluate these kind of things?
If you gather confidential information do not hand them over to some third party for transcription, storage, sync, cloud office solutions, etc.
I feel that this is quite a basic truth and packaging it in such a long essay makes me think that most people on HN might not be the articles target group.
It was still fun to read though.
The relevant key elements for the topic are all there: undisclosed third parties, outsourcing to cheap labor, spy agency backdoors, governments wanting blanked access, subpoenas against sources and journalists, bad tech, bad laws, data breaches, hacks and overly broad clauses in the service contracts. But it is too shallow, not intended for those who already knew that and hoped for details and depth.
For a moment the article even hit my advertisement detection due to the way the "list of transcription services" at the end is formatted, how brand names are sprinkled in, how the author tries to avoid making it too one sided (which is a good thing!) and how the investigative journalists approach is related to overpacking and old fashioned motels. I suspected a bait and switch: bait with a cautionary tale, switch into a sales pitch. But the author says he does not endorse any of the named services and the short reviews focus on problems with their policies.
Overall i find it ok, even if it is a bit too heavy on the narrative of this journalist protects his sources, be like him for my taste.
EDIT: i ran the literal quotes through google and i think these are interviews backing the article as primary sources. This is no bullshit, it's good and honest work.
...I immidiatelly close the site and make a mental note of the "dont open links from this site again".
Sometimes I wonder if I am the only one. I came to read info, not authors attempt at a short novel.
Personally this narrative oriented style has me on edge about rhetoric subterfuge. It feels like a sleight of hand, like a Kansas City shuffle, like a trick to fly fiction, an advertorial, or opinion below the radar.
Fyi, it's not a fad. It's what's called a "human interest" angle and some publications (like this rrj.ca website of this thread) have that editorial allowance for personal storytelling. (I made a previous comment about this: https://news.ycombinator.com/item?id=24270673)
More examples of that include "The New Yorker", "Harper's" and other literary style magazines.
The opposite examples of "just the facts" type of publishers would be "lwn.net" for Linux tech articles or The Economist magazine for business stories. You won't find articles beginning with personal things like, "When I was a little boy, my father took me to discover ice...blah blah blah...."
The problem is that HN is an aggregation site that gets the [human interest] articles when some readers don't want it and we have no convention for tagging them to avoid annoying that subset of the audience.
This particular investigative journalist is over-preparing for interacting with sources. That is a literary device to color his avoidance of the services named in the headline.
This particular cliche has grandkids by now.
I loved my grandmother, we’ve spent such a great time together as we used to do our daily walks in the green park in front of our house and we watched all those dogs running around and birds chirping, while enjoying our ice-cream that we had acquired from the shop around the corner.
Then I wake up, and I go on Hacker News to read about all those interesting and innovative things that happen around me all the time, while my little dog wags his tail and watches me as I click every link on the front page.
Thank you for reading this.
I try to add as much “human” as possible into my writing and tech interaction.
I feel that we have allowed technology to leech our humanity, and the consequences have been devastating.
When we have people running multi-billion-node datasets, that have no empathy for the essential “humanity” of each node, we can have Jurassic-scale disasters.
Overpacking a story with junk (how ironic!) makes it harder to spread the important ideas of the article, which hurts the author's goal.
Unfortunately, many large media companies have adopted the story-first-facts-second strategy. Are those who prefer otherwise such a tiny minority?
To me, these articles look like those SEO recipe sites that are stuffed with random content because Google won't rank them as well if they just provided what the user is looking for.
This assumes information can always be cleanly severed from the story. That strikes me like cutting a paper down to the abstract and conclusion. Yes, the methods may be tedious to get through, but someone with an understanding of them sees the problem with more depth.
But for a story about how chemical X is bad for your health, you don't need to read about Suzy and Michael, their new house, what they're going to name their baby and how they decided on painting the house in Suzy's grandmother's favorite color... to learn that some paint includes a chemical that is bad for your health.
I get the impression that there are multiple things at work: some people really like those side-stories, and they're very easy to write. You'd have a much, much harder time filling a newspaper with facts, which makes it much more expensive.
But the story isn't chemical X is bad for you. It's about how Suzy and Michael, within a specific set of circumstances, had bad things happen after being exposed to chemical X.
Depending on who you are, those details could be important. If chemical X is known to be harmful, it opens up details into how Suzy and Michael got exposed and who exposed them. Is this a local problem or a national one? Are there alternative explanations for the bad things that have nothing to do with chemical X? The human interest details, meanwhile, clue you into socioeconomic and demographic factors which may be at play. (But also may not.)
These sorts of stories are surfacing anecdotes. Hopefully fact-checked anecdotes. But not peer-reviewed papers, either. (All that said, I personally prefer publications that tend towards terseness.)
That's a very specific story then, and that's not what I'm talking about.
I'm talking about "eating poison is not good for humans". Suzy and Michael are humans, therefore eating poison is not good for them, but they could be replaced with any other human: they specifically don't add anything to the story. They're an emotional connection for the reader at best and a filler at worst.
I found it's very liberating to consciously decide "this isn't intended for me" and focus my attention elsewhere. Still, it's very easy to lower my guard and fall into the mindless consumer trap again. But it's an indulgence, and I try to limit my intake.
Perhaps if someone wrote a blog in nested-comment style I would read it.
Fun though experiment: could Harry Potter be rewritten in comment style?
I would not comment, if it were somewhat disrespectful to the source and context of the article. It's a journalism oriented blog, and in this case, its target audience is practicing or student journalists. So the story is painted to relate to that audience's concerns.
Also giving such a 'Pro tip' is equal to teaching HN readers how to use scrolling...
I'd also note that ML transcriptions are useful in the context of this article, i.e. to get rough cheap transcripts. But, if you value your time above minimum wage, human transcription is a lot better if you actually want to publish a transcript of an interview.