Social network ad targeting can “listen into” conversations using “live photos”
twitter.com
twitter.com
I'd actually probably get a lot of use from a feature whereby live photos displays a quick transcript of any words it picked up in the recording over the top of the image, or in a prompt, any time you try to share one with clear button to remove it. Similar to the local voicemail transcription feature, then I wouldn't have to waste time playing the recording. You could probably add the same detected words to the photos search index too.
Here's my script for the curious:
#!/bin/bash
#extract mp4 files from Google Pixel 6 .MP.jpg files
set -e -u -o pipefail
extract () {
dd if="$1" of="$1.mp4" bs=$[$( strings -td "$1" | grep ftypisom | head -n1 | cut -d' ' -f1 ) - 4] skip=1
}
for fn in "$@"; do
extract $fn
doneMore info: https://support.apple.com/en-us/HT207310
... _and_ a still photo. You don't lose the full-resolution still photo, and the still photo isn't simply a frame from the video. In case anyone's curious.
In other words the video is exactly the same quality as the photo in each frame so you can pick a moment a second later for example if you film something fast moving.
Does anyone know how it works and isn't gigabytes in size?
You can change the keyframe on a live photo (ie the image displayed first) from the still capture to a video frame, but I generally would never do this due to the quality drop. It adds a little more memory context to a still photo, its not a burst capture tool.
There are actually some features with live photos I hadn't seen before though that are kind of interesting, for example you can fake a long exposure shot via a feature that interpolates information from the video stills into the main image:
https://osxdaily.com/2018/01/10/take-long-exposure-photos-ip...
So these Live photos are supposedly not prerecoded - except they actually are. They are a prerecorded video that you send to someone. There is nothing live about it.
You just fell for the marketing.
You’re right that there’s probably a video format under the hood but the experience to meant to be closer to a still picture with a bit of added magic (a la Harry Potter). The clip is too short to catch much action or have (intentionally) meaningful audio.
bemoans?
So that's how it is in their family... /Rooney
1) photos permission includes access to full file for each photo, including geotagged metadata and any audio from "live photo"
This seems desirable and working as expected. You might legitimately want to share the live photo, or the photo with metadata. However, it would be nice if the photo picker asked you whether you wanted to share with metadata. (Note: unlike the photo picker in native apps, the Safari file picker, ie the HTML5 <input type="file" />, does strip EXIF data by default.)
2) it's easy to take a live photo without realizing it, and similarly easy to share it
This could be improved when selecting photos within the photo picker. There should be an option to separately select the static or full file (maybe with a long press).
2) granting access to "all photos" doesn't alert you when an app accesses a photo without you selecting it for some reason
This used to be much worse, when the default and only option was to grant access to all photos, or no photos. Ever since Apple added the option to select which photos an app can access, I think this is less of a problem. But there is an education issue; people don't realize that "all" means the app can truly read all your photos, including geotagged metadata and audio of live photos, even without you selecting one to "upload" (or "do something with," depending on the purpose of the app). Even as a relative expert, I never understood this until I saw a demo app (to prove the privacy issue) that looped through all your photos and displayed your geotagged locations.
But even if people did realize this, there is no indication that an app has accessed a photo, like there is with the icon indicating the microphone or camera was used recently. Perhaps a solution to this could be changing the photo picker to overlay a green dot on any photo the app has accessed.
I mean, it makes sense, but ouchiewauva, how much is that misstating the typical conscious intent of the user :O
Any link to this demo? Would love to use it as example
EDIT: Found it: https://krausefx.com/blog/ios-privacy-detectlocation-an-easy...
Also worth checking out the author's other projects: https://krausefx.com/projects
I actually don't doubt it, I'd just like to see some kind of evidence it's being done.
And when you think about what people take pictures of (their parking spot, selfies, nudes, landmarks, birthday cakes, sunsets, cats), what's heard is likely not even relevant to the picture taker's life or interests. If I look at all of the photos I've taken in the last two weeks, I've got:
- Cat (2) - Building (1) - Stuff in my home (6) - Selfie (4)
All thirteen photos were taken in ~silence.
I kinda agree, still worth it to mention to your family or whatever some kinda safe word or whatever. I speak several languages so I guess any scammer impersonating me would focus on one, but who knows you can also make it speak any language I guess.
In order to be safe you'll have to pre-establish a safe word only you and the other person know in order to avoid fake scams etc
Q: Why would anyone do that? A: It's to train an ML model, maan.. to train an AI, maaan... to target ads, maaaan..
Ultimately this seems like far too much hassle and processing for very little gain, given the 2 seconds of audio in a live photo is bound to be useless for determining product relevance. Advertising algorithms are clever enough to target well without having to go to the expense of analysing real time audio from your mic or your photos & videos.
"Caused by a bug"
I've used them deliberately perhaps twice, and wish I knew how to force all the far more numerous accidental uses into normal jpgs as I don't want to waste any of the limited iCloud storage space on video I'll never watch.
"Basically, Live Photos on iPhone is a regular digital photo and a short 2 seconds video recording the few moments before the photo. Essentially, the iPhone camera captures 2 media files, one photo and one video. When viewing a Live Photo, the operating system, iOS plays the video file first and then it shows the picture."
The part that makes me feel old is that I have literally no idea why this sort of thing would be useful or desirable.
It also lets you pick a different key frame, so if e.g. you get someone blinking, you can pick a different frame from the video to use instead.
And if you capture photos back to back such that the live portions overlap, you can convert the group of photos into a single contiguous video.
Note that this significantly lowers the photo quality.
I don't use them personally, though, but I'm a curmudgeon about my camera controls ;)
I think Google added something similar in their Pixel camera app.
It’s a fun gimmick especially with kids as 2 seconds goes from not seeing you to noticing and lighting up with a smile.
That's what I was picturing. I'm glad my impression isn't far off. Perhaps it's a "you have to use it to understand" sort of thing, but I truly struggle to see the value.
It doesn't matter, though. I don't use iPhones and even if I did, that I don't understand it is meaningless.
https://en.wikipedia.org/wiki/Burst_mode_(photography)
Adding sound can make these bursts more pleasant. But there is no easy way to disable or remove the audio.
Some apps are friendly about it. I appreciate that Apollo gives me the option of which browser to use for opening links, and that it distinguishes between "in-app browser" just as it does between Chrome, Firefox and Safari.
For a long time I've had an idea for an app that can take advantage of this, but for the user's benefit. Imagine a newsreader app with modular news sources, where each "source" is a client side script that runs within the in-app browser, to navigate to a URL and then "parse the content" by getting rid of ads, paywalls, etc. So unlike a typical RSS feed reader app that makes you rawdog ten articles in an in-app browser before you hit a paywall, it would inject user-defined scripts into the in-app browser to make sure you never see the paywall in the first place. Philosophically, it's still a web browser, but with a more rigid interface for browsing between websites. The client is in control. It would be like Gopher for the modern age.
There are two separate web views that apps can use.
SFSafariViewController - content blockers will work with this. It’s an out of process web view that the host app has very little control over.
WKWebView gives the developers a lot more control over the web view and it runs in process. Content blockers don’t work in these views.
Sure, users mostly remember that they’re recording audio when shooting a video; but do they realise that advertisers have access to that as part of their photo library?
Even if you don't mind it accessing locations from individual photos, nothing prevents the app from scanning your entire library and getting all the locations from there too.
Same risk with timestamps - the app can get a list of all the timestamps and use that to confirm/refute other fuzzy datapoints collected elsewhere (web tracking, etc).
[0] by tagging the photo location or suggesting location specific things etc
“Haha I made that xD. Oh well 600k tc lol”
I've had too many "mentioned an idea, then got targeted ads for it within a day" incidents to not think some combination of things on the Google Home or iPhones or other smart devices is spying. Yeah, yeah, familiar with all the "someone else on the same wifi probably searched for it" or "Baader-Meinhof" or what have you, but I don't live with a lot of people and Instagram ads in particular are VERY focused these days in a way that Baader-Meinhof doesn't really fit - if I'd seen the ad for [specific thing] a couple days before I mentioned it, it would've been noticeably weird and out of place in a different way.
I was in ad-tech for a bit in the last decade, and even at a tiny company there were some data sources we bought or heard about that were pretty spooky, so I would not be shocked if there are some wild ones out there today.
Nobody has caught continuous audio feeds being transmitted from smart devices to the cloud (which would be noticeable due to increased network traffic and bandwidth usage) nor identified any secret speech recognition code on the client (which would be noticeable due to severely shortened battery life). Nobody who's worked in adtech has come forward to blow the whistle or admit that they shipped this feature for a big tech company.
I get why it's an appealing conspiracy from a gut instinct perspective, but it really makes no sense. When you're observing the behavior of billions of people and using machine learning algorithms trained to get the best results possible, some uncanny shit will naturally result. Look at how effective LLMs like ChatGPT have gotten without an obvious route to profitability, then think about how much more money has been invested into ad targeting algorithms just in the last couple decades alone.
Sneaky apps would be another source, obviously the phone OS/computer vendors wouldn't want this, but I imagine there's some cat and mouse. It's just a new version of browser toolbars, not something hard to imagine some unscrupulous 3rd party data collection company building.
I definitely wouldn't expect Facebook or Google to be doing it directly.
[1] - https://www.samsung.com/us/business/samsungads/resources/tv-...
I'm glad that you got out! I tend to be shy about calling things "evil", but I really do think the entire ad-tech industry qualifies.
Unless you assume that Facebook et al are running ASR models on-device all the time and somehow making it invisible to the end-user.
Same problem if they were surreptitiously streaming audio to their servers. You would see it from the outgoing packets and streaming that amount of data would also be fairly expensive.
Some voice codecs get their rate down to 2-3kbps. Maybe it could store and forward with content loads without raising flags?
Just spit balling here.
People are incredibly predictable using only a handful of demographics, there's simply no need to invest the astronomical amount that would be necessary to process these conversations when there's already many simple ways to track/generate user interest.
Interestingly enough, we find ourselves getting unusually excited when this happens, primarily because English is not our native language. When the spyware manages to comprehend our conversations despite our heavy accents, it gives us a sense of improvement in our language skills.
Conversely, our sentiments take a complete turn when we unintentionally activate Apple’s Siri while discussing unrelated topics. For some reason, the dedicated chip for detecting “Hey, Siri!” interprets one of us uttering that command and starts listening in. We exchange glances, shed a few tears of frustration over the misunderstanding, and then burst into laughter.
I’m not entirely sure how they manage to accomplish this, but perhaps there’s an idea for an app hidden in these experiences. For instance, an app that consistently listens in the background and occasionally shares a relevant joke based on the ongoing conversation. That would certainly add a humorous twist to things.