They don’t even need to MITM the traffic. Just the fact that an app makes network requests when using a supposedly offline feature should immediately get them rejected.
And iOS should introduce a visible network activity indicator that can’t be manipulated by applications, like they do for location tracking.
I think this is already a thing in China.
> Further, the user should be able to control what domains/hosts the app has access to. The user should be able to have feedback indicating that the application is communicating over the network at what times and with how much data.
This is difficult to do, since it's easy to swamp the user in prompts for every little thing.
I don't see this as a bad thing. It will let me quickly see what apps are built not following best practices.
I would even surmise that a user's patience threshold for that kind of annoyance these days is pretty low.
In principle it's nice to be able to manually allow / reject individual requests, but in practice you can't get anything done if you do. Popups keep popping up all the time, and it's rarely clear what the request is good for. Then you get random failures, because you accidentally blocked an important request, or you just allow anything anyway, because there really is no way to know if that request is good or bad.
And that's from the perspective of a software developer who understands what protocols and ports and addresses are.
A firewall like that would be absolutely useless for non-technical users.
Basically I want Little Snitch built into iOS as a core feature. Hiding it behind “Advanced..” is fine.
A typical setup would work like this: when you launch an "instrumented" app, it generates a UUID. Then whenever a user interacts with the app, it would send messages like "UUID 1234... launched app version 3.14. UUID 1234 clicked the 'home' button. UUID 1234 viewed their profile page. UUID 1234 searched for a video. UUID 1234 played a video. UUID experienced an OutOfMemory exception in module foo, line 942." These were aggregated together so that you could run reports like "among people who experienced the OutOfMemory exception in module foo, line 942, how many viewed their profile page first?" That allowed developers to very quickly focus on the exact steps required to reproduce a specific problem.
So sure, apps were gathering a lot of information about what you were doing, but it really was genuinely for your benefit. There was no way for customers to run queries like "what video was heywire watching?" or the like. Everything was 100% focused on being able to quickly and accurately identify the cause of crashes. Now, that was just one company and it was several years ago. Maybe every other company was creepy? Maybe Apteligent is, too, now? I don't know. I don't have any insider knowledge into the current state of things. But at the time I personally witnessed it, I would have felt very comfortable at an EFF meeting explaining how every byte of metrics information was being used.
No, you've only talked about what currently happens when everything is working properly. What happens if the company ends up in financial trouble; to they have a Ulysses Contract[1][2] on record that binds their future ability to monetize all of this data? Without legal enforcement, we just have to hope this company will somehow resist the temptation that most other companies are not able to resist.
> what specifically about that data flow bothers you?
> it generates a UUID
That's obviously personally identifying, which it's a header in all of the analytics you describe. Just because it's synthetic doesn't make it anonymous. Once it's mapped back to other information - which is trivial if you correlate IPs[3] or event timestamps[4] - this type of analytics is only an INNER JOIN away from being merged into someone's pattern-of-life[5].
The problem isn't what happens when everything works as intended. You need to also prepare for when (not if) your data is merged into other databases, and what others might do with the data in a future.
[1] https://en.wikipedia.org/wiki/Ulysses_pact
[2] https://www.youtube.com/watch?v=zlN6wjeCJYk
[3] https://news.ycombinator.com/item?id=17170468
[4] Take a set of "UUID 1234... launched app" events for a common app that is regularly launched e.g. when someone wakes up (or whenever). Correlate those times to other times that also happen to be launched (or webpages/email visited) at similar times. What are the odds that two unrelated people just happened to open different apps [..., 2019-02-04T10:11:22, 2019-02-06T10:17:44, 2019-02-07T10:14:52, ...] (+/- maybe 30 seconds)? A unique identifier and a few high resolution (seconds) timestamps can easily identify someone uniquely when you have enough data points.
I can say that at the time I was there, it was not possible for a developer logged into our system to suss out any fine-grained information about a particular user. They just got aggregated data like "92% of people who experienced this symptom did this other thing right before it happened".
For my benefit without giving me a clear understanding of what was being collected or the option to opt out. Gee, you really shouldn't have.
Yes I know that you still have baseband + binary drivers, but at least then all of the apps are open source, and so is the OS.
Out of curiosity, what iOS features do you use? I ask because one thought you can do is have that as a phone and an iPad, so you can segregate your very private stuff and still use an iPad for iOS features
One thought for you, I just had my Nexus 5x die, but I got a Sony Xeperia XA2 for $150 and I flashed it with lineageos. See if you can use it and ween yourself off?
Alternatively setup Pi-Hole for testing.
Microsoft Edge (on Android) was the worst offender, contacting vortex.data.microsoft.com almost constantly, on almost every UI action I made. Other notable were the amount of apps contacting analytics servers in the background, when the app as (from an end user perspective) not even running.