iOS Apps Using Private APIs
sourcedna.com
sourcedna.com
We take a different approach to understanding code than the traditional antivirus world. Rather than try to hunt for a needle in a haystack, we've created a system for finding anomalies in code that's already published. For example, you can build a set of signatures for "bad apps" and then repeatedly search for them (AV model) or you can profile what makes an app "good" and then look for clusters of apps that deviate from it (SourceDNA).
Consider an ad SDK like Youmi here. They weren't always scraping this private data from your phone. There are some apps that have this library but that version is a typical, only sorta intrusive, ad network.
But, over time, they began adding in these private API calls and obfuscating them. This change sticks out when you track the history of this code and compare to other libraries. There was more and more usage of dlopen/dlsym with string prep functions beforehand. This is quite different from other libraries, where they stick to more common syscalls.
By looking for anomalies, we can be alerted to new trends, whatever the underlying cause. Then we dig into the code to try to figure out what it means, which is still often the hardest part. Still, being able to test our ideas against this huge collection of indexed apps makes it much easier to figure out what's really going on.
The iRiS paper we mentioned in the blog post describes a really great approach to doing this. They do "forced execution" using a port of Valgrind to iOS. They also do the exact right thing by resolving as many call targets as possible statically, then using dynamic analysis only for the call sites that can't be resolved. This saves on runtime and complexity, though you might notice that even this approach didn't resolve 100% of them.
Ultimately, you're dealing with a variant of the halting problem, where the app uses a specific value only on a full moon, on iOS 4.1, where the username is "sjobs". And that's why computers are still fun.
PDF: http://www.cse.buffalo.edu/~mohaisen/classes/fall2015/cse709...
Apple can defend against unauthorized calls to even runtime-composed method names though. I can think of a few ways.
They could move as much "private" functionality as possible outside of Objective-C objects entirely, which requires that you know the C function name and makes it obvious when you've linked to it. This should probably be done for at least the really big things like obtaining a device ID or listing processes.
Even if they stick with Objective-C, they could have an obfuscation process internal to Apple that generates names for private methods. Their own developers could use something stable and sane to refer to the methods but each minor iOS update could aggressively change the names. If the methods are regularly breaking with each release and they're much harder to find in the first place, that may be a sufficient deterrent to other developers.
They could make it so that the methods are not even callable outside of certain framework binaries, or they could examine the call stack to require certain parent APIs. At least that way, if you want to call a private API, you have to somehow trick a public API into doing it for you.
And, I think Apple does say somewhere that developers shouldn't use leading underscores for their own APIs. They could hack NSSelectorFromString(), etc. to refuse to return selectors that match certain Apple-reserved patterns in all circumstances.
I am not sure if you have to explicitly link against certain frameworks/dylibs though, someone with more knowledge feel free to correct me
I've always wondered why Apple doesn't run all apps against a "debug" build of iOS that asserts that the caller of private APIs is itself internal/private to Apple, but instead relies on something akin to grepping the output of strings(1)
[1] https://developer.apple.com/library/ios/documentation/Genera...
There's even a paper on intentionally inserting security flaws into your code and then exploiting them from your own server to change execution patterns:
https://www.usenix.org/conference/usenixsecurity13/technical...
Ultimately, you need to enforce access control instead of just trying to detect problems a priori. Apple's sandbox is a great start to that, and I expect they'll keep improving it to block apps like these.
We scan code on both iOS and Android, identifying libraries and patterns of behavior. OpenSSL, for example, is the same code, regardless of platform, but the versions of it may vary depending on packaging and porting. We found one guy who had compiled OpenSSL for Android and was distributing binaries from github. Not the safest way to get your software.
As a third-party, you can move quickly and correlate info across app stores. But I also think we've developed a pretty unique code analysis engine that neither Apple nor Google has. :-)
So while Apple is motivated to catch this stuff to an extent, they aren't that motivated, and they may well be perfectly happy with the current level of sophistication.
There's still no clear reason why Apple choose to enforce the security via app reviews, rather than the system call returning 'permission denied' to apps without the right security level.
But it is not a security thing at all.
Apple is probably in the unique position to actually implement such a system successfully, controlling the hardware, os, and even language choice.
Or they could continue with the current sandboxing system where security- and privacy-critical functionality is performed out of process, and plug the remaining leaks, of which there aren't that many.
This way your code and most legit code that is not trying thousands of URLs works, but apps are trying to do this fail but don't crash.
open() would be a much simpler syscall if it didn't bother to enforce access controls. The code would be a lot easier to maintain too...
e.g. the advertising ID could be stored in a file with appropriate permissions. All the API needs to do is open the file and read it out.
I think we're talking at cross-purposes. My original question was why Apple don't restrict this restricted information. It appears we both agree that putting it behind an access control (like a syscall) would prevent this.
Because Apple wants to use this information, but they don't want anybody else to use this information.
If Apple requires security permission to get at the information, then they hose themselves as well.
I believe the current sandbox design makes it more complex to promote system interfaces to require individual access control. It doesn't provide the fine-grained access, so more work is needed to separate code with the granularity that sandbox policies can control it.
This does happen over time (e.g. UDID in iOS 8).
If there are any security implications to making a private API call then this implies that Apple's sandboxing system is broken. If so then the solution is to fix the sandbox. Better detection of private API usage might be good for them, but is neither necessary nor sufficient nor particularly useful for security.
http://stackoverflow.com/a/27686125
The reason they don't just do this for every API is due to the sandbox design, which makes this cumbersome.
The trend took big steps in iOS 6, when Apple created a mechanism for remote views, and moved things like sending SMS and email out of process, in addition to the mechanism for determining what process gets what touches (backboardd).
As you mentioned, MobileGestalt was another step in this direction, when Apple blocked access to the MAC address and UDID.
iOS was not originally engineered with this focus of security and sandboxing in mind. For example, CVE-2015-5880 (accessing contents of screen from anywhere prior to iOS 9) existed because the screen framebuffer was needed for QuartzCore to function, and Apple didn't take the time to re-engineer how things work.
The number of things you can still do from the sandbox is mind boggling, and Apple is aware of it, they just don't have the time to re-engineer everything.
toucharcade.com/2015/09/23/bioshock-removal-2k-support-response/
It has, and very early on Apple didn't check for their usage so developers started using class-dump results to look for interesting private/undocumented APIs.
Things went south when Apple started checking and rejecting applications which did that: https://www.cocoanetics.com/2009/11/forbidden-fruit-apple-ap...
The youmi stuff is creepy, but generally speaking it's more a question of there being no guarantee about these API's stability or continuity, so the application may break between even minor OS updates, which is undesirable.
More ambitious programmers can scan executable regions of memory for the instruction stream of the target function and call the address directly without any help from symbolic names at all.
The runtime isn't as straightforward as SourceDNA makes it seem.
A statically-linked function would have to be referenced directly in the executable somewhere.
Some of the things developed during the arms race have become best practices, though. If you are an iOS game developer and you release a game update without the ability to adjust the game substantially from the server side, you are really in trouble. It takes another long review period from Apple to change anything, so best practice is to control the functionality from the server where if it makes the game too difficult for the users or something it can be adjusted without suffering the Apple review delay.
It has happened years ago. About a year ago I left a company that was doing this for at least the last two or three years. Their motivations were strictly a technical workaround to get functionality that Apple didn't present to a public API, not to nefariously gather user info, but the technique is the same, and not particularly difficult to figure out.
Probably a good five years ago Apple rejected an app of mine for use of a private API (this was about the time their static analysis was implemented). Another company that sold a buttload of copies of their app had to have used the same API, and their app came out a little bit before mine. Could be Apple just missed it, but I suspect the company in question was doing something similar. (They made millions, I languished in obscurity, but I'm not bitter. <g>)
And those are just the ones I directly know about, or strongly suspect. I'm sure there are/were plenty more.
https://stackoverflow.com/questions/6530701/is-the-function-...
Like what ASLR does for memory addresses, but for selectors.
Conversely, it would be nice if the user could grant access to email and iMessage sandboxes. This would allow us to apply machine learning to personalize services. Ironically, by allowing opt-in, Android is an easier platform for creating private services.
Normally, these private API selectors would show up in a class dump and Apple will reject you app. But, if you're clever, you can hide them from a class dump. You could encrypt the strings, then decrypt them at runtime and Apple could no longer find your private API usage in a static scan.
At runtime they could detect you calling private APIs, but it would be easy enough to code it so that you don't call any private APIs for a few days after first launch or make them so they're turned on with a server side flag. That way Apple would never notice the private api usage during an App Store review.
We're moving ahead with detecting this too as it's an obvious next step.
> We believe the developers of these apps aren’t aware of this since the SDK is delivered in binary form, obfuscated, and user info is uploaded to Youmi’s server, not the app’s.
Know your binaries?
I think you need a appstore-wide view of the entire software world and the ability to query it for arbitrary behavior. But I'm biased since that's what we've spent years building. :-)
Because as far as I, a user is concerned, that developer put that software on my phone.
Maybe I as a user have a similar due diligence burden. One thing I might do is never download an app from the affected developers again. But that doesn't seem like a desired outcome. Nevertheless, I don't see how a user could do anything more fine-grained and nuanced than that.
I'm biased, but I think developers should be using our service to track the code that they're putting in apps (including their own code) for security, quality, and app review problems. We watch third-party code for them, which is how we found this issue. I think it's unreasonable to expect developers to reverse-engineer every library they include.
I agree users really can't do as much about this. We are considering ways to distribute the list of vulnerable apps/versions to help users find if they have them.
It's not an easy problem and not something you can do by hand. We've indexed the code in millions of apps and then learn patterns from it. For example, you can discover a new SDK being adopted by apps in one region purely from the co-occurrence of binary signatures.
We're working hard to give developers this kind of insight, helping them detect problems in their apps before they are affected.
(Fair warning: I am not just a disinterested observer of SourceDNA).