Why I’m not enabling Bitcode
medium.com
medium.com
I'm not clear on how Bitcode changes this, apart from expanding the already huge attack surface of the app store. Code execution inside app store servers is already game over for end-user privacy, isn't it?
The bit about how Bitcode makes backdoor injection harder is also very hand-wavy, since backdoor injection into fully compiled binaries is a teenaged pastime going back to the late 1980s. About the most you can say here is that Bitcode makes a trivial task trivial-er.
The concern about introducing side-channels into crypto apps seems legitimate (though: if you believe this is a real issue, and I do too, you're also saying "I will never trust Javascript cryptography"). I wouldn't want Signal submitting Bitcode.
If you're relying on compiler optimization levels to enforce stuff about crypto, you're gonna have a bad time. What happens when compiler behavior changes from minor revision to minor revision? Are you going to stare at the assembly code to make sure that something you did that was fair game for the optimizer to play with changed, or are you going to write your code so that it does the same thing no matter a (relatively correct) compiler and optimizer?
If your code had UB in it, you're already hosed and it's a matter of when. Fix that. There are primitives like RtlSecureZeroMemory to deal with the intersection of cryptography/security and compiler optimizations. We've mapped this terrain.
I guess the side channel crew can be appeased in the llvm bitcode view of the world - write your side channel resistant stuff in inline assembler. The bitcode optimizer can't touch it then. (edit-apple probably keeps you from doing this though)
Sounds like a good case for a unit test.
Inspecting generated assembly is a strange middle ground between too fragile for the paranoid and too much work for the not paranoid.
I don't think constant time C is feasible.
Then you're verifying the microarchitectural details of every CPU it executes on. (Intel's OpenSSL countermeasure to an RSA timing attack was dependent on CPUs not coming out with narrower cache lines than a predefined constant, but one did later.)
It's difficult to get high assurance in software. Best bet is to design algorithms to avoid data-dependent accesses (see: djb).
It massively expands the attack surface, since Apple is running untrusted IR input through LLVM, server-side.
> The bit about how Bitcode makes backdoor injection harder is also very hand-wavy, since backdoor injection into fully compiled binaries is a teenaged pastime going back to the late 1980s. About the most you can say here is that Bitcode makes a trivial task trivial-er.
It makes it WAY more trivial, expands the scope of what can be inserted by allowing essentially unlimited modification of the source without concern for linking/symbols/fixed offsets/section sizes, and allows for a degree of automation of complex changes across applications that would be difficult otherwise.
None of that is strictly earth-shattering, but add to this Apple's shipping a binary that the developer can't trivially verify by comparing against their local build, and you have what amounts to total opacity coupled with trivially easy application-wide code modifications.
Apple already runs your binary through a bunch of binary analysis tools, I'm not sure why LLVM is significantly different. A remote exploit is a remote exploit, and LLVM doesn't need escalated permissions to compile things. Besides, it's probably safe to assume Apple sandboxes all of this stuff.
> Apple's shipping a binary that the developer can't trivially verify by comparing against their local build
Who actually does that? I've never heard of anyone even trying to do this before. Besides, if you are that concerned about Apple tampering with your binary, it's a very small step to believing that Apple would still serve you the original binary and only send the tampered version to a limited audience, thus making any such verification meaningless.
Massive increase in attack surface.
> Who actually does that?
Anyone that debugs an issue in their shipped code.
Anyone doing security research.
Any users investigating their own system.
You don't know how large the existing attack surface is. I suspect that it's already large enough that adding LLVM to the mix isn't terribly meaningful. Especially once you assume that all this stuff is sandboxed anyway.
> Anyone that debugs an issue in their shipped code.
You're going to have to do better than that. Speaking as someone who's been developing iOS apps since the moment the app store opened, and who knows a lot of other iOS developers, I've never heard of anyone downloading an un-DRM'd version of their app store app for debugging purposes. The closest I've come to that is verifying what entitlements file the App Store version was built with (but this was done with iTunes Connect, not actually downloading the ipa).
The only benefit that you can get from downloading the binary is inspecting the assembly, which itself is only useful if you're hitting an optimizer bug, but even then, you'd actually want to work with a non-app-store build because you can't debug app store builds. So you won't actually be downloading the app store binary anyway, you'd just use the existing archive you uploaded to begin with, or even build a new one from the source.
Which means that, given all that, the only case here where you'd have a legitimate reason to want to download the app store version is if Apple manages to introduce an optimizer bug when recompiling your bitcode that you can't reproduce yourself. Not only is this likely to be extremely rare to get such a bug, you should still be able to reproduce it yourself if it happens because any new optimizations should be made available in the latest Xcode.
> Anyone doing security research.
Irrelevant to the topic. The case here was a developer verifying their own product. Bitcode changes absolutely nothing with regards to people downloading other developers' apps.
> Any users investigating their own system.
See previous paragraph.
That is because it has never been necessary before, because up until now you have always had a binary copy in Xcode -> Organizer -> Archives. Please think.
Also note that I am refraining from judging you or anyone you know on the telling account that neither you nor anyone you know have, for years, had the need to understand your binaries.
If your binary has a chance of functioning differently on someone else's device than on your testing devices, then you should care. But then again, your business may not depend on having an app that works, but on producing a lot of apps.
If you are religious in trusting Apple's decisions despite all your years as an Apple developer, then you should at least acknowledge that some developers and their companies may have different (less faith-based) values than yourself.
[1] https://cryptocoding.net/index.php/Coding_rules#Compare_secr...
I guess with LTO the result could be the same, but asking to always inline seems like asking the compiler to fuse operations, and then it sees shortcuts. Dunno.
The trick here is that -march=i386 predates CMOV, and that LLVM specializes code emitting for bool (and _Bool). If the secret bit were an uint32_t, there wouldn't be a branch anymore.
There is one class of cryptographic code, however, that is entirely unsuitable to distribute in Bitcode---DPA/EM-protected code. EM attacks on middle-end ARM chips have been demonstrated recently [1, 2].
Protecting against these attacks usually involves splitting the computation into 2 or more "shares" (see, for example, [3]); these require strict control of which register each word goes into, and which registers overwrite which. This cannot be enforced in Bitcode---or any other bytecode, for that matter---and direct assembly must be used.
[1] https://eprint.iacr.org/2015/561
[2] http://cr.yp.to/talks/2014.09.25-2/slides-dan+tanja-20140925...
I think the point is that it would be harder to detect if this happened, as, with Bitcode, the author loses the ability to compare binaries downloaded from the App Store to those produced by Xcode on their machine.
> FairPlay encryption already made it really difficult for developers to make sure that Apple distributes the build you think they are distributing. But at least, with jailbroken iPhone, I’m currently able to get some idea of what the binary I’m distributing looks like and compare it to the one I submitted to Apple.
But even stipulating that you're right, this still doesn't make sense. The argument here is that someone with control over Apple's code distribution infrastructure could not, in the absence of an all-bitcode submission system, surreptitiously backdoor iOS apps. But of course, there's a million ways Apple (or someone who owns up Apple) can backdoor apps in ways that don't alter the app binaries themselves directly at all.
I actually think it would be better if Xcode did both encryption and code-signing at the developer side, then Apple just added another signature on top of the full binary once it was approved for App Store distribution. This way, developers could easily verify that their signature was intact and that nothing had been modified in between.
Why is this a legitimate concern? I've never heard this even suggested as a remote possibility of a concern before. If you're already so paranoid as to suspect Apple of inserting malicious code into your application, it's a very tiny step to believing that they'd still serve you the original app and only serve the modified copy to certain targets, thus making any binary verification you can do meaningless.
> I actually think it would be better if Xcode did both encryption and code-signing at the developer side
Apple doesn't serve the same encrypted binary to everyone. Each user has an independent set of FairPlay keys and Apple serves up content encrypted for that particular user. That's kind of the whole point of FairPlay. If Apple served the same encrypted blob to everyone, then the encryption would be completely useless as everyone could decrypt it.
I used to think this as well. However, (my memory is fuzzy on this so take it with a grain of salt) a year ago I downloaded the same app from two different iTunes accounts. Both encrypted apps were identical, and had the same encrypted blob. So it seems Apple is not distributing a unique encrypted blob to each user, at least as far as apps are concerned.
It's not that I don't trust Apple, it's that there's no reason to trust them in this case. Either stop forcing developers to deal with code signing for App Store apps or stop destroying their signature, allowing verification and logging of builds.
And yes, the same encrypted binary IPA is served to everyone, but it gets customized by the device/iTunes.
iOS devices are already running a closed source proprietary OS, which can easily set an hardware breakpoint in a binary and then do anything that could be done by modifying the binary itself without leaving traces anywhere (other than timing and cache effects).
Apple even designs and manufactures the CPU so they could even modify the CPU to detect instruction patterns and siphon off private keys without any detectable timing or cache effects at all.
Of course, that doesn't mean that it's all fine, but rather that the system is already completely broken to begin with in this respect.
As much as I'd like to be upset about the lack of binary verifiability throughout the chain, as the article even points out, without jailbreaking it is already impossible. We have no idea what code Apple is telling our devices to run, regardless of the form in which it is submitted to the App Store.
It's a bummer, but rms was right all along.
It's going to be really really really Fun with a capital PH when Apple enables some new optimization that contains a subtle bug, and recompiles everybody's apps, and us third parties are left holding the bag trying to figure out why our apps are crashing with no visibility whatsoever into the process.
I don't quite get why Xcode couldn't generate both the LLVM bitcode + the fat binary for all architectures, along with any flags. Then Apple could regenerate the same binary (to verify the two match) but use the bitcode for their program analysis tools. Then, the developer's signature could be left intact on the binary and Apple would add an additional signature over the whole thing that indicated it was approved for distribution via the App Store.
If I created an app, both distributed internally (say, to my corporate users) and via the App Store, but the App Store version has a bug because they're using a different ARM backend than my Xcode, that will be a real pain to debug.
One thing that is different is that there's no transparency into the process. If your Java code is crashing, you can look at the code the JIT is generating. If your Java code is only crashing on certain hardware or OS revisions or whatever, you can see what's different about the JIT in those circumstances and at least make an attempt to investigate. With Apple's system, you have no idea what compiler version they're running, when they make changes, what machine code your users end up running, etc.
There's no guarantee an ARMv99 will run even identical code at a rate directly proportional to an ARMv7.
Bitcode does add potentially new and exciting failure modes, though, since you might start failing even in the absence of new hardware, and you might encounter problems other than timing (e.g. failing to zero out sensitive data because dead store elimination got smarter, or buggier).
(1) Apple has the signing master keys for app store certificates. It would be easy for them to binary-patch your app with anything they want, generate their own signing certificate in your name, and re-sign the app as they push it out to users. Since app store transactions are on a per-customer basis this could be done only for some customers -- such as only residents of a certain country, people whose names are on a list, etc. It's very unlikely that anyone would notice.
(2) It would also be easy for an attacker to do this if they compromised Apple's core infrastructure as described in this article. Then they could generate signing certs in other developers' names and trojanize apps in exactly the same way. A "doomsday hack" that trojanized all Apple devices in the world within hours is not impossible. If they all started, say, DDOSing core BGP and DNS systems at the same time, such an attack could take down the entire Internet.
(3) Apple, since they control the platform and can remotely patch it at any time, needn't bother with trojanizing apps. They could simply trojanize the OS kernel or user-space OS components and accomplish anything a trojanized app accomplishes.
The same also applies, by the way, to Windows and also to Linux if you trust remote code repositories and do not compile everything yourself from verified known-safe source code. It also applies to from-source-compiled OSes and distributions if you did not painstakingly check every line of code yourself. Google the "underhanded C contest."
If you run code that is not yours, you are trusting someone else. It's a subset of a larger principle: civilization implies trust. Without trust you cannot have specialization of labor, commerce, or law. Without those there could be no advanced technology or infrastructure, so none of this would matter.
To use an Apple device is to trust Apple, at least insofar as that device and any data you keep on it goes. To use Ubuntu and its repositories implies the same level of trust in Canonical, etc.
The privacy problems of the digital age are not entirely technical in nature and do not have entirely technical solutions. They can only really be solved legislatively and socially.
Curiously, Google/Android has developers sign the app, and that exact signed package is what's distributed, so at least the app developer can verify that the distributed app is identical to the one shipped to the store.
On the other hand, Amazon/Android signs apps with their own keys, so they suffer from the same issue as Apple.
So even if apps are signed by the developer, the app store owner can throw away the signature, modify the app and then sign it with their own key. People downloading and running the app won't see a difference.
But more to the point, I published an app on the Play Store that actually verified its own signatures (as an anti-piracy measure) and it worked correctly. Not only that, but it used the signature as a key to decrypt some of the assets, so a changed signature would mean the app would completely fail to work.
So it's a solvable problem on Android. But not at all on other ecosystems.
Not true. Not using bitcode allows you to select optimization options and have a "receipt" on what CPU:s your code will run, by providing those slices. Now, Apple can also (in theory at least) change endianness of the target code. Instead of users being unable to install the app (a bad idea), they will install an app that breaks (an even worse idea). Endianness is one example. Optimizations another. There can be more. Hiding even MORE stuff from a developer who in turn is responsible to a customer and to users can only raise the development cost (because maintenance contract is usually a part of development cost), not lower it. I do not need MORE hidden issues and LESS control over the work I put my face on.
LLVM bitcode also limits the kind of tamper resistance and obfuscation you can put into your apps. Fairplay offers almost no security if your attacker has a jailbroken device. He can simply dump the decrypted app from memory and reverse enegineer it. If you care about protecting your secrets, or preventing your app from being pirated and redistributed on Chinese app stores, you could at least apply fairly sophisticated obfuscation & integrity checks to the app binary. With LLVM bitcode that all becomes much harder, and applying integrity checks to the app becomes impossible.
Also, Apple doesn't really need the LLVM bitcode of the app to carry out app thinning. All they really need to do is extract the proper architecture from a FAT binary, and send that to the user's device, instead of the entire FAT binary.
It makes me wonder what will they even be using the LLVM bitcode for? Will it become a requirement for all apps submitted to the app store sometime in the future? If so, why?
LLVM bitcode is much easier to reverse engineer than machine code, and you will be handing Apple all your symbols and metadata as well.
With LLVM bitcode, it doesn't matter if you don't hand over your .dSYM file to Apple, all the symbol information (and more!) is available in the bitcode.
It's 2015. Lack of source code is no longer a meaningful obstacle for attackers.
Also, if Apple requires apps to be submitted as LLVM bitcode, then certain things like integrity checks and certain useful types of obfuscation will become impossible, making apps much easier to reverse engineer, or to pirate and repackage on another app store.
However, I'm talking about applications for who the source code is not available. If you properly use integrity checks, signature verification, and obfuscation, you can make repackaging and piracy of applications significantly more difficult. With LLVM bitcode this is no longer the case.
Are you concerned about Apple reversing your code? You can still use arbitrary types (e.g. int64 instead of a pointer) and give them no additional data about the structure of your code. As you pointed out elsewhere, tools like OLLVM perform obfuscation at the bitcode level already.
I'm guessing Apple made this change to make it easier to do program analysis of iOS apps being published. You can certainly find bugs easier in bitcode that has proper type information and clear differentiation between safe and unsafe branches. One of the first steps in program analysis tools is to lift native code to an IR in order to determine its correctness.
With bitcode distribution, Apple has potentially made it easier on themselves to skip the lifting and type recovery steps for conforming apps in order to look for bugs. But unless they start requiring bitcode conform to certain additional standards, you can always transform bitcode to obfuscated bitcode, destroy type information via aliasing, etc.
I agree also that Apple probably made this change to better analyze programs being submitted to the app store. That and to recompile programs to use intrinsics more efficiently, as new intrinsics become available.
And finally, while I would not be concerned about Apple reversing my code, certain companies are and have to undertake steps to make that as difficult as possible.
If you are concerned about anyone reversing your code (Apple or otherwise), you must obfuscate it. Either you work at the ARM level or the bitcode level (watchOS for now), but the basic techniques are the same. Or, you can avoid changing tools by doing source-level transformation.
One of the companies probably impacted by this change is Arxan or other obfuscators. They have to change both front and backend to be LLVM -> LLVM.
[1] Except for the case I mentioned before, where you predict the generated code for known architectures and use that to generate your constants. This still breaks if Apple generates code for a new architecture without your help.
It's also possible for you to remove the ObjectiveC metadata from a binary entirely, obfuscate/encrypt it, and then add it back to the ObjectiveC runtime as needed. With LLVM bitcode this becomes much more difficult (but maybe not impossible...)
Meanwhile: LLVM's ARM backend is not an obfuscating compiler backend.
If you are obfuscating your code, yes, Bitcode is a problem. But Signal is not obfuscated.
This entire time I was referring to code obfuscation/integrity checks, and why bitcode is a problem for people who want to obfuscate their code. I was not referring to apps like Signal at all.
As an interesting aside, LLVM's ARM backend can be made into an obfuscating compiler, and furthermore some people have started obfuscating the LLVM bitcode itself: https://github.com/obfuscator-llvm/obfuscator/wiki http://0vercl0k.tuxfamily.org/bl0g/?p=260
Actually, I should raise the 'NateLawson alarm; he can probably tell us precisely how many mobile developers obfuscate on either platform.
First, throw out Proguard on Android. That's not really obfuscation, it's an optimizer. Sure it renames variables and removes dead code, but that's only a problem if you have a naive system that relies on class & method names. Proguard is used in 20% of Android apps though.
The most common issue with Android is the malleability of its bytecode it inherited from Java. It's common to patch apps or repackage them for other app stores (after swapping advertising API keys). There's no work required to rebase addresses or anything after patching, unlike native code. This drives the need for obfuscation, but also provides a useful springboard for implementing it.
We've found that many of the obfuscation schemes on Android are custom. This is because of the liberal policy towards using native code via JNI and the ease of inserting a custom Java classloader. With iOS's code signing and memory protection, you can't write a custom loader in your app that decrypts and executes new code.
However, every other obfuscation measure is available in iOS. You can build opaque predicates, mangle control-flow, do arbitrary pointer arithmetic, and lots of things that are impossible in Dalvik. (However, you can just write a thin Java wrapper around your .so in Android and you've got the same capabilities and more.)
With LLVM bitcode distribution, you're really potentially losing only one thing in your obfuscation toolbox: self-checksumming. You couldn't due self-modifying code already due to code signing, so nothing changed there.
While I haven't implemented it myself, I believe you could still even do self-checksumming via clever use of intrinsics. That is, you carefully lay out the instructions and data and predict what you'll expect to see when it's translated to armv7, v7s, arm64, etc. That's the value you'll check for.
Self-checksumming via intrinsics and clever layouts is a very interesting idea, but it's impossible for you to checksum a program when you're not 100% sure what it will look like after compilation. For example, even if you manage to correctly 'guess' what machine code your LLVM bitcode would be translated to and adjusted your checksums accordingly, Apple could just update their LLVM bitcode compiler at some point in the future, and invalidate your checksums.
Also Android is a much more interesting platform than iOS, because it allows for self-modifying code, which various commercial obfuscators use in various interesting ways :) Also there are very interesting things you can do with reflection or the JNI in Java.
Regarding Android's flexibility, I've written an obfuscator (on another Java-based platform) that took a DSL and generated half the code in Java and half in C, intertwining the computation between the two processors via IPC.
That being said, there is also a very sizable minority of companies that do obfuscate their mobile applications.
And no, the transition to 64 bit or Swift are not that.
If you meant, Apple imposing newer stuff to developers, then yes Apple does that.
Though, I'd say most of that is for the platform's good. Even stuff like deprecating Java -- it never got popular for OS X apps as Apple intented, the situation with SUN had changed, and keeping it going forward would be dead weight.
And Metrowerks might be good for early '00s, but not going forward with iOS, and multiple languages, and Storyboards, and all those things.
And I find it disheartening that most developers seem so accepting of unnecessary costs to their time and health. Much like the worker class before the unions - any work conditions go. I for one am pretty tired of years of iOS-update introduced incidents and overtime work to manage new products, while juggling undue maintenance on old products. I also find it disheartening developer companies don't take a good look at costs and demand improvements, instead of accepting a steady trickle of punishment for success.
If it is for the platform's good, I wish the platform ended so we can develop for platforms that deliver better working conditions and less motivated uncertainty, fear, and doubt.
A good argument against the bitcode is the huge amount of open source libraries in C used by thousands of popular apps (such as openssl), which may not, or do not, particularly enjoy the bitcode treatment at AppStore.