Mac Apps That Use Garbage Collection Must Move to ARC
developer.apple.com
developer.apple.com
Apple NEEDs to do this for their own sanity. I'm surprised they waited this long actually.
It's a serious PITA to support GC and non GC at the same time in a framework. Having to add finalizer methods instead of relaying on dealloc methods and dealing with non-deterministic object lifetime as well as deterministic was a huge problem.
ARC doesn't have these problems because it wraps traditional manual reference counting (to pick nits, it actually bypasses it when it can safely but this is an optimization detail that you shouldn't need to worry about).
To Apple and 3rd party framework developers, it's painful to support GC and non-GC paths in your code at the sametime. It's not really possible to use a non-GC compatible framework in a GC app.
It's also impossible to use ARC and still support the GC as well in the same framework. Apple has been completely prevented from using ARC in it's own frameworks because of keeping GC compatible apps working, and making this change will allow them to start ARC-ifying their own code.
It also means that a significant chunk of the ObjC runtime can be simplified.
There are also some great advancements in the ObjC runtime that Apple is doing (like tagged pointers and shadow objects) that the reference boehm garbage collector can't really deal with and was all completely disabled in the ObjC runtime when you had a GC app.
This is a good thing. Not many app actually shipped with the ObjC garbage collector and ran well.
Now you are free to use your GC for your own needs in your app. That has nothing to do with ObjC. Feel free to add Java or Ruby or whatever language you like to your own app and submit that to the Mac App Store.
Garbage collection is deprecated in OS X Mountain Lion v10.8, and will be removed in a future version of OS X. Automatic Reference Counting is the recommended replacement technology.
[1]: https://developer.apple.com/library/ios/releasenotes/Objecti...
As a programmer I don't want to litter my code with disgusting crap that only serves to tell the compiler something it should be able to infer anyway.
My suspection is that Apple is just doing it because they don't want to put the necessary RAM in their phones (Android has something like twice the amount of memory Apple has at this point).
ARC is compatible with manual reference counting entirely, but internally, when ARC detects the classes in a particular objects hierarchy are not doing anything fancy with MRC, it's super clever and bypasses it for a bit of a speed boost. It means that the developer doesn't have to do anything extra to be compatible with ARC from MRC or the other way around.
That would be like saying the only problem with the Titanic is that it had a few extra holes in it.
Of course what Apple really should do is dump its old tech and more to something new and more modern. I hate to say it but when even Java is more modern and can do stuff your system can't, it is time to upgrade (preferably to something other than Java).
They did. It's called Swift... and it uses ARC under the hood as well.
That hood is called Objective-C runtime and allows for interoperability with existing Objective-C libraries.
So from the technical point of view only ARC made sense.
You can, although I haven't seen systems that do this, enqueue objects to be deallocated then have either the main thread every once in a while or a background thread periodically pop objects off of the queue and do the normal reference counting dance.
Weak references take a while to figure out but once you have, refcounting is so much more predictable than GC.
The other problem with reference counting, which is I believe insoluble, is that it is inherently O(alldata); that is to say you must perform an operation each time a piece of memory dies, at the very least. This makes it good for collecting older generations in a generational GC (where O(alldata) ~= O(livedata)) but bad for the nursery (where the generational assumption means that O(livedata) << O(alldata)).
But this is always the case. Refcounting, GC, manual memory management, etc. But I'd say you've gone one step too far in saying that there is little or no temporal linkage with refcounting. If you're using a queue, objects will be cleaned up as soon as they can be while remaining inside the time limitations you've allowed.
If you want to get fancy, you can even have the compiler rearrange objects such that pointers that are likely to point to objects with destructors are first in line.
And yes, refcounting is O(alldata). But that's ignoring that freeing is relatively fast, and that you can easily implement bunches and bunches of optimizations to prevent refcounting from having to work at all.(For example: if you acquire a reference to an object on the same thread and then release it, you don't have to do anything. If you create an object, but at runtime you can determine that you never released the reference to anything, you can free it immediately. If you acquire multiple references to an object, you only need to check for references reaching zero once. Etc. All of these are things that theoretically can be done with a standard GC, but good luck actually doing so.)
It is usually presented as a quick solution for when the resources for doing more powerful GC algorithms aren't available, either in computing power or engineering.
Yes, with reference counting you must perform operations per allocation. But you already need to do this. Malloc / free, even arena-based solutions, all do this.
(Sorry for the question, I am not an iOS, Mac dev.)
All solutions are not optimal, but unless you have really many allocations and dealloctions going on, have tight resource constraints or hard timing constraints, current garbage collectors are really good enough. Not that you can not shoot your feet by accidentally keeping large object graphs alive, but that is easy to detect, though sometimes hard to fix.
Some developers seem to have some kind of resentments against garbage collectors because real developers of course manage their resources on their own, but many of their arguments are not well founded. Garbage collectors nondeterministically pausing program execution to free unused resources is a common point, but they usually don't realize that automatic reference counting has the same problem - you never know when a reference count drops to zero and triggers the deallocation of the object possibly with a large object graph attached.
Also, its not really fair to sweep "hard timing constraints" under the rug as a minority use case when these are essential for smooth modern GUIs.
Sure, GC is pretty much ideal in terms of absolving developers of resource management responsibilities, but it comes with some significant downsides which can easily effect the final product and should be understated. Rust is a great example of a new language which avoids GC for this reason.
If Java had a placement syntax like
new (HeapReference) SomeClass(); => allocate in custom Heap
Then you could just turn off GC for that heap, and then blow the whole thing away when it got full. Or, you could have a more fine grained API to allow GC on these custom heaps. Perhaps you could even allow copying objects back to the main heap. You'd have to ban cross heap references or have a smart API that lets users pin an object in the main heap while it is known to be referenced by an object in custom heap.
Most of the time in client applications, you only care about GC pauses on the main UI rendering thread, but GC pauses in other threads are less obvious because there's no jank, just that latency may go up for some operations.
You could have a kind of best of both worlds with languages that support non-GC allocation for your UI thread, but GC everywhere else.
The thing is that most discussions tend to be reduced to RC vs GC, as if all GCs or RCs where alike, whereas the reality is more complex than that.
RC, usually boils down to dumb RC, deferred RC or RC with couting elision via the compiler. Additionally it can have weak references or rely on a cycle collector.
Whereas GC, can be simple mark-and-sweep, conservative, incremental, generational, concurrent, parallel, real time, with phantom and weak references, constrainted pause times, coloured....
As for frame drops during animations, I think pauses in a missile control system is not something one wishes for:
http://www.militaryaerospace.com/articles/2009/03/thales-cho...
http://www.spacewar.com/reports/Lockheed_Martin_Selects_Aoni...
Heap compaction is another part of GC which results in long non-deterministic pauses. But the real problem is having fragmentation in the first place, this is something which manual allocation strategies can go a long way to mitigate.
For any sufficiently complex real world application you will not in general know when a deallocation occurs and that is one of the reasons you decided to use automatic reference counting in the first place. Some short lived objects within a function are non-issues to begin with and are easily handled manually if you want to, but the interesting bits and pieces are objects with unpredictable life time and usage pattern because they are heavily influenced by inputs.
I think you've confused determinism with regards to a program and the inputs to that program. Any retain/release program is deterministic because it always behaves in the same manner. Given any possible input I can tell you exactly when and on which lines of code the deallocations will occur. You can't know that with GC, especially when other libraries and threads enter the equation.
Which deallocations occur and in which order depend on the input to the program, but we still know in advance all possible lines of code where specific deallocations can occur and we have complete control over that. With GC we wouldn't know that and we couldn't control it.
Knowing when deallocations can occur and having control over that is the reason to choose ARC over GC in the first place. Not the other way around. With respect to manual retain/release, ARC doesn't take anything away: I can still retain an extra count of a large object to prevent its untimely release and share that object via ARC at the same time.
It's also worth noting that GCs rarely does anything because it feels like it, but because you are allocating past a certain limit. Knowing the size of objects, and given that the allocator can provide you with the current size of allocated objects, you should be able to determine if a GC is likely to incur in the next lines of code. This determination becomes harder when threads enter the picture, unless you're using a GC scheme like Erlang, which have one GC per process (lightweight thread). Some GCs also allow you to pause and resume them, for when you really want to avoid GC cycles, this can of course lead your program to crash if you allocate past your limit.
As always in computer science, 'it depends'.
After a while working with refcounting, I found myself writing code whose natural flow meant I was actually pretty sure, at least for deallocations large enough to care about. I don't even think about it any more, really, I just continue to enjoy my objects going away when I expect them to.
(Edit: Perl is garbage collected since version 6, and Visual Basic is garbage collected since version 7. I am not aware of any language besides Objective C moving in the opposite direction.)
It is hard to get right with existing code bases.
Additionally C++11 got a GC API, but requires explicit programmer control.
http://www.stroustrup.com/C++11FAQ.html#gc-abi
C++'s memory model is tainted by C's compatibility, as such it is quite limited in the GC algorithms that can be implemented.
How can you guarantee that the component library that your company just bought, available only in binary form isn't doing something like this?
void cool_stuff_with_widget(const std::shared_ptr<UI> ptr)
{
UI* evilPtr = const_cast<UI*>(ptr->get());
// now store evilPtr somewhere else and try to use it any time the library feels like it
}
No one is going through disassembly to check such behaviors.C++ RC is a very welcoming addition to the language, but it only works if everyone plays by the rules, 100% of the time.
It's just better documented and easier to spot the bugs if you're doing RC.
It is impossible to prove that all memory access to a given address are done via the RC wrapper object, specially in the presence of third party libraries in binary form.
Sadly due to its C compatibility, there is no way around this in C++.
This is mostly a problem in large projects with various skill levels across team members, having high attrition levels.
There is no Visual Basic 7. Visual Basic .NET is garbage collected as it runs on the CLR, but is quite a different language to VB6.
Similarly, VB vs VB.NET.
In the general case, it is estimated that GC uses on average twice as much memory (high water mark), but direct comparisons are of course very application specific. GC also has the dreaded collection pauses which can make performance unreliable. (Eg stuttering animations)
Biggest drawback of ARC is that it doesn't detect cycles, so the developer needs to have a global understanding of an app to avoid hard to debug leaks.
Objective-C is a superset of C. If you think about what it means to have a garbage-collected version of C for just a few minutes it should become immediately clear why abandoning GC for Objective-C is necessary and appropriate. The GC can't always reliably detect a reference (pointer) so it has to be conservative. It also can't compact the heap, which means you don't get the nice "allocation is just advancing a pointer" behavior that makes allocations nearly instantaneous in managed languages.
ARC is essentially just the old rules for manual reference counting (MRC), but the compiler analyzes your code and inserts the retain/release/autorelease calls automatically. This gives you similar deterministic behavior as MRC but while writing the code it feels like a garbage collected language because you mostly ignore management of memory and everything "just works".
The downside to ARC is the inability to detect retain cycles and a slight performance hit compared to MRC in some scenarios. The performance is actually pretty good... What ARC loses compared to hand-coding it gains by using special runtime functions (not slower ObjC message sends) that elide unnecessary retain/release calls in many situations.
objc_retainAutoreleaseReturnValue is inserted in a caller and callee (typically a property getter). It examines the return address on the stack to detect if the caller is going to invoke objc_retainAutoreleaseReturnValue on the returned value; if so it just does a retain and sets a flag the caller uses to skip doing retain+autorelease, leaving the caller's release only. Instead of the object staying alive until the next run loop iteration and getting autoreleased it can safely be released immediately, plus at least one extra pair of function calls can be eliminated.
Microsoft's C++/CX introduced in 2011 is also very similar; the hat syntax[1] is to COM's AddRef/Release as ARC is to -retain/-release. In both schemes, what used to be manually refcounted objects are now managed by the compiler. I don't recommend C++/CX though, it is confusing to C++ devs. Smart pointers are a better recognized idiom.
[1] Unrelated to the hat syntax from MS's previous attempts, Managed C++ and C++/CLI, which are GC and not refcounting as in C++/CX.
And my guess: GC probably doesn't play nice with recent/future changes in OSX memory management (e.g. compressed memory) so the less apps use it the better the whole system is overall.
Hence, they probably want to get rid of the whole thing ASAP, forbidding it _for apps submitted to the mac app store_ is a reasonable step in that direction.
There is a precedence for Apple doing this usually related to a new hardware platform e.g. x86, 64-bit and perhaps ARM Macs.
GC in theory is a good feature but it does not guarantee that constant time is used for each mark-and-sweep cycle. And together with multithreading that means uncertainty for rendering every frame of video, which is done in the main thread of an App 60 times per second. And for modern system, anything you see on screen is part of this rendering. Try to think about how could memory shared by main thread and a background thread be GCed. It could turn out to be: halting the process for GC, retaining some memory forever, or implementing some complicated algorithm that is too difficult to debug/maintain.
E.g.: http://michaelrbernste.in/2013/06/03/real-time-garbage-colle...
I've had to deploy lots of soft real time apps on GCed environments over the years, and it's always a problem. You can work around it with things like object pools, but some library or API will assume that the GC is OK and will be quietly spitting out objects continuously which will lead to a GC pause.
It's worth pointing out the Android devs finally started noticing this for Lollipop (probably due to their animations) and the API now has lots of places where it passes Java primitives instead of objects, which is the distinction between passing by value and by reference. Even if you're in C++ modern compilers can only make the most out of it if you pass by value, as this enables all sorts of other optimisations to kick in.
The key benefit of reference counting is it's predictable. Real time systems are also not strictly the lowest latency, they are defined by predictability. This becomes a preoccupation with minimising your worst case scenario.
This is simply not true. The API has always been heavily based on Java primitives. They didn't even use enums in the older APIs, preferring instead int constants. (I hate that one personally). GC pauses have always been a point of focus for the Android platform.
Notice the introduction of methods with API level 21 that recreate existing functionality without RectF objects being allocated. Touch events, for example, still spit ludicrous amounts of crap on to the heap.
All this is why the low latency audio is via the NDK as it's basically impossible to write Android apps which do not pause the managed threads at some point. Oddly this is stuff the J2ME people got right from day one.
Doesn't Android use a different flavor of Java anyways, allowing them to make these changes?
The problem Google has is that Dalvik wasn't written all that well originally. It had problems with deadlocking due to lock cycles and was sort of a mishmash of C and basic C++. But then again it was basically written by one guy under tight time pressure, so we can give them a break. ART was a from scratch rewrite that moved to AOT compilation with bits of JITC, amongst other things. But ART is quite new. So it doesn't have most of the advanced stuff that HotSpot got in the past 20 years.
If you look at the most advanced garbage collectors like G1 you can actually give them a pause time goal. They will do their best to never pause longer than that. If pauses are getting too long they increase memory usage to give more breathing room. If pauses are reliably less, they shrink the heap and give memory back to the OS.
Reference counting is not inherently predictable and can sometimes be less predictable than GC. The problem with refcounting is it can cause deallocation storms where a large object graph is suddenly released all at once because some root object was de-reffed. And then the code has to go through and recursively unref the entire object graph and call free() on it, right at that moment. If the user is dragging their finger at that time, tough cookies. GC on the other hand can take a hint from the OS that it'd be better to wait a moment before going in and cleaning up .... and it does.
It gets even worse when you consider that malloc/free are themselves not real time. Mallocs are allowed to spend arbitrary amounts of time doing bookkeeping, collapsing free regions etc and it can happen any time you allocate or free. With a modern GC, an allocation is almost always just a pointer increment (unless you've actually run out of memory).
The problem Apple has is that their entire toolchain is based on early 1990's era NeXT technology. That was great, 25 years ago. It's less great now. Objective-C manages to excel at neither safety nor performance and because it's basically just C with extra bits, it's hard to use any modern garbage collection techniques with it. For instance Boehm GC doesn't support incremental or generational collection on OS X and I'm unaware of any better conservative GC implementation.
Some years ago there was a paper showing how to integrate GC with kernel swap systems to avoid paging storms, which has been the traditional weakness of GC'd desktop apps. Unfortunately the Linux guys didn't do anything with it and neither has Apple. If you spend all day writing kernels "just use malloc" doesn't seem like bad advice.
If you're blocking your UI thread with deallocating a giant graph of doom then you have other problems. Deferring pauses, however, is not a realistic option.
Deferring pauses is quite realistic for many kinds of UI interaction and animation. If your animation is continuous/lasts a long time and requires lots of heap mutation then you need a good GC or careful object reuse, but then you can run into issues with malloc/free too. But lots of cases where you need something smooth don't fit that criteria.
Apple's GC implementation wasn't a Boehm GC [1].
It's true that it's hard to use a tracing GC with Objective-C, because of the C. But, if you want interoperability with C, you're kind of stuck.
[1] http://www.reddit.com/r/programming/comments/2wo18p/mac_apps...
This only happens if you choose to organize the data this way. This is a big difference from GC, where the whole memory layout and GC algorithm is out of your control.
Go doesn't restrict you to stack/heap allocation. You can create structs which embeds other structs. This simplifies the job the GC has to do, even if you don't allocate on the stack.
You can do something similar with Struct types in C#.
Over on the Reddit discussion there was a comment from Ridiculous Fish who at least was an Apple developer (and probably still is) and worked on adding GC to the Cocoa frameworks,
http://www.reddit.com/r/programming/comments/2wo18p/mac_apps...
Basically, because of interop with C, there's only so much you can do. Plus, the tracing GC wasn't on iOS so if you want unified frameworks (for those that make sense cross-platform), supporting the tracing GC along with ARC is added work.
Apple failed to produce a working GC that didn't crash all the time, because it required all Objective-C libraries to be compiled the same way.
Additionally the C subset only allows for conservative GC, which is not optimal when performance matters.
So in the end they went with ARC, with isn't nothing more than having the compiler produce the retain/release calls that Mac OS X frameworks were already expecting anyway.
So this has more to do with how technically feasible it is to have a proper GC in Objective-C, than GC in general.
Android developers tend to not think so much about lifetimes which leads to memory leaks because Dalvik apparently can't figure out some of those cycles by itself. Just do a search for "Android memory leak" to see some examples.
That Dalvik can't handle the cycles is stupid but an issue with Dalvik (that is likely to be phased out with the new runtime) not an issue with GC. Computers are really, really, good about executing trivial tasks repeatedly without ever making a stupid mistake, humans not so.
And as Terminator taught us: never send a human to do a machines job.
We did this for DBIx::Class in perl, thereby keeping full liveness but still getting timely destruction for connection objects. It was a trifle insane to get right, but it works extremely well.
The standard use case for weak references in Objective-C is a child-to-parent reference. The parent holds a strong reference to the child and children hold weak references to their parents, avoiding potential cycles.
In fact, part of good memory management in a reference counted system is that you should always have a hierarchy to your data structures so it never makes sense to have an actual strong reference loop.
Because that does not at all sound like a huge exploit.
> To aid in migrating existing applications, the ARC migration tool in Xcode 4.3 and later supports migration of garbage collected OS X applications to ARC.
EDIT: Or at least they had such tools 5-6 years ago, regression is certainly possible. For unrelated reasons I've been playing in other playgrounds for the last 5-6 years so I can't say for sure.
I've assumed anybody with a almost 10 year old computer doesn't get that much new software. I tend to release supporting the last 2 or 3 OS releases and 64-bit only and haven't received any complaints.
However, you are right, there is little demand for 32bit software; my newer apps require 64bit. But I see no reason to cut off support in an old product.
Had one or two complaints from existing 32 bit users but not much really.
It is nice to have a single complication target now.
I'm not sure if this was a typo or intentional but I like it.
This announcement is more about a feature in Objective-C being actively discouraged by refusing to distribute your app on their store. You can still use it if you distribute your app on your own.
This is somewhat misleading. Apple bans all use of JIT compilation on iOS, which de facto eliminates languages that depend on it.
[0]: https://code.google.com/p/chromium/issues/detail?id=423444
Finally Apple mentions the following in it's migration document [0]:
Is GC (Garbage Collection) deprecated on the Mac?
Garbage collection is deprecated in OS X Mountain Lion v10.8, and will
be removed in a future version of OS X. Automatic Reference Counting is
the recommended replacement technology. To aid in migrating existing
applications, the ARC migration tool in Xcode 4.3 and later supports
migration of garbage collected OS X applications to ARC.
Based on the above statement, my guess is that's Xcode has some build-in tools to convert GC code to ARC code, making the transition easier. ---
[0]: https://developer.apple.com/library/ios/releasenotes/Objecti...