M1 Memory and Performance
blog.metaobject.com
blog.metaobject.com
Is it as fast as having more DRAM in all possible scenarios? No, but with good memory management the real world performance might very well be identical to a system with the same capabilities and even greater than traditional memory systems with larger non-unified pools.
https://developer.nvidia.com/blog/gpudirect-storage/
It’s the same fallacy as RAM size impacts video editing, it doesn’t.
8K RED raw is 4.374 terabytes per hour... at that point it doesn’t matter if you have 8, 16 or even 256GB of RAM there if your storage is too slow to support it you’ll drop frames and or have huge seek times, more memory won’t really help you there is no amount of prefetch you can do to bridge gaps that big.
There is a huge advantage of having a unified memory architecture where the application can address a single memory address space to access your data and you do not need to perform memory copies, allocations and all that stuff between your storage, CPU and GPU...
Computers are used for much more then just Video Editing...for Fluid Simulation Ram definitely has a impact. And Talking about SSD (QLC)...if they run out of SLC-Cache they are sometimes slower then HDD's. And i bet your Video editing is faster if you have 4TB of RAM.
It also doesn’t help you at all for random access.
Is it as fast as having more DRAM in any scenario?
Essentially all the big compute vendors moving towards memory unification and cache coherent interfaces.
We have cheap consumer SSDs today offering 5GB/s of read performance and more than enough I/O to satisfy caching.
While I haven't seen benchmarks yet, the claim was 2x relative to their current systems, IIRC, which were already crazy fast at > 2GB/s.
So it would probably do OK even just plain swapping, but really, really well moving these large blobs in and out, for example using mmap() and madvise().
In theory this can also be extended over TB/USB4 and ofc it can be extended over the network (however that’s the slowest part) this going to be quite a big thing especially for use cases like video editing if they offer a true unified memory as a lot of the latency doesn’t necessarily comes from the interface latency but from memcopy and translating between multiple address spaces.
Also the big culprit in shortening the life of an SSD are writes, unified memory doesn’t mean that you have to write more often to the SSD in fact it means the opposite, unified memory isn’t a page file.
https://en.wikipedia.org/wiki/Shared_graphics_memory
I'm curious if the same drawback mentioned for shared applies to unified
> A side effect of this is that when some RAM is allocated for graphics, it becomes effectively unavailable for anything else, so an example computer with 512 MiB RAM set up with 64 MiB graphics RAM will appear to the operating system and user to only have 448 MiB RAM installed.
Yes. Previously shared was used in the context and goal of cost savings. But* Unified* now has a goal of higher performance.
Apple's talk about RC over GC is just marketing speak for the failure to have a tracing GC in Objective-C that wouldn't crash left and right when integrated with C like code.
It was a very sound decision given Cocoa semantics, and the difficulty to make anything written in C not to fall apart with segfaults, but lets not oversell it.
Likewise Swift RC makes sense from having to integrate with Objective-C runtime and existing ecosystem, but again that is all about it.
There is no RC implementation with comparable performance to tracing GC languages that isn't just yet another tracing GC from the amount of runtime support needed to make it actually fast.
Swift is a disaster, also agreed. So?
For the rest: actual research disagrees with your forceful but unsubstantiated assertion:
https://2013.splashcon.org/details/oopsla-2013-papers/21/Tak...
And once again, tracing GCs do well in microbenchmarks where you only check the cost of local operations. They are horrible when you take the global effects into account, with those local/benchmarking advantages not translating into real world use.
https://people.cs.umass.edu/~emery/pubs/gcvsmalloc.pdf
Very similar to JITs, which also do massively better in microbenchmarks than in production code:
http://blog.metaobject.com/2015/10/jitterdammerung.html
Oh, and integrating well and without high cost to fast languages where you have even more control is actually an important feature.