An example that might solidify the idea: pack a Wikipedia snapshot into it for search, and serve ~1m queries per second on it (12k qps per core).
The dataset is 2-3B records with 5-12 64-bit values each, stored in a few dozen files using the Apache Arrow format. If we take the midpoint of this range that's 170 GB just with raw data. With the overhead of data structures, I was running the process with ~400 GiB of RAM and could have done more on a beefier machine.
It took about 20-30 minutes to run the full algorithm on these tens of billions of data points and this approach was perfect for this use case. No overhead of Spark and all of its dependencies, just one program, a bunch of input files, and it's done when I get back from lunch.
If the vendors were ready for something like this on the software side, this would be great for edge compute when low latency response is required - remote utility substation handling and reacting to a large array of sensors feeding at 60 data points per second. In some use cases going to the control centre and back would be too slow to benefit. Basic grid control is well handled, but I could see optimizations benefiting from this. Vendors and utilities are way behind on this though.
2 Intel(R) Xeon(R) CPU E5-2690 @ 2.90GHz
768 GB of RAM (384 GB per Processor)
18 TB of SSD Storage (Intel 3 Series or Samsung 840 Pro Enterprise Series)
Naively, assuming identical instruction sets (I know they're not), 16 threads at 4 GHz is less than half as good as 80 at 2 GHz. But that can't be the whole story.
ARM only has 128-bit SIMD through NEON. Its reasonably well designed, but nothing beats the brute force of just doing 256-bits at a time (or 512-bits in the case of Intel's AVX512)
No. This core is terrible at encoding.
EDIT: And encoding is limited to ~16 cores in practice. It seems like after that, the communication between threads get too much to be useful. Unless you plan to be doing 5-simultaneous encodings at a time, then you're gonna have to find something else to do with all those cores.
For a decently sized video (say a TV episode) there's usually like 100 split points to divy out to encoders.