The M1 has been out for a month and we already have hundreds/thousands of natively compiled binaries—and Windows and Linux are up and running.
Choice is not a problem and competition is good.
The M1 has been out for a month and we already have hundreds/thousands of natively compiled binaries—and Windows and Linux are up and running.
Choice is not a problem and competition is good.
Can you create a single build for all ARM chips for the same OS? If not, its all the same. At work we're using binaries built for XP that work without modification on W10. No doubt, its a testament to the massive backwards compat. effort by MS, but also from Intel. This is a massive, massive benefit to actual customers who just want to keep running software that they purchased/developed/commissioned on the replacement hardware when their current hardware fails. ARM/M1 brings a lot of benefits too, but as always, each person has to decide what they're willing to give up.
>The M1 has been out for a month and we already have hundreds/thousands of natively compiled binaries—and Windows and Linux are up and running.
Assuming the vendor is still in business, you have to rely on the good faith of the vendor to gift you the new version for free.
The OP raises very valid points, I don't know what is 'utter nonsense' about it. It's only tech, not a tribal war :)
I would add that mac osx does have universal binaries for some years.
The m1 _does_ have one weird non-ARM extension; it can adopt an x86 memory model on demand.
https://developer.arm.com/documentation/dui0801/g/A64-Floati...
Not Apple Silicon specific though
>Firestorm can do 4 FADDs and 4 FMULs per cycle with respectively 3 and 4 cycles latency. That’s quadruple the per-cycle throughput of Intel CPUs and previous AMD CPUs, and still double that of the recent Zen3, of course, still running at lower frequency. This might be one reason why Apples does so well in browser benchmarks (JavaScript numbers are floating-point doubles).
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
>Apple has confirmed that it’s a massive 192KB instruction cache. That’s absolutely enormous and is 3x larger than the competing Arm designs, and 6x larger than current x86 designs, which yet again might explain why Apple does extremely well in very high instruction pressure workloads, such as the popular JavaScript benchmarks.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
A major reason it does so well is the native Javascript support added in ARMv8.3-A.
Specifically, FJCVTZS (Floating-point Javascript Convert to Signed fixed-point, rounding toward Zero)[1] which is a Javascript-specific variant of FCVTZS implementing the overflow and exception handling behaviour that Javascript wants.
Javascript needs to do these conversions a lot, since it doesn't have integer types, and it's much faster to have it implemented in silicon rather than using the old instructions and having to handle the overflow/error checking.
[1] https://developer.arm.com/documentation/100076/0100/a64-inst...
https://mobile.twitter.com/saambarati/status/104920213252247...
I'm sure it doesn't hurt, but the performance advantage predates the instruction's use.
Here's another theory.
>Finally doing perf counters on a dedicated test bench... 9900K having 60% worse branch misprediction than Apple's A12.
Also apparently most JS engines do not actually use the FP unit but can do most computations on the integer units (unless is really FP math) which normally have much lower latency (but also lower bandwidth).
Another difference between x86 and ARM is that, for historical reasons, on x86 there is no need to invalidate the instruction cache explicitly when writing instructions to memory. That is, on x86, when a core writes to memory, the corresponding line in the instruction cache has to be invalidated. Since the instruction cache is usually VIPT for performance reasons, it has to be indexed only by bits which don't change in the virtual to physical mapping, otherwise there's a risk of cache aliases. For an instruction cache, an alias should not be a problem (it just wastes space with duplicated data), except that all aliases have to be flushed when invalidating by physical address.
IIRC, in 64-bit ARM user space (EL0) the only available instruction to invalidate the instruction cache is an "invalidate by virtual address" instruction. Since calling that instruction (after calling an instruction to flush the data cache to the point of unification) is required on ARM, there's no need to be able to invalidate all aliases of a physical address, like would be required on x86. That means it would be easier on ARM to have much larger instruction caches than on x86.
1. If you can increase the performance of a language by adding language specific instruction sets to the processor, how much of a boost would we see on average?
2. Instead of building in instructions for say Python, C#, Java, etc, what if you built it in a language designed to create other languages I.E Racket? (I like Racket but another language that specialized in this would be fine). Since languages built on Racket ultimately compile to valid Racket code, you'll still get the speed up, and can program with the features and syntax you want (within reason obv constrained by the limitations of the base language).
3. I wasn't around for this, but isn't this what made Lisp faster on Lisp Machines? Since they had dedicated hardware for interpreting Lisp instructions?
Hardware can parallelize type checks. For instance, you can have an add instruction which proceeds on the assumption that the two arguments are numbers. In parallel, a type checking unit in the hardware can abort that instruction and cause a branch to some handler if the operand types are wrong.
Function calls in dynamic languages can be expensive partly due to the dynamic checking that there are not too few or too many arguments. This affects code that doesn't otherwise require type checking (like logic that is doing nothing more than just passing arguments through several layers of functions and capturing return values). Hardware can help here also.
The industry moved toward general-purpose hardware. The Lisp people figured out ways to compile Lisp well to general-purpose hardware.
The thing about hardware is that it also needs optimization: a machine designed for ideal execution of Lisp is going to be expected to produce new revisions that are faster and faster, keeping up with advances in general-purpose machines. If it fails to do so, its performance will be overtaken by Lisp that is compiled to the general-purpose machines.
A PlayStation 4 is x86_64 but is NOT a PC. Watch the Fail0ver video on their PS4 kernel port to understand how different it is.
ARM is a god damn clusterfuck of random pins soldered to random shit and every Android kernel patched to hell in non-upstreamable ways. Look at all the work PostmarketOS has to do in order to get mainline Linux to work on all the random ARM e-waste out there.
DeviceTree is a joke. Microsoft at least forces ARM+UEFI, but they have locked bootloaders and even though people have found exploits to unlock them, there's virtually no reversed engineered drivers. All those hundreds of thousands of Lumina phones? Now they're e-waste. Worthless.
Linux grew because IBM created the PC. Compaq reversed engineered the BIOS and everyone started making compatibles. Over time, BIOS and later UEFI became solid standards.
ARM may be fast, but the non-standard Basic Input/Output and device detection makes it nothing but potential e-waste.