Big-Endian “is effectively dead”
thread.gmane.org
thread.gmane.org
What is really alarming to me is that I occasionally run into middle-endian systems on 64-bit chips (two little-endian doubles in big-endian relative order, to signify a single quad). This is an abomination and must be killed with fire.
So the big endian gene lives on in the JVM and its relatives.
... except char.
You also forgot to mention their odd utf-8-ish encoding of strings in class files.
Java 8 has unsigned types[1]
> boolean that cannot be converted into integer types
Don't know why you would want to do that when a boolean is a primitive type... and also has an object variant, but...
int myInt = (myBoolean) ? 1 : 0;
> no typedefs
A typedef is really like a java bean or class object... it's just a custom data structure.
> It's really freaking hard to write optimised code in this stupid language.
Not true, some of the most highly performant systems on the planet run Java... HFT, stock exchanges, banking, nuclear plant control systems, etc...
There's also the JVM with it's optimizing compiler... one of the best (the best?) optimizing compilers around. Long running Java applications eventually compile hot paths down to native machine instructions, achieving C performance without a lot of the hassle.
But then again, Java wasn't intended to do bit-twiddling, it's a higher abstraction.
[1] https://docs.oracle.com/javase/8/docs/api/java/lang/Integer....
Right. Now try that with a long, and you're replacing every piece of code with something that's, what, 20x more verbose? If not more? And becuase of how bad the JVM is, substantially slower.
Look at JGit.
> int myInt = (myBoolean) ? 1 : 0;
Now try to do that throughout the code. Not to mention which, that's a conditional branch for what should be (and is in bytecode) a no-op. (Well, assuming your bytecode is well-conditioned. But the JVM being the JVM, there's no such thing as a boolean, which means you can have the "boolean" value 2, for instance, which seriously messes all sorts of things up.)
Sure, the JVM is supposed to optimize that out. But you can't rely on it being able to do so. At least not if you're not going to use a mainstream JVM.
And no, the JVM is not "one of the best" optimizing compilers around. It's not even a good optimizing compiler. Nowhere near. For a quick counterexample, this: https://github.com/RS485/LogisticsPipes/commit/bb8a57665c4f8...
Halving the time taken. Why? Because the JVM wasn't smart enough to realize that copying an EnumSet for a readonly foreach loop was unnecessary. Oh, and there's more low-hanging fruit there w.r.t. the amount of work (read: reflection) EnumSet has to do behind the scenes to work around Java's type system. Simple. But no, the Java compiler doesn't optimize it out, and the JVM doesn't either.
And that's with adding an additional unnecessary layer of abstraction (unmodifiableSet) that you wouldn't need if Java had a sane type system.
Copying an EnumSet of a small number of elements would be, in a sane language, quite literally just a register move. Ditto, containsAll a bitwise-not and a bitwise-and. But no, "one of the best" optimizing compilers around cannot even do that.
Trying to write high-performance code in Java generally requires tossing out all the supposed advantages of Java. Write your own object pools, whee! Hardcode your own primitive types, whee! Avoid temporary objects, whee! Avoid using polymorphism, whee! Manually unpack arrays inside objects, whee!
BUT
In this case, to be fair, I took his statement to include Hotspot, which can do some pretty cool stuff provided you meet the requirements for it. ie, be long running, have enough memory and horsepower available on the machine to run hotspot, and have paths through the code that rarely jump around (meaning most executions go through the same path).
If you can meet those requirements, my understanding is that the Java ecosystem does a damned fine job.
The issue is that a lot of javaheads will extrapolate that out to the rest of the language and tech and start making claims about Java being the best overall at X, or as fast as language Y (C or C++, take your pick).
And it still didn't optimize something as trivial as avoiding an unnecessary copy that it was taking ~50% of the time doing.
It's too bad - Java is a fine language in many ways (though it tends to be rather overly verbose for no good reason, but meh. Looking at you getters and setters and lack of operator overloading), but it's saddled with a reliance on the arcane to actually get non-hideous performance out of it. I mean: 13ms per copy of what should be a single integer? (Milliseconds! I'm not joking. 595ms inside 47 calls to EnumSet.copyOf (mainly inside Object.clone))
That's, and I'm saying this quite literally, more than a million times slower than what it should be.
Assuming it needs to be done at all, and you can trivially show that it doesn't.
(That being said, I need to explicitly check that Hotspot does do it's full optimization pass on that chunk of code. I see no reason why it wouldn't, but maybe Hotspot doesn't want to. Though that'd be a WTF in and of itself.)
(On a side note: does Java cache hotspot optimizations? I think it does, in which case there's definitely no excuse. And if it doesn't that's a wtf in and of itself.)
(On another side note: is there a Java bytecode-to-bytecode optimizer that'll do optimizations based on the code you've got in front of you now?)
Java as a tech is strong, but Java as a community was full of pretentious assholes who had a complex about performance (in my opinion of course).
I have no doubt your example was most likely due to some technical issue preventing hotspot from doing what it should have. When hotspot can do it's work it's amazing, you just have to enable it, and you're right about doing arcane things to get performance. That's true in any GC'd language though, even .Net has it's boogeymen.
This is purely an optimization issue, and one that can be done regardless of if a language is GC'd or not.
http://geekandpoke.typepad.com/geekandpoke/2011/09/simply-ex...
(Edit: which is precisely why I prefer BE for anything, unless there is a pressing - usually hardware- and performance-related - requirement for the opposite. Computers are good at reading bytes in swapped order. I suck at it.)
That said, it really didn't matter because all our software had to build and run on x86-64 anyway, so all the endian-specific code was wrapped in preprocessor directives. We could have changed the endianness with a single makefile change.
Yeah, I said it badly. IBM wouldn't guarantee Apple that Power would always support BE, so Apple pushed (or tried to) developers not to assume that.
Apple kept screaming at devs that IBM couldn't be guaranteed to keep the arch bi-endian, so Apple couldn't guarantee a big-endian platform and developers shouldn't assume one.
How expensive do you think endianness-swapping is?
About 1 clock cycle per word.
On the iSeries it has no bearing on application programmers except when using transformation products to other platforms.
This is worse than the HP-UX PA to IA transition, the mac 68k to PPC, or PPC to x86, transition/etc.
The least they could have done was some kind of elf loader, and some fat libraries that flips the execution mode for one generation while everyone transitions.
The stated claim was to make it easier for people to transition to POWER from x86 (aka little porting effort), but what they really did was screw their existing customers by requiring them to port their software (big porting effort). I question, why if I'm going to spend the effort to port from BE POWER to LE POWER why I wouldn't just port to x86...
He specifically mentions cases where the data is stored in BE and the program reads it, swaps it, operates with it, swaps it back again and writes it.
The BE comment seems to me to be of much less importance than the other message (kind of like a side comment). And as you say, BE is very much alive today.
Byte endianess is about binary content, which isn't supposed to be read by humans. Humans read only interpretations of the binary data. You shouldn't come nowhere near to the question of if you are or aren't good at reading binary.
As somebody who had to generate and read back tokenized files decades ago, I can vouch that little endian sucks.
You can tell a lot about someone's programming history from statements like this.
Anyone who habitually works on low-level code has plenty of experience working with raw binary data.
P.S.: You are a bad judge of characters. I worked on low-level code (microcontrollers programming, for a living).
But regardless, I agree that no one is really going to be looking at raw binary. Atleast I can't think of a good reason why anyone would.
I do read hex dumps & disassemblies, and there LE is a huge pain in the backend. Everything is in reverse. Why would a sane programmer voluntarily use LE?
God Save the Network Byte Order!
When Motorola 68000 was out, it was doing 32bits linear addressing.
The mov operation would load directly from the memory by using a direct addressing.
Wiring 32 wires is expensive.
Wiring 2 overlaping address bus of 16 bits partially ovelapping was less. Intel CPU ASM works big endian because we all prefer to write AND 0xFF, DSI to get the 8 lower bit of a number.
The first register would hit segment, and the second the offset.
Hence you would first need to load the page (that would wire the multiplexing) to get the offset in the page/segment.
The addressing bus would be accessed with a stack. So first you push the lower 16 bits, and then the highest 16 bits. Which would revert since its a FIFO in first loading the segment address, then the offset address. The cost of this "savings" in terms of CPU cycle was marginal while the gain in money was tremenduous.
So little-endian is (slightly) more convenient for machines and big-endian is (slightly) more convenient to humans reading memory dumps.
The unfortunate thing is that "network byte order" is big-endian so it's still the traditional endianness to use for wire protocols or on-disk formats. In my own designs I've switched to specifying fixed-width integers as little-endian which makes more sense for 99.9% of computers today.
Contrast with a big-endian representation, where the byte at offset x would have weight 256^(length - 1 - x).
Also, on a 64-bit big-endian architecture, you don't need specialized string comparison instructions to get efficient string comparisons. The code is only slightly more complicated on BE architectures that only support aligned 64-bit loads.
This is literally the opposite of how most programmers in the world write numbers naturally.
See "1234". The digit at offset x from the right has weight 10^x. That corresponds to big-endian.
Considering that the only important difference between little and big endian is when people have to read or write it by hand, we should probably model it after common human representation...
Think about the irony :)
I've become fond of thinking about little-endian becoming known as a universal "CPU byte order", to complement "network byte order". Each order makes the most sense for its domain.
Beyond that, it's called after two stupid factions in Gulliver's Travels for a reason...
Actually LE is the natural order and the question is why anybody sane would prefer BE. The reason you think otherwise is because you view hexdumps the wrong way.
A hexdump shows hexadecimal bytes. Isolate one byte and number the bits -- 76543210, right to left. Number the nibbles -- 10, right to left. Now number the bytes -- 0123..., left to right?
Display your hexdumps with addresses increasing right to left and LE makes perfect sense. This is how numbers are written, so the convention should be followed when displaying numbers.
For showing strings it's better to use left to right, as that's the convention for text.
I've always wondered WHY this flag is present - why would an implementation wish to change the byte-order? I could understand if the file contained machine code, and would vary with different architectures - but given it's running in a VM, why bother?
I guess it could be argued that if you're that concerned about performance and know your target architecture, that it might be worth going native rather than running on a VM.
The original idea was that on big-endian machines, the resulting .odex would be endian-swapped compared to the original .dex, and this field would provide a reasonably-blatant indication. This would probably help with debugging more than anything else (maybe help with security, but not much), because there's no reason anyone would drop a little-endian .odex file on a big-endian machine (or a big-endian one on a little-endian machine).
We wrote (most of?) the code to get the vm working on big-endian systems, but within Google we never shipped any big-endian hardware AFAIK. I have my doubts that anyone ever did.
Do you know of any modern browsers that use BE typed buffers with full WebGL support for a cheap BE machine that we could buy for testing? Does anybody actually use such? Is it worth the bother?
My gut feeling is that it would have been less so than the current solution.
Our current representation is great for skimming in a language written LtR. Placing the most significant digit all the way to the right would be like putting the lead paragraph of an article last.
Yes, they can; that was my entire point. Human-sized numbers fit within the eye's fovea. Numbers that don't fit easily within the fovea require effort under either scheme. Your model of having to go digit-by-digit is a computer's scanning model, not a human perception model; we do not scan that way, we scan in fovea-sized chunks, which are smaller than people may think because the brain is very good at interpolating before the information hits the conscious mind, but is still large enough to fit "millions" in quite comfortably.
The number system we use was imported to Europe from Arabia. In Arabic, the number 1234 is written with the most significant digit on the left -- but Arabic is written right-to-left, so it's actually little-endian. When they were introduced to Europe (with the help of Fibonacci), the convention of putting the most significant digit on the left was retained, but since European languages are written left-to-right the notation magically became big-endian.
The fact that the Roman numeral system that it replaced was also big-endian was likely an influence.