1) People use Apache Storm and Cassandra daily. These need Unsafe to function at current performance changing its functionality may change these tools. This can be considered existiential threat to their job.
2) The one big selling point of the JVM is that JVM 1.0 code will run in JVM 8.0 without issue, and often much faster. With this change that isn't necessarily true, and potientally when your .jar was compile will weigh on which JVM is used or HOW that jvm is started.
There's a lot of software out there that uses MD5 or DES right now. Saying that the software needs MD5 or DES is a different claim entirely.
As there is no alternative to unsafe serialization in terms of performance. It goes without saying.
>There's a lot of software out there that uses MD5 or DES right now. Saying that the software needs MD5 or DES is a different claim entirely.
I agree, but you are putting the trees before the forest.
Saying implementation X is flawed is one thing. Saying We are removing implementation X and offering no alternatives is another. While there are plan to offer an alternative, they are simply plans not an immidate alternative, thus a problem. When DES/MD5 were removed it was clear that they were already inferior implementations, and superior alternatives existed. That isn't true for java.
I don't think that's true, though perhaps I don't understand what you mean by "unsafe serialization".
First, I suspect one can use JNI and get at least as good performance. There are other tradeoffs, like the ease of packaging, but in order to evaluate those we need a clear problem statement.
Second, premature optimization is the root of all evil, and we should be looking at realistic benchmarks to determine that there is a performance gain. There are many documented examples of software safety checks adding negligible performance overhead because they get branch-predicted away. Alternatively, new hardware support like Intel's MPX adds bounds-checking at the instruction level, so if we're talking about raw writes to a bounds-checked buffer (which is a vast improvement over raw writes to anywhere), this can probably be implemented efficiently on new hardware. On old hardware, the JVM could implement these as unchecked writes, which preserves the exact same performance, but the API would now have information to add safety where it is performant.
Finally, sun.misc.Unsafe covers a lot of ground. If we can restrict it to the specific uses that these projects make of it, that's still an improvement. I don't think these libraries use the whole thing, do they?
Very few things in science go without saying, and software performance and correctness is science.
http://www.javacodegeeks.com/2010/07/java-best-practices-hig...
Is the specific Unsafe use here to serialize a Java object into a byte array? The linked article doesn't describe using Unsafe, so I'm still not sure what functions are at issue.
This is what I was referring to about a bounds-checked buffer in my parent comment: it seems like you could state that accesses to an area of memory are unsafe and unchecked, except to check that the reads/writes are within the buffer. This preserves safety for the JVM as a whole, but gets you native performance within the buffer. Intel MPX should be able to implement this efficiently, and the standard techniques in other languages for efficient bounds-checked memory access should all apply.
Am I completely off-base here?
Java VM isn't native code. Serialized objects are "Field"(s) within the JVM. Which is short hand for a "Growable Page File" more or less. Which starts at a fixed 32bit address. How does that work? JVM implementation takes care of that, so you decide.
On top of that whats inside the Field itself doesn't actually matter to anyone provided the same read/write opcodes do the right thing(s).
The serialization process starts by jumping to your initial field, copying the serialized data into another buffer. While checking integrity, and assuring that your data conforms to the serialization standard b/c its in memory representation doesn't. The fun part is when ever it encounters a RetAddr type it has to jump to that Field, and serialize THAT Field also, as that Field is part of the object your serializing. Of course you don't actually know whats in the Field without consulting its ClassField which gives you a basic prototype of what-is-where.
Furthermore what ever optimizations the JIT makes, it'll end up treating Fields more-or-less like stack frames. So you have to keep an active record of how it'll mangle memory, so you know how to untangle that when you want to do serialization.
Understand now?