Java 10 – Specification for Value Types
cr.openjdk.java.net
cr.openjdk.java.net
These will have direct impact on the JVM. Is there a test version of the JVM with these bytecodes already implemented? Not sure if Java follows similar rules as the IETF, show us working code with the specifications.
Also, why do you consider the introduction of these new bytecodes as breaking changes for the JVM? Backward compatibility has always been one of the highest priorities for the Java architects.
No and yes.
No, it doesn't really matter for dynamically typed languages (just use Object for generic types).
Yes, it matters for statically typed languages. One of the factors for Scala .NET abondonment was the difficulty in creating a interoperable advanced type system in a fully reified environment.
IMO, full type erasure (like C/C++) is the right way to go. Java erasure is weird only because it didn't do it completely, not because they did it at all.
This line is disproven by all the dynamic languages that happened and failed on the CLR.
At which point any call to the underlying environment fails with a type error. Sorry your List<Object> is not and will never be a List<int>, trying to pass it as one wont do. Either you provide a way to construct a generic type with the right type parameters in the dynamic language or you accept that you can only inter-op with a limited subset of the underlying environment.
Of course that also gets ugly. You now have to deal with the fact that you not only have a List but several incompatible List<T> floating around, adding the contents of two lists suddenly includes the question what sort of list to return, List<Int>,List<Double>,List<Number> or an error? In a type erased context the answer is always the same, you return a List<Object>.
I would sincerely doubt they're considering breaking back-compat in Java 10 (that is, making it so older bytecode targeting something <10 cannot run on 10+).
My understanding of generic type erasure is that it is more of a language issue than a VM issue. Nothing in the VM prevents reified generics, but when generics were introduced in Java 5 they decided to use type erasure in order to avoid having to produce a whole new collections standard library or break compatibility with existing code.
They could have simply left the old library, deprecated it, and added in their genericed one as a a new package. It had to be modified anyway.
Type assertions have no ability to recurse. RN you assert both classes are a UTF8 value in the constant pool. Or really 1 class is a UTF8 value, which the compiler know the first one is.
But with Generics, you may not know what it holds at runtime so this doesn't work. You'd have to constantly rebuild the constant pool (Slow) each time a new _kind_ of generic is used.
So either a new type+byte code would have to be introduced, or the type assertion instruction would have to become recursive. The later is a breaking change.
Also interfaces break this b/c those are class specific. So a Lorg/my/company/myClass$5; and Lorg/your/company/yourClass$0 might both implement the same interface. But that interface information isn't in the collection class, but those to class's class files. So really it just needs new byte code.
There's something seriously wrong with the type system when you want "instanceOf T" checks or "new T".
Instantiating a generic type is like a function call at compile time, where the type arguments are substituted for type parameters in the body of the generic type. It's a variable at compile time, but it's a concrete type at runtime for any version of the code that can execute.
I haven't given it a ton of thought but I thought that apart from reflection it would actually be nicer if the runtime erased types. I.e. at runtime a type should just be either a stack/copy/move type or a reference type. So erase all reference type to objects.
It feels like it's the job of the compiler to handle these. Obviously at runtime you can get some extra safety guarantees from reified types, but at the cost of languages on the runtime being more difficult to adapt to other paradigms etc.
I imagine a runtime that doesn't have strong opinions about types to be easier to make type level language changes for (i.e. the language(s) can evolve without the need for runtime changes), and it should also be easier to support entirely different type systems.
I know - but I'm not familiar with the situation where I'd have to to that in java. Do you mean e.g.
Foo(object x)
{
if (X is List<int>)
...
else
...
} void AddAnItem<T>(List<T> aList) where T : new()
{
aList.Add(new T());
}
Or for instantiating generic types: public TColl CreateACollectionAndAddAnItem<TColl, T>(T item)
where TColl : ICollection<T>, new()
{
TColl aList = new TColl();
aList.Add(item);
return aList;
}
// Usage
List<string> myList = CreateACollectionAndAddAnItem<List<string>, string>("hello"); <T> static void addItem(List<T> list, Class<T> type) {
T item = type.newInstance();
list.add(item);
}There are some ways to introspect generic types in Java, but you need a concrete binding. For instance, if a method returns List<Integer>, you can in fact see that it returns List<Integer> and not just List. Method.getGenericReturnType() would return a ParameterizedType with a List raw type and Integer type arguments. But that requirement for it to be a concrete binding means it's not really helpful from the context of writing a generic class or method in the first place.
Using the above example, I'm not sure how you could desugar that without having the type known to the runtime. The generic method is going to be the same code no matter the type provided, but the type provided is necessary in order to know the type to construct. So either the runtime must provide the type, or it must be given as a parameter.
Additionally, glossing over the issue like that creates a huge trade-off. Now programmers must build a mental model of when the compiler can and can't do the type binding. The difficulty of building such a mental model accurately is one of the central complaints against Rust's borrow and lifetime checker(s).
public Object[] toArray()
public <T> T[] toArray(T[])
You cannot turn an ArrayList<T> into a T[], for example, without passing a T[] into the function so that you can grab the type from the passed parameter at runtime. C#, which retains the generic information at runtime doesn't need this so you can just do public T[] toArray()
The other place I've seen this come up is with Exceptions- you cannot catch a generic Exception. It's pretty annoying now that Java 8 has streams because you cannot have a checked exception abort out of processing a Stream unless you either 1) catch Exception (rather than the specific subclass) and deal with everything or 2) catch the checked exception in the lambda, wrap it in a runtime exception, rethrow it, and then catch the wrapped exception.When Dotty and Java 10 land many annoyances will go away thankfully.
x match {
case _: List[Int] => 1
case _: List[Char] => 2
case _: String => 3
}
This used java's semi-erased class tags and doesn't report a TypeTag constraint in its type. It will also return 1 when x is of type List[Char] since partial erasure means that the first two patterns end up being identical. The compiler will warn you about this situation, but generally it shows up all the time for various reasons. Super bad news.That all said, the TypeTag system could be very nice someday. Especially if asInstanceOf were dropped eventually or, better, relied upon TypeTag.
All runtimes need to erase types at some level. Otherwise ArrayList<Foo> and ArrayList<Bar> would end up compiling identical versions if both Foo and Bar are reference-only types, which just wastes memory. At some level the compiler and runtime need to merge duplications - in C++ that feature is called COMDAT folding, or used to be.
.NET has had serious problems with code duplication in the past. Here's an excellent blog post by a Microsoft engineer on it:
http://joeduffyblog.com/2011/10/23/on-generics-and-some-of-t...
"... instantiations over reference types are shared among all reference type instantiations for that generic type/method, whereas instantiations over value types get their own full copy of the code."
Just about the only thing it could do better is to reuse the same instantiation for all value types of the same size.
Duplication exists with generics and value types, e.g. List<long> and List<DateTime> are entirely separate code. It's just a thing to keep in mind when mixing them with value types.
Look at C# on the other hand and you see they had to add entirely new types which means that the .NET framework has an ugly split between APIs based on the old container classes and those based on the new container classes. That fits the trend that C# is a better language than Java, but Java has better class libraries to work with.
They took a risk, and it paid off. In my opinion, C# is both a better language and has better class libraries.
As for better class libraries, I found the .NET BCL to be excellently designed and thought through. It's also very consistent throughout. Now, parts of the FCL, like System.Windows.Forms are another matter ...
There's still some warts, such as there not being an ISet<T> before .NET 4, but Java's standard library has its share of quirks and historical weirdnesses as well. And as libraries age there are always old ways of doing things you can never really remove, and newer ways that are better. None of the two is as bad as C++, but depending on what you do you can stumble around in a swamp of old APIs for a while before finding what you're actually supposed to use.
That's covered by a sibling valhalla feature called 'generic specialisation'. See http://openjdk.java.net/jeps/218.
Disclaimer - I help run the adoption group and maintain the Valhalla wiki.
Does anyone with the slightest familiarity with JVM languages not understand that each major Java language version is also associated with a major JVM version?
What this does provide is a mechanism by which the use of value types can be experimented with by library and language authors. I believe these value type features are intended to be optional and are not guaranteed to remain unchanged at future JVM releases.
Java SE 10, the platform, will support value types. One can interpret the ambiguous term "Java 10" as either "Java 10 the language" or "Java 10 the platform". Applying the principle of charity[0] will yield the right interpretation in this case.
There is even more: invokeDynamic, Java (the language) doesn't use it in its compiled bytecode and the JVM instruction is targeted at JVM based languages. (InvokeDynamic is still available from Java via java.lang.invoke)
while invokedynamic was introduced to enhance the support of dynamic languages on the JVM, invokedynamic is now used as a kind of macro instruction to avoid to add new instructions in the bytecode set when the instruction can be decomposed in a set of already existing instructions.
I felt the same slight sadness when I saw the complexity of the planning involved in Java value types, and the many-year path to get there. Intuitively, a sufficiently smart compiler should have been able to take care of this, and in some cases it does do. So it's worth reflecting on why adding new opcodes and such is necessary.
HotSpot and some other JVMs can do an optimisation called "scalar replacement", which converts objects into collections of local variables which are then subject to further optimisation. So for example a Triple<A, B, C> type class could be converted into three variables, then the optimiser notices that the second element of the triple was never actually used anywhere, and eliminates it entirely.
Scalar replacement is one of several optimisations that relies on the output of an escape analysis. Escape analysis is a better known term so people often use the name of the analysis to mean the name of the optimisations it unlocks, although that's not quite precise. Sometimes people talk about stack allocation, but that isn't quite right either. Only objects that don't "escape" [the thread] can be scalar replaced.
There are several reasons this isn't enough and why Java needs explicit support for value types.
Firstly, the JVM implements two different EA algorithms. The production algorithm that is used with the out of the box JIT compiler (C2) is somewhat limited. It can identify an object as escaping just because it might escape in some situations, even if it often doesn't. There is a better algorithm called "partial escape analysis" implemented in the experimental Graal JIT compiler, but Graal isn't used by default. In Java 8 it requires a special VM download. It'll be usable in Java 9 via command line switches. PEA can unlock optimisation of objects along specific code paths within a method even if in others it would escape, because it enables the un-optimisation of escaped objects back onto the heap.
Graal unfortunately can't be just dropped in. For one, you don't just replace the entire JIT compiler for a production grade mission critical VM like HotSpot. It may take years to shake out all the odd edge case bugs revealed by usage on the full range of software. For another, Graal is itself written in Java and thus suffers warmup periods where it compiles itself. The ahead-of-time compilation work done in Java 9 is - I believe - partly intended to pave a path to production for Graal.
Implementing PEA in C2 is theoretically possible, but I get the sense that there isn't much major work being done on C2 at the moment beyond some neat auto-vectorisation work being contributed by other parties. All the effort is going into Graal. I've heard that this is partly because it's a very complex bit of C++ and they're afraid of destabilising the compiler if they make major changes to it.
Unfortunately even with PEA, that still isn't enough to replace value types.
(P)EA can only work on code that is in-memory when the compilers optimisation passes are operating, and moreover, only in-memory in the form of compiler intermediate representation. In a per-method compilation environment like most Java JITCs that means it can only work if the compiler inlines enough code into the method it's compiling. Graal inlines more aggressively than C2 does but it still has the fundamental limitation that it can't do inter-procedural escape analysis.
It can be asked, why not? What's so hard about inter-procedural EA?
That's a really good question that I wish John Rose, Brian Goetz and the others had written a solid full design doc or presentation on somewhere. They've said it would be incredibly painful, and obviously view the Valhalla (also incredibly painful) path as superior, so it must be hard. In the absence of such a talk I'll try and summarise what I learned by reading and watching what's out there.
Firstly, auto-valueization - which is what we're talking about here - would necessitate much, much larger changes to the JVM. Methods and classes are the atoms of a JVM and changing anything about their definition has a ripple effect over millions of lines of C++. HotSpot isn't just any C++ though - it's one of the most massively complex pieces of systems software out there, with big chunks written in assembly (or C++ that generates assembly). The Java architects have sometimes made references to the huge cost of changing the VM. They clearly perceive changing the Java language as expensive too, but compared to the cost of changing HotSpot, it can still be cheaper to complicate the language. Frankly it sounds like Java is groaning under the weight of HotSpot - it's fantastically stable and highly optimised, but that came at large complexity and testing costs that have to be considered for every feature. Part of the justification for doing Graal is that the JITC is the biggest part of the JVM and by rewriting it in Java, they win back some agility.
As an example of why this gets hard fast, consider that the compiled form of a "valueized" version of a parameter is different to the reference form. Same for return value (a value can be packed into multiple registers). So you either have to pick a form and then box/unbox if the compiled version doesn't 'fit' into the call site, or compile multiple forms and then keep track of them inside the VM to ensure they're linked correctly. And a class that embeds an array may wish to be specialised to a value-type array, but then you need to detect that and auto-specialise, detecting any codepaths that might assume reference semantics like nullability or synchronisation, and then you have to be able to undo all that work if you load a new class that violates that constraint.
Secondly, detecting if something can be converted into a value type (probably) requires a whole program analysis. Java is very badly suited to this kind of analysis because it's designed on the assumption of very dynamic, very lazy loading of code. Not just because of applets and other things that download code on the fly, but it's quite common in Java-land for libraries to write and load bytecode at runtime too. Also Java users and developers expect to be able to download some random JAR and run it, or a random pile of JARs and run it, with essentially no startup delay. HotSpot can run hello world in 80 milliseconds on my laptop and the very fast edit-run cycle is one of the things that makes Java development productive relative to C++.
All that said, applets are less important than they once were, and Java 9 does introduce a "jlink" tool that does some kinds of link-time optimisations:
https://docs.oracle.com/javase/9/tools/jlink.htm#JSWOR-GUID-...
https://gist.github.com/mikehearn/930e3001e76415ae1e2d85dcdf...
Note that they solved their problems using exactly the features being added in Java 9.
Java has two fundamentally different kinds of types: objects and primitives. Value types are user-defined primitives. Seems pretty straightforward, from a language design perspective. I'm not sure why we're trying so hard to avoid this.
You'd need to generalize stacks to full-fledged region analysis to see any real benefits, but only MLKit has taken it this far IIRC.
Nothing's impossible I guess, but this kind of global datastructure reshaping is definitely far beyond the state of the art in Java JIT compilers.
My mistake--I thought "regions" meant some sort of interprocedural "code regions".
I suspect the state of the art in GC has moved in the last 15 years since this paper was written, and they might have trouble beating modern techniques, but that's speculation.
Escape analysis is a degenerate region analysis, and many of the cases where it fails because it's not general enough would be handled quite well by regions.
I worked for years on escape analysis in IBM's Java JIT compiler. We struggled to find any actual programs that showed any benefit at all. The real benefits of escape analysis were second-order effects like eliminating monitor operations on non-escaping objects, or breaking apart objects and putting their fields into registers (especially autoboxed primitives). The actual stack allocation wasn't really any faster than heap allocation, and a GC operation in the nursery doesn't even look at dead objects.
EA is basically a microbenchmark-killer. For real software, it's not often worth the trouble.
Seriously though, the JVM is killing it. Can't wait until I can compile my .jar to verilog and run it on my FPGA, or send the verilog to TSMC and get some ASICs printed out.
C++ is built on these things. The "value type" thing informs everything about the language. I love it, and it's my go-to language when something just has to be fast, but it comes along with some hefty baggage. Java gets a lot of mileage out of its reference-first philosophy that C++ would be wise to try out -- the "everything must be possible with no overhead" philosophy of the language bleeds into the programming community in unhealthy ways, and the amount of language machinery to make the value stuff work "as well as possible" (https://stackoverflow.com/questions/3601602/what-are-rvalues...) is absurd.
> Seriously though, the JVM is killing it
As for the JVM and Java, this seems like a guilty admission that the CLR and C# ate their lunch a decade ago (perhaps not wrt performance, but for language features....) I've never worked in the Microsoft stack, but the stories have been the same since the dawn of the age -- "C# is Java with real generics, value types and lambda functions." First they laughed...
They are immutable so you have no way to know if they are stack allocated, split over several registers or on heap.
This is important because it means that JITs will not have to be changed, they will not have to manage stack aliases like in C#/C++.
One example that I run many years ago was having a struct Foo { int x; } which was nested inside a List<Foo> myList. Trying to mutate one list element with syntax like myList[0].x = 27 silently failed, because the indexer of the list returns a copy of Foo and you mutate that inside of writing it back. Nowadays its a compiler error, but back then it wasn't. The workaround is something like "Foo f = myList[0]; f.x = 27; myList[0] = f;". Which is quite weird if you are coming from C++. From a C# point of view it's actually obvious, since you can't pass around references and pointers to structs inside of other things - and if you also add this functionality, you get lots of the complexity from C++ back (references to structs, functions/properties that return references, questions around reference lifetimes, etc.). Without having also reference semantics for structs, structs also can't achieve the kind of performance which "value types" in C++ have, since they sometimes need to be copied fully.
Another nice gotcha is the surprising behavior of readonly structs, which e.g. is documented here: https://bytes.com/topic/c-sharp/answers/261922-readonly-beha...
For those reasons I think it's quite fair that Java and other languages did not immediately jump onto the value types train. I think introducing them carefully, and trading off the performance gains against the amount of new complexity in the language is fine.
Eiffel, Oberon, Modula-2+, Modula-3, VB, are just a few examples.
Hence why I always saw as lost opportunity not to provide them in first place.
I disagree. The JVM continues to eat the lunch of the CLR in performance, developer mindshare, pervasiveness and ecosystem. And with each new Java version those language feature differences continue to decline. And if you do need that syntactic sugar then there's always Kotlin.
I agree with you that C# is a much nicer programming language than Java. The HotSpot JVM, otoh, is a masterpiece of engineering in terms of its JIT and GC optimizations.
Ah, my fault. You expected more than the audience deserved.
The reason why this never happened was that Java was good enough. A quick look at TIOBE shows C# stuck at around 3% while java continues to reign supreme.
I do think C# is a better language in many ways, but not, it seems, in any ways that really matter for purposes of language adoption. Which is a shame, since I enjoy writing code in C# more than in java, but I don't have much opportunity to do it, and I've never been in a situation in which C# was the language and it wasn't an MS-only solution.
http://www.jesperdj.com/2015/10/12/project-valhalla-generic-...
> Finalization. Finalization makes no sense for values and should be disallowed (though values may hold reference-typed components that are themselves finalizable.)
This is a little disappointing. I was hoping that the lifetime of essentially stack based variables could be tracked, and finalization could be bound to the stack. That would be really powerful; though mind bending after beating so many Devs to not rely on finalization.
Imagine wrapping a Closable object in a Closing tuple for instance, and that is guaranteed to run when the tuple goes out of scope.
Edit: I'll add that it's possible that the proper implementation for even a bog-standard, assign-to-a-variable copying would need to work this way (make another finalization-delaying handle) but I think that's less-certainly true.
Of course, since value types are immutable, that would be a bit memory wasteful I guess.
I guess it doesn't typically make sense to capture a RAII object and have the copy outlive the original, but it also may not make sense to delay the lifecycle of the original since it breaks the expectation of "i put this on the stack, so expect it to finalize when it goes out of scope". Probably best not to mix the two. That is, a stack-allocated RAII object whose finalizer is called when it goes out of scope, which references some data that can be closed over, with a lifetime separate from the RAII object.
At the very least, you would call finalizers on both (the stack one when it goes out of scope and the copy when it gets garbage collected) and then its up to the programmer to make sure that shared resources are finalized correctly (eg by keeping a shared reference count...). (I'm thinking a use case where the object represents some external resource that should be freed only once)
Sounds pretty error prone though.
I suppose that Java's solution is KISS: simply don't support finalizers on value objects.
I did notice one point that struck me as odd (under "Details, details"):
> Can a value class implement interfaces? Yes. Boxing may occur when we treat a value as an interface instance.
I don't normally think of values in other languages as having methods directly or an interface like this. I wonder what the primary reasons for this would be. Backwards compatibility might be one reason.
[1] https://blogs.msdn.microsoft.com/abhinaba/2005/10/05/c-struc...
If I understand you correctly, Haskell has this, and it's basically all you have to implement generic interfaces in Haskell.
A simple example is the Bounded class, whose interface exposes two functions: minBound and maxBound. So, for example, all integers (int8, uint8, int16, uint16, [...], uint64) have upper and lower bounds, and can implement this interface, like so (Word8 is uint8):
instance Bounded Word8 where
minBound = 0
maxBound = 255
This forces all instances of a class to be a type that contains all information needed to implement the interface (as opposed to some construct that needs to accumulate state at runtime in order for it to work properly), which means the compiler can check almost everything at compile-time.This is where monads come in, which is simply a value that describes a computation, e.g. fetching a value of a specific type from a web server at a given URL. So, in Haskell, all that exists are values, and they're also used to represent an operation which, at runtime, performs a certain action.
So, for example, the following function:
concatStr :: String -> String -> String
is a pure function which only modifies it's two arguments (String and String) and produces a String, whereas printReadStdIn :: String -> IO String
printReadStdIn str = do
putStrLn str
getLine
is a function which, when given a String, will -- at runtime -- print out a String and read a String from standard input, and this String will be available inside the IO monad (which can implement interfaces, because it's just a value).It's been moved a lot and AGAIN, but I believe "the target is starting to move slower" :)
I suppose that at some point in the future, the Java language spec will be updated to let us play with that.
e.g. You could define a tuple-like type that will get passed and stored by value. Now you can make a billion member Array of 'em and not incur the speed and memory overhead you'd get from using Objects that need to be tracked, dereferenced, garbage-collected, etc...