Value Types for Java – A sketch of proposed enhancements
cr.openjdk.java.net
cr.openjdk.java.net
Generics over values. It is persistent irritant that generics do not work over primitive types today. With value types, this irritation will increase dramatically; not being able to describe List<Point> would be even worse. Still worse would be having it mean “today’s list, containing boxed points”. Supporting generics over primitives and values, through specialization, is the topic of a separate investigation.
The need for boxing/unboxing is for me one of the most common sources of facepalm in Java. I guess it's difficult to fix without breaking existing code (I guess that a new implementation of generics may not break java code, but it would break existing bytecode), but it's really annoying.
Python learned it the hard way. The value of backward compatibility got a bit underestimated there, with Java it's going the other way (the value of progressing, giving better tools to developers is overlooked). The JVM is a big bag of crazy code already (HotSpot with C1 and C2 JIT compilers, the many garbage collector algorithms/implementations and the other stuff that is currently deemed the JVM's responsibility), so there's plenty of space for new opt-in features.
I think you're missing the point of generic programming.
Sort of off-topic, but isn't that dilemma also the one preventing Go from having generics?
Meanwhile, in Haskell, there already are value types and newtype wrappers are erased at compile-time, eliminating all overhead.
Scala does true specialization of primitives by means of the @specialized annotation: http://www.scala-lang.org/api/2.10.3/index.html#scala.specia...
This functionality is old and because it tends to generate many class files, there's also a newer project that fixes that and that should replace the @specialized annotation soon: http://scala-miniboxing.org/
But wait, there's more - Scala 2.10 also brought with it "value classes", as in zero-overhead wrappers: http://docs.scala-lang.org/overviews/core/value-classes.html
Speaking of the JVM and overhead, the JVM also does escape analysis and it can choose to allocate values on the stack in case they don't escape their context: http://docs.oracle.com/javase/7/docs/technotes/guides/vm/per...
This is actually a big problem. The key word is "flow-insensitive". Because current JVM escape analysis is flow-insensitive, it misses lots of opportunities.
Thankfully, this is changing. Oracle recently published "Partial Escape Analysis" paper http://dl.acm.org/citation.cfm?id=2544157 and I hope it will be integrated in future versions of JVM.
I follow Graal since the Maxime days and am looking forward to the day it might replace Hotspot, specially if Truffle comes along.
What might already happen in Java 9, given the Graal/HSAIL ongoing work.
But
It's a bit sad how java struggles to become a C# from five years ago. Would make a lot more sense to leave java alone, add reified generics and value types at the bytecode/vm level and instead launch a new language that could use all of it.
Throw in pattern matching, algebraic data types and a bit more and I'm sold.
Add the Sun troubles into the mix, and slow the features go.
Actually type-erasure is one of Java's best features, because it didn't cripple the runtime for other languages. For example Scala's type system is too expressive to be built on top of .NET's generics.
First of all, reification only helps in terms of type-safety if you have a weak type system. The only value that reification, as implemented in .NET brings is specialization for value types.
Second of all, specialization can be done by the compiler. That's how many languages have always done it. Scala for example has a @specialized annotation for specializing type parameters for primitives [1]. And because this functionality has some gotchas, there's also an up&coming project plugin that's meant to be a replacement in future versions of Scala and that works really well [2]
I myself have used the @specialized annotation on many occasions and in general it does what it's supposed to do. And if that means I can develop in languages like Scala, Clojure and JRuby on top of the JVM, then I really, really love type-erasure. And oh, apparently one barrier for implementing type-classes in F# are the generics in .NET ;-)
> It's a bit sad how java struggles to become a C# from five years ago
What I find sad is grown men that can't see the forest from the trees. Java, the language, is totally irrelevant and uninteresting and I like it that way, because it's the kind of language you can rely on in case you want backwards-compatibility (yes, that's a feature). On the other hand, the JVM is light-years ahead of the CLR, because that's what the Sun, now Oracle engineers have been doing with all of their time.
> pattern matching, algebraic data types
See Scala. It also does type-classes and higher-kinded types, amongst others that F# is not capable of.
[1] http://www.scala-lang.org/api/2.10.3/index.html#scala.specia...
...
People keep repeating this stuff and they're totally wrong. F* is dependently typed and is a CLR language.
bad_user, have you ever actually attempted to write a type unsoundness proof for your claim or are you just repeating something you heard somewhere?
If you have no such requirements (e.g., Scala, where there's no requirement that Java be able to consume or interpret all of Scala's types), then you can most definitely define a language for the CLR which does this.
F* is not a typo -- that is not F#, it is a separate language.
Personally I don't like F#, the result, because on one hand you've got the Hindley-Milner type system and on the other hand you've got OOP with .NET generics. It resembles Ocaml in philosophy (I'm sure this wasn't by accident), in that it feels and smells like 2 type-systems in the same language. It's just an opinion of taste of course - but would have anybody used F# if it didn't integrate well with .NET's standard assemblies? So F#'s integration with .NET's generics is very understandable and one reason Scala.NET died was because it was too hard to keep the semantics of Scala, while also integrating well with .NET, which meant the project was interesting only for people wanting to port their Scala code to .NET - for which approaches like IKVM.NET were far better options. I think somebody wrote an interesting paper on Scala and .NET's reified generics, sadly I can't find it right now.
You may have gotten annoyed about hearing this thing about .NET generics not being expressive enough - I also get annoyed hearing about how Java sucks because it doesn't have reification for generics or other such things, when in truth the compiler can handle specialization just fine and both approaches (runtime versus compile-time) have merit and make different compromises.
BTW, I see that you're working on the Roslyn team - you guys are doing great work lately. Keep it up.
I happen to agree with you -- I think the technical limitations in both languages/VMs are highly overrated. Java, especially, has been a victim mostly of insufficient leadership and resources, not any fundamental technical issue.
C/F# and the CLR, meanwhile, are probably most limited by backward compatibility.
The LAMP team actually got pretty far with Scala on the CLR. The reason it didn't go farther was more due to a lack of interesting and funding.
Specialization for value types (including ones not known at compile time) is the main point imho of reified generics, yes.
I like scala and agree that F#'s typed sometimes feels like the bastard child of ML and C# (although I do like ML syntax better).
I agree with you that the backwards compatibility of java is a feature now, since so many other languages incorporate what we'd otherwise want to see in java.
I thought that compile time specializations wouldn't reach all the way, just like C++ templates don't.
That being said, not being Java programmer, I'd like to know what's the state of the libraries now? Do libraries handle the collections of the primitive types and the collections of the objects differently now, using the most efficient infrastructure, or is it something that you have to select as the programmer manually even now on case-by-case basis?
(edit: tried to adjust the terminology for Java)
So, in general, most libraries just ignore collections of primitive types.
Background: I did a lot of 3D app dev in Java in the 1998-2001 time frame and then switched over to C#. C# was a huge improvement because it had value types. When you have 5M vectors, you really do not want each to be its own allocation.
My assertion is simply that compared to the cost of serializing to/from database/user, it is probably not the pain point of many programmers. Hence, most programmers probably don't think about it.
Just ask a few developers about the ridiculous overhead of HashSet<Integer>. And then see how many places people still use it.
"Wrapping. Just as the eight primitive types have wrapper types that extend Object, value types should always have both an unboxed and a boxed representation. The unboxed representation should be used where practical; the boxed representation may be required for interoperability with APIs that require reference types. Unlike the existing primitive wrappers, the boxed representation should be automatically generated from the description of the value type."
So there is some boxing planned. What does it mean for the programmers, the users of existing libraries? Ah, I believe gordaco hints to that in his comment.
edit: And yes, I see it says they will be auto generated. It is the "how to distinguish" question that is non-obvious to me.
Unboxing can also produce counter-intuitive typecast behavior.
If I want to cast an object-boxed int into a short
Object myBoxedInt = GetInt();
short myShort = (short)(int)myIntObj;
because casting directly to short is a run-time exception. I've run into this problem in C#, but my understanding is that Java shares it. Integer boxed = ...;
short myShort = boxed.shortValue();You can't use the built in collection classes on unboxed primatives, but there are a few different popular third party libraries that help overcome that. For example, Google guava has utility classes to work with arrays of primitives as if they were lists (https://code.google.com/p/guava-libraries/wiki/PrimitivesExp...). There's also GNU trove, and Goldman Sachs' java library. Maybe somewhere in the apache commons library too, I don't remember.
(defstorage (file-header)
(version fixnum)
(logical-size fixnum-bytes 2)
(bootload-generation fixnum) ;for validating
(version-in-bootload fixnum) ;used in above
(number-of-elements fixnum-bytes 2) ;for later expansion
(ignore mod word)
(* union
(parts-array array 8 fh-element)
(parts being fh-element ;and/or reformatting
fh-header ;this header itself
info ;header-resident info
header-fm ;file map of header
dire ;directory-entry-info
property-list
file-map
pad-area ;many natures of obsolescence
pad2-area)))I will be curious to see how this affects code targeting android and friends.
There was a wave of excitement about doing numerics in Java around 2000ish, but it seems like everyone has since moved on.
* Has the proposal progressed beyond this 'initial' version?
* What were the circumstances that led to it not being public originally?
* What circumstances have led to it having been made public now?
The single-field restriction is there because there are limits in how aggressive one can fight a non-cooperating runtime, not because it was their principled choice.
Their idea is, IIRC, to flatten the members of a type into an array of primitives.
Here they want a discussion by implementers about the technical side and avoid bikeshedding about syntax (for now).
1.Defining Value Types? A. What about forcing a primary constructor to indicate that it is a value type, there is some thing like it in C#6, but in C# it just to make easy to define a primary constructor. So in Java we can use it differently. So we can say that if we use it that we are defining a value type. It is possible to have more than one explicit constructor. But I don't think that we have to force developer to write a constructor if it is not necessary. Why? According to the document Value-types constructors are really factory methods, so there is no new;dup;init dance in the bytecodes.
So I think that to define a value type we can only do: <p><pre><code> final class Point(int x, int y){
boolean equals(Point p){ return this.x==p.x&&this.y==p.y}
}; </code></pre></p> x and y are automatically final and public members. So that can exactly be compiled like: <p><pre><code> final __ByValue class Point { public final int x; public final int y;
public Point(int x, int y) {
this.x = x;
this.y = y;
}
public boolean equals(Point that) {
return this.x == that.x && this.y == that.y;
}
}
</code></pre></p>
Simple, concise and contributes to productivity. I also I thought about something new.
We can also do:
<p><pre><code>
final class Point(int x, int y){
private int c;//if I wont
boolean equals(Point p){
return this.x==p.x&&this.y==p.y}
public static Point getFunPoint(){
}
public static void maFunction(Point p, boolean b, double b){
}
}
</code></pre></p>
B. Why not an implicit default equals. So if I do :
<p><pre><code>
final class Point(int x, int y){}
</code></pre></p>
that means if the JVM doesn't find an explicit equals, it cans do logical compare.
<p><pre><code>
public boolean equals(Point that) {
return this.x == that.x && this.y == that.y;
}.
</code></pre></p>
So a Point can be defined in one line:
<p><pre><code>
final class Point(int x, int y){}
</code></pre></p>
C. If we don't want to define something else in a value type, braces can be optional as in lambda expressions :)
So to define a value type we can only do:
<p><pre><code>
final class Point(int x, int y);
</code></pre></p>The fact that a final class cannot be represented and optimized to a simple struct in a non-dynamic language is really disturbing.
A final class can be represented as a simple (plain) struct if the scape analysis tells so, and allocated on the stack. Else, that inmutable object can be used as a lock, inserted into a collection (think hash) or whatever you can do with an object.
I wouldn't have moved from Java to C# for desktop apps back in 2001 if Java had stated they would have adopted value types quickly, but they didn't.
History has moved on though and Java isn't really used for desktop apps at all now so this is a moot point. (And C# is on its last legs as well as a desktop development tool, except in large organizations...)
Downvoted to -1 for stating a fact. That has never happened before. Edit again, seems I've been upvoted again. Rollercoaster ride today.
What's used instead?
Projects in the chemical industry, desktop software to control the manufacturing machines and process result datasets.
But now I do JavaScript + WebGL (which is admittedly slower than both C#, Java and C++): http://clara.io
The only real downsides are scaling issues that npm is running into, the fact that Google moved much of its V8 team to Dart, and no WebGL on iOS.
Node.js is just an implementation of the Reactor pattern with some package management.
Google are smart people, they understand that dynamic weakly typed language is a dead end, hence Dart for client and Go for server.
https://xamarin.com/ imho a much better alternative to C++. I really like C# as a language, it's got that general purpose excellence, but when you get on with things like LINQ,Rx,TPL it's very hard to go back to a language without such first class citizen frameworks.
I agree that Microsoft abandoning XNA development was a very disappointing decision, it was a pretty good tool.
Besides being mostly off-topic (not entirely, since it does talk about value types), this comment is arguably a middlebrow dismissal, since it dissmises a serious piece of work without engaging anything interesting in it, and for a superficial reason.
That reason (Java moves slowly) may be valid or may not be, but it isn't interesting in the HN sense of the word, because the route from the article to there is entirely predictable. Such comments get upvoted because they elicit the pleasure of recognition ("I've noticed Java moves slowly too!" or worse, "Yeah, Java sucks!"), and that is a reflex reaction, not a thoughtful one. We all do it—but if we're to have reflective discussions instead of reflexive ones, we need to inhibit it.
The larger problem here, though, is the dreadful subthread the comment spawned. HN threads are sensitive to initial conditions, and middlebrow dismissals lead to lowbrow rejoinders. It's the online comments equivalent of "B players hire C players".
All: please re-read what you're posting and reflect on it. If it goes on a tangent, make sure that tangent is not a predictable one. But if there aren't many comments in the thread, try not to go on a tangent at all—sensitivity to initial conditions means it may well derail the discussion.
That's a very interesting point: the lack of this very feature closed Java off from an entire field. That implies that if this feature is implemented, it could open up that field to Java. (And, presumably, others with similar requirements.) I wish that bhouston had expanded on that subject.
This, I think, can make the difference between a middlebrow dismissal and an insightful contribution. Instead of framing a comment in terms of what is wrong, frame it in terms of "what would this enable"? You can still make the same fundamental points ("Java lost 3D visualization because it lacked this feature"), but the ensuing discussion can be very different. The post changes from a dismissal to the start of a discussion.
After 14 years since C# was introduced, it still doesn't do escape analysis: http://docs.oracle.com/javase/7/docs/technotes/guides/vm/per...
I think stack-allocated values by means of escape analysis done at runtime is just an extension of that.
Also, escape analysis, theoretically at least, goes beyond what one can do with value types. For instance functional/persistent data-structures are guilty of generating a lot of short-term junk and because internally these data-structures are represented as trees, you need references for linking the nodes together. On the whole I am pleased with how well the JVM handles short term junk, whether escape analysis plays a part in it I don't know.