State of Valhalla
openjdk.java.net
openjdk.java.net
I have been waiting for these fixes to the language since I was first introduced to Rust which got the type system “correct” where Java had left a lot to be desired.
The big thing that Java will get beyond generics and such being cleaned up for all types, is true Optional types. That is, a stack based Optional that is itself not null, but tracks something that might be null in the interior.
It will take a long time to update the ecosystem to correct this, but when it comes it will make Java a much “safer” language to develop in, that is a class of runtime errors can be avoided and caught at compile time. This will be great, and I look forward to the day when it comes.
Apartment from that I think the type system power of Java/Kotlin in the near future is good enough and a nice blend of complexity and getting things done. And with Valhalla, Loom and modern GCs a lot of performance concerns should also be removed
https://wiki.openjdk.java.net/display/tsan/Main
But the types of races I struggle with in java I suspect are similarly hard in Rust.
Here's two pieces of concurrent code I struggled with:
https://github.com/google/guava/blob/master/guava/src/com/go...
https://github.com/google/guava/blob/master/guava/src/com/go...
I'm not sure reasoning about the correctness of those would be easier in Rust, but if there are any rust experts reading this I'd love your perspective.
Don't get me wrong - Kotlin is a very fine language with lot of promise and the comparison to Scala is not exactly fair - Scala is just a completely different league.
I really like what Odersky and the Scala community did with Scala 3; they lifted the language onto a sound new theoretical framework, streamlined the way implicits are used and got rid of puzzling irregularities. Did I mention the new sound macro system? But I'll stop fawning now.
I'm currently working on a project where we use Kotlin instead of Java. Probably wouldn't have taken the job if it was with Java and Kotlin makes it quite fun. But damn - I do miss Scala.
True "primitive" objects which are packed densely in memory and generics over such objects (to use ArrayList or equivalent instead of bare array[]) is main reason what allows .NET code to be faster than JVM code, though .Net compiler (JIT) and VM are much less optimized. But JVM (HotSpot) doesn't have choice now :-( Valhalla, when fully implemented, will gibe a huge and long-awaited performance boost.
All these hyped non-nullable Option<> maybe good, but is not a show-stopper. Performance is.
Edit: for grammar.
In traditional/non-FP code bases it's not a problem at all to check for null when a data structure uses it to signify "missing". There are usually not that many places where it's necessary.
Getting Java on par with C++/STL would be great though. No need for trove4j/colt. Anything that reduces memory usage without making people think hard would also be great to compete with golang.
Kotlin does have the !! operator for practical purposes like Java interop. Equivalent to this Scala 3 code:
extension (t: T)
def !!: T = if(t != null) t else throw NullPointerException()
I haven't checked after release, how is Scala 3 smart-casting now? E.g. Kotlin and Typescript convert a T? to a T inside an if(t != null) branch, and Scala's left much to be desired when I checked.Scala's and Typescript's union & intersection + smartcasting approach is much more general than Kotlin's T? of course, but Kotlin could introduce that backwards-compatible if they wanted.
> "I haven't checked after release, how is Scala 3 smart-casting now? E.g. Kotlin and Typescript convert a T? to a T inside an if(t != null) branch, and Scala's left much to be desired when I checked."
See: https://docs.scala-lang.org/scala3/reference/other-new-featu...It works perfectly well in a number of ways (IE, you can use "=="/"!=" null, use a "match")
https://scastie.scala-lang.org/s853XsLoQEKTVosebJHelA
Kotlin is still a better experience when interacting with Java because of inbuilt interop with Java types like "List" though, in my opinion.
In Scala you need to use ".asScala", ".asJava" etc.
Have you experimented with converting entire projects of Java to Kotlin?
As a trivial example people would rather use ScalaPB instead of officially supported Java library. When I worked on a Kotlin-based system people were adamant to use github.com/JetBrains/Exposed. Which had a couple pages of documentation and was missing features.
And then again there's politics. Big/successful companies always end up re-inventing PL wheels. The whole GOOG/ORCL mess gave GOOG incentive to fragment the Java ecosystem too.
I’ve converted several projects from Java to kotlin incrementally and it’s been very smooth.
I rather use the platform languages without extra tooling and wrapper libraries.
So, if `Optional` will continue to be a reference type, you should be able to use `Optional.val` to use the associated value type.
Optional will be retrofitted to a value class not a primitive class, you get flattening but only on stack (parameter type or return type) not on heap (field).
I don't really understand the end of the part3 (the VM model), so i may be wrong.
I don't know why the glossed over small values of Integer's sharing identity (effectively behaving like values) when discussing the complexity of these things.
From 5.1.7. Boxing Conversion[0]: "If the value p being boxed is true, false, a byte, a char in the range \u0000 to \u007f, or an int or short number between -128 and 127, then let r1 and r2 be the results of any two boxing conversions of p. It is always the case that r1 == r2."
[0] https://docs.oracle.com/javase/specs/jls/se7/html/jls-5.html...
I like this bit:
> These so-called primitive classes can combine the expressive power of classes with the runtime behavior of primitives. The slogan for Valhalla is:
Codes like a class, works like an int.
I can immediately get a sense of what I'm getting--very much like Go using a non-pointer receiver for a function.It is mentioned in the second part if I remember correctly regarding changes to identity-dependent operations (namely ==).
Basically, p = p with { .x = 3 } or something like that. But that will also planned to work for other kind of classes and will be a basis for advanced pattern matching as well.
Also, in case of a primitive class with no guarantee of non-tearing, in practice you will get the same performance as with a mutable struct, maybe even better as the JVM can better reason about it.
Given how well the language is defined, they could almost certainly compile Jv1 classes to Jv2 classes anyhow with minimal impact probably.
They could fix a couple other things along the way.
It took python what, over a decade to get people to switch from 2 to 3 didn’t it?
I’m just not sure you could ever pull it off. It seems way harder than that. Even with some sort of compatibility layer. I think you’d really end up in more of a perl 6 situation where it’s a different language that would eventually have to be renamed due to confusion.
A lot of isolated Java could remain V1 no problem.
Allow new starts, separate modules, and those that want to port a J2.
Python 3 was a disaster because they never bothered to think about it, provided no facility for it, didn't care about collisions at the platform level.
Oracle could literally pay for the most important OSS modules to be ported, give architectural guidance scripts to transform old codebases, special test units, training.
Python and Javascript are ironically two of the biggest messed in all of tech, that gain popularity and drag everything into the mud with them.
Unfortunately the more compatible you are with "V1" the more compromises you have to make. See: Kotlin, F#, TypeScript bending over backwards to maintain compatibility with the dominant language of their respective platforms. Not to mention C++. You have to admire the hutzpah of Python 3 level of changes. Some software should simply have a "sell by" date and be thrown out when it starts to stink.
That said I'm surprised more languages haven't attempted the approach of Rust with 'editions'.
> There are times, however, when it is useful to be able to make small changes to the language that are not backwards compatible. The most obvious example is introducing a new keyword... Editions are the mechanism we use to solve this problem. When we want to release a feature that would otherwise be backwards incompatible, we do so as part of a new Rust edition.
[https://doc.rust-lang.org/edition-guide/editions/index.html]
Two major initiatives for many years in Java have been Loom and Valhalla. A cynical take on that is that Java is attempting to copy something that Go and C# (and others) have had and excelled at for a decade or more, and were designed to do from the start. Sometimes, at some point the expense and complexity of dragging decades of code forward (with all accumulated mistakes, design assumptions made based on now obsolete hardware, backwards compat constraints, etc.) outweighs the cost of going back to the drawing board to start again.
(Also, do note that we still rely heavily on numerical libraries written in fortran)
State of the art GCs are incredible - but there is a reason you see Java pioneering the GC area, that being that Java has the most to gain because of its tendency to create enormous amounts of garbage, which is a problem other language designs do not suffer from nearly as much, letting them get away with simpler GC while maintaining high performance.
And Fortran isn't the only game in town for high performance numerical libraries ;) There is large incentive to reinvent there, as legacy doesn't matter if the new thing has more performance - Eigen, cuBLAS/CUDA, and others are in C++, not Fortran as far as I'm aware. Fortran has momentum no doubt, but it has been on its way out for a long time.
Also, due to the JVM’s flourishing multi-language support all these changes will greatly benefit all the other JVM languages as well. So in that view, I do think that the JVM has quite a future ahead of it because Valhalla will further improve the already killer performance.
For Kotlin and TypeScript, compatibility with the dominant language is the main reason for their success. For Typescript especially, the reason that it's so popular is how easy it is to migrate from JS to TS: JS is valid TS. I don't think the same is true for Kotlin. That combined with less "game changing features" in general (Java is already statically typed for example) probably explain why TS is way more popular in the JS world compared to Kotlin in the Java world.
> You have to admire the hutzpah of Python 3 level of changes.
You mean breaking backwards compatibility and dividing the ecosystem for relatively small changes? While I admire the self-confidence that it takes, I'm happy other languages are more reasonable in their approach. And the Python people seem to be more reasonable these days. Async was added without needing Python 4, and they're talking about performance improvements and multicore too. The rolling release probably helps a lot here, they can deprecate stuff slowly but surely.
> That said I'm surprised more languages haven't attempted the approach of Rust with 'editions'.
Editions can only affect some part of the language. From the Editions Guide [1]:
> The requirement for crate interoperability implies some limits on the kinds of changes that we can make in an edition. In general, changes that occur in an edition tend to be "skin deep". All Rust code, regardless of edition, is ultimately compiled to the same internal representation within the compiler.
Of course, a way to easily make skin deep changes is better than no way of making changes at all. But often, when a language changes, idioms do too and thus code has to be changed. For example, OCaml 5.0 will have effects for direct asynchronous IO. This will be backwards compatible, current monadic asynchronous code will still work, but people might want to rewrite their code in direct-style. C# introduced nullable reference types in C# 8.0. It's backwards compatible too, but you might want to update your code here too.
[1]: https://doc.rust-lang.org/edition-guide/editions/index.html
Rust editions are no different from selecting language versions, on their current state they require compiling from source with the same compiler, a very tiny set of language changes that might be supported.
> Each project can opt in to an edition other than the default 2015 edition. Editions can contain incompatible changes, such as including a new keyword that conflicts with identifiers in code. However, unless you opt in to those changes, your code will continue to compile even as you upgrade the Rust compiler version you use.
Now granted it’s entirely possible that a given compiler version accidentally broke comparability when building in a different mode - that’s a bug. It’s also possible that this model does have a limit as to how breaking of a change the editions can have (eg probably the ownership model can’t change drastically) which is maybe what you’re trying to say? Certainly if I recall correctly there have been breaking changes in editions.
[1] https://hkalbasi.github.io/mdbook-ra/appendix-05-editions.ht...
Other languages, like Java, C, C++, C#, F#…, also embrace binary libraries and multiple implementations.
Rust editions currently don't have an answer for such scenarios.
Which isn't really much of a problem given that the ecosystem is built around cargo compiling from source.
Aren't they? Can you choose to compile Java 17 with Java 7-8 syntax, but still make use of Java 17's other features which don't break backwards compatibility? As far as I have understood, that's possible with Rust (though I might be wrong, I've never really tried using an older edition).
One lesson for Borland days is that I rather use the languages that are on the box than dealing with extra complexity of third parties in the platform.
My point of view is not always what wins out, and there is good to learn new approaches in programming that can be latter applied to the daily tools, but that is how I see things.
The Borland example is confusing. The way I remember it their Pascal and C++ support was equally official. BC++3 was comparable to TP6 though the latter was less resource hungry (and at <700KB fit a single FD). Then for many years Delphi was well ahead of C++ Builder. But in the end professionally C++ skills were much more valuable.
Ten years ago I hoped that ORCL would acquire Typesafe and make Scala supported officially the way MSFT did with F#. Instead they left the space open for Kotlin to appear. And then antagonized GOOG into making it a juggernaut. There must be some Innovator's Dilemma-style explanation here but as a developer I regret it.
As a side-note I believe Martin made a huge mistake by allowing Haskell fans to highjack the narrative. His original strategy reminds me Elon's a little. Make industry finance and popularize high-impact type system research by using its impressive end result. Investing in Scala ten years ago prepared me well for Kotlin and modern Java (not to mention a few years of data engineering jobs around Spark and ML).
It's production tested in the largest and hardest java ecosystem out there - android apps.
The most unique Scala feature seems to be implicits. While I'm not convinced by their usefulness, the feature will be available in Kotlin in a few months: https://github.com/Kotlin/KEEP/issues/259 I'd even argue it is more useful in Kotlin since the language is state of the art regarding it's ability toward making DSLs.
The major features that Scala 3 has that Kotlin does not that I can identify are: 1) Pattern matching -> Kotlin is almost there, it has destructuring, a when control structure, and can express ADTs via sealed classes/interfaces. The limitations are: name based destructuring is not implemented. No destructuring in when blocks. However since java will be getting extensive pattern matching support in jdk 19/20 I expect Kotlin to catch up there and I recognize pattern matching can be a very useful feature.
2) union types and intersection types. While those are useful in a structurally typed language such as typescript, I don't see much usefulness for them in Kotlin. You can express union types via sealed classes somehow. The compiler plugin arrow meta brings union type support but I haven't tested it yet. As for intersection types it is on the roadmap. The main usefulness of union types I can see are the union of exceptions in catch blocks, that BTW Java supports, surprisingly.
3) higher kinded types I don't see much usefulness in higher kinder types What else could their use be beyond abstracting over collections? Every Kotlin collection implement the interface Collection<T> which enable this use case. Again, higher kinded types are supported by the arrow-meta compiler plugin.
To me the very few pain points of Kotlin have nothing to do with functional features but are: 1) Inability to deflare a function parameter as mutable. Inability to declare for loop element as mutable. Which in fact are because of IMHO religious functional toxic beliefs about immutability everywhere. 2) no package private modifier (although java 9 modules are supported and they are working on the issue)
I really don’t see the point of coroutines, especially when a) project loom is on the horizon, b) higher kinded types make it quite trivial to implement as a library (answering your 3rd point)
I’m not sure about reified generics, they are very seldom a problem and frankly, type erasure is most often just misunderstood. Also, with pattern matching overloading is used less often so it is even less of a win.
Regarding compile time safe recursion, I’m pretty sure that Scala was actually the first bigger language implementing that (with an annotation), so it definitely has that.
So don’t get me wrong, Kotlin is a great language with some nice features, I just don’t see it doing anything novel to be honest (and that’s not a bad thing! Non-research languages rarely do anything truly novel!). I just feel like Scala does all of those with less special cases (eg. extension functions can be done with implicits so you get one feature doing many things)
I am merely pointing out that Kotlin has a much larger community adoption. and it was never backed by corporates.
Ultimately - adoption matters.
https://spring.io/guides/tutorials/spring-boot-kotlin/
in fact, this data is fairly accurate - check out the stats on rise of Kotlin's popularity.
https://www.developernation.net/blog/infographic-programming...
1. No object value types; and
2. Boxing and primitives.
On (1), all objects in Java are references, being a pointer to a pointer effectively. This indirection has a cost but has been incredibly useful historically because it enables moving allocation around the address space because none of the references change. They're effectively just pointers to a lookup table and it's the lookup table that changes.
As for the example of packing an array of Points into a contiguous block of memory to avoid the dereferencing, I'll be curious to see what the proposal is here. Now, for example, if you have a collection of Points (x,y), you can put in a 3DPoint (x,y,z) that extends Point. You can't do that if you've flattened storage, not without some overhead at least.
For years it's been the dream in Java to have true value Object types. You can do this in C++, which also allows you to allocate onto the stack. My position is that the complexity cost of this is massive. Copy constructors, move constructors, assignment operators, implicit constructors, implicit casts, etc. I mean you can have class A with instance a1 on the stack and a2 on the heap and pass &a1 and &a2 to a A* parameter and the callee has absolutely no idea of the lifetime of those or where they're allocated.
But still, it would be nice to pass around an SHA1 hash as a 20 byte value instead of a reference to an object that contains a reference to an array that contains 20 byte elements.
As for (2), this was a deliberate design choice in Java 5 (1.5) when generics were added. The decision was made, for better or for ill, to retain backwards compatibility through type erasure such that a List<T> is a List. This decreased the pain of upgrading and legacy code but created this ugly corner of primitive types.
C# 2.0 came along later and decided to go the other way such an IList<T> is not an IList (IIRC the types; I'm not a C# guy).
There are pros and cons.
The fact that certain values are guaranteed to satisfy reference equality is kind of weird and (IMHO) confusing ie new Integer(127).equals(new Integer(127)) is true but new Integer(128).equals(new Integer(128)) isn't necessarily true. I guarantee you a lot of people don't know that.
There's only one level of indirection. References are simply pointers. You can't see the raw bits of the pointer without Unsafe, but that's how the Sun JVM implements references. (The implementation calls them "ordinary object pointers".) GC rewrites these pointers when it moves objects; there is no separate lookup table of all objects.
You might have been mislead by the existence of Object.identityHashCode(). The identity hash code is not a memory address and not guaranteed unique. It's an arbitrary value stored in some bits of the object's mark word and copied around when the object moves. That's how it remains stable across GCs.
Or you might be thinking of "compressed oops". Those are still pointers, but encoded as (object_address - start_of_heap) >> 3 to save space.
I think there is a global lookup table for interned strings and symbols loaded from class files, but slots in that table are not themselves pointed to by references.
If a program compiles with no unchecked or raw warnings, the synthetic casts inserted by the compiler will never fail.
Huh. That was written in 2020, four years after it was shown how to write a very small program that "compiles with no unchecked or raw warnings", and yet "the synthetic casts inserted by the compiler" will fail at run time [0]: class Unsound {
static class Constrain<A, B extends A> {}
static class Bind<A> {
<B extends A>
A upcast(Constrain<A, B> constrain, B b) {
return b;
}
}
static <T, U> U coerce(T t) {
Constrain<U, ? super T> constrain = null;
Bind<U> bind = new Bind<U>();
return bind.upcast(constrain, t);
}
public static void main(String[] args) {
String zero = Unsound.<Integer, String>coerce(0);
}
}
[0] N. Amin, R. Tate "Java and Scala's type systems are unsound: the existential crisis of null pointers", https://dl.acm.org/doi/pdf/10.1145/2983990.2984004I imagine C# was able to learn from what was going on with Java at the time and not have to suffer the same fate. Or maybe it was just less popular and could force the issue (I’ve never dealt with it so I don’t know the history well).
In Java, List<Cat> may or may not be the subclass of List<Animal>, but it is up to the List implementor. This way Scala/Kotlin/another JVM language is free to define their own variance model independent of the host language. C# did limit their language ecosystem with it quite a bit. (Afaik Scala for CLR stopped in part due to this).
https://docs.microsoft.com/en-us/dotnet/csharp/language-refe... https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
You probably know this, but I'm replying in the hope it clarifies things to other readers. The way it's expressed in Java is List<? extends Animal>
Also, unfortunately arrays are covariant so Cat[] is a subclass of Animal[] (both in Java and C# actually), where your mentioned example indeed introduces a “poisoned” value, waiting for a classcastexception.
What makes you think this is the case? Pointers to objects are direct, and garbage collectors freely change pointer values whenever it decides to relocate objects. Some collectors might use temporary forwarding pointers, to support concurrent collection.
This is not how Java works. I stopped reading after that.
I can see their reasoning to add value classes in addition to primitive classes, but the difference between them is going to create the most confusion, since they are identical in terms of requirements to classes themselves.
As for value classes - Valhalla is a reason (Scala's) value classes weren't removed from Scala 3 (in favor of opaque types), since they seem to be a good candidate to be implemented using Valhalla's classes.
std::vector<void*> ?
What the JVM ends up seeing is something like Java 1.4 bytecode would look like, and there is no need for allocations, because it is only references and generics don't support primitive types.
While C++ templates are monomorphic so a few more tweeks are indeed needed all the way down to machine code.
The principle of cleaning the underlying type applies though.