JEP draft: No longer require super() and this() to appear first in a constructor
openjdk.org
openjdk.org
It's easier to just say "this() must ALWAYS be called first in constructor" and everybody understands... than to try and say "you can call this() after other statements, but not in a try block and not if those statements reference the instance under construction" which will certainly allow you to do more interesting things, but also be more confusing.
Foo(long x) { ... }
Foo(double y) { this(Math.round(y)); }
In the second constructor, the first call is Math.round(), followed by Foo(int). I am talking about evaluation order, not lexical order.I found myself needing to work around the limitation recently: https://github.com/nayuki/PNG-library/blob/b92cedc23b1fc32d5...
In all other cases, either super(...) or this(...) MUST be called exactly once. Additionally, you can't catch Exceptions thrown by a call to super(...) or this(...) (so you can't call them inside a try{} block). Finally, you can't reference any non-static members of the "this" object except to write to write to a field.
So yes, the rule is significantly more complex.
Bison from Buffalo coerce other Bison from Buffalo.
Generally, making a language more expressive makes it harder to parse. You have to know a lot more stuff. Computer languages often have a formal spec. I think java does, but I don't think perl does. Regardless, you just have to remember that footnote on page 26 of the manual that the construct X is an option.
I haven't touched java in a while, but I do remember some difficulty with parent class constructors calling methods implemented in the children, that relied on member variables that hadn't been initialized yet. I wish I could point at a concrete example, but I think it was something deep in the heart of swing (part of the jre so not changeable).
I can't imagine a way this makes it easier to parse in my head as I'm reading through code, or communicate failure. but there probably are libraries that become more usable with this new syntax.
It sucks because you're guaranteed super is called, the parent class can implement stuff that _has_ to happen. there's no way for me to force a subclass to call, something like pleaseInitializeAfterSuper(). and there's no way for me to do anything before super, right now.
Regardless, there's no easy way out. Making the language more expressive is a damn big hammer to use though. you'd think you could use a builder to guarantee initialization, But sometimes you can't.
To answer your question, no. I don't think this makes it easy to communicate why a program was rejected. it makes more programs possible though. At the cost of having to remember that footnote on page 26.
> Bison from Buffalo coerce other Bison from Buffalo.
The words don't have different meanings, though, at least in idiomatic English. No-one uses the word "buffalo" to mean "coerce". I'd be surprised if you could find even a handful of examples of it being used that way.
File the "many buffalos" thing under the same heading as the "code golf / obfuscated C" type things where you've got a chain of half a dozen ternary operators wrapped in a loop to make a state machine, where a switch() statement would suffice.
Vaguely related: www.marchmarchmarch.org.uk
(a badger being someone who makes badges, and thus badger badger being someone who makes badger badges)
Also, if you can have a badge depicting a badger badger, then badger badger badgers badger badger badger badgers.
PS. Mushroom! MUSHROOM! Panic, a snake!
I am not sure what your definition of "expressive" is. To me, effectiveness to convey thought or ideas in this context. And I don't have the faintest idea of what that sentence means.
Still, after I experienced the constructor-less paradigm in Rust, where you only can only construct structs entirely all at once, and every "constructor" is just a static factory function, it's harder to justify the traditional OOP paradigm in Java, C++, etc.
Moreover, given that Java calls overridden methods when super() calls a non-final method, the existing rules prevent some reasonable code patterns (again, see the proposal for the example) and the new ones allows to address those in reasonable and readable ways.
What's pushing for more possible correct programs?
- super() must be called first: the rule is simple but restrictive - non herited elements can be used before calling super(): the rule is more complex but less restrictive. This can lead to more complex error messages, etc
Technically, this does introduce new functionality when extending a class, specifically it's possible to choose a constructor to overload depending on arguments, but in most cases there is usually one constructor that all other constructors use.
C# was so much more joy to program in for a long time, even though they were equivalent languages in safety/speed tradeoff.
Having said that, from the link posted by carimura [0] it looks like the author of this one is doing an excellent job of engaging with Brian Goetz and others at oracle.
[1] https://mail.openjdk.org/pipermail/amber-dev/2023-January/th...
[2] https://github.com/openjdk/jdk/pulls?q=is%3Apr+is%3Aopen+arc...
public A(String s) {
this(s.isEmpty());
}
is allowed but public A(String s) {
boolean helpfulName = s.isEmpty();
this(helpfulName);
}
isn'tYes, in your particular case, the compiler could figure out that you're not using "this," but you can only make a compiler so clever.
public A(String s) {
this(helpfulName(s));
}
with a static method.Likewise for "super", because Java allows subclasses to modify parent fields, which need to be initialized to fulfill the same contract.
Swift, which has the same contract, splits the "initialization" phase from the "construction" phase, so you can indeed modify internal state prior to calling "this()" or "super()".
This contract doesn't hold anyway, because you can call object's methods from the constructor before the object initialization finishes, and those methods can even call the overriding code in the derived classes. So even though `super` is guaranteed to be initialized fully, `this` is not.
It is complicated and I'm sure it has caused some bugs, but it has also solved some bugs.
While the idea is reasonable on its face, it's convoluted enough in practice that it used to be a common source of interview gotchas, so I'm not sure that it fixes more problems than it creates in practice.
It was a painful learning process from there to where we are now, where we consider any internal state that isn't passed to the constructor in an already immutable shape a considerable smell that you'd better have very good reasons for and we prefer a clear distinction of preparation and operation phases separated by that one constructor which sets (and in most post-OOP languages also declares) the values of the object's constituents.
(PS why isn't "post-OOP" a far more established term? It feels considerably more useful to me than the usual mind gymnastics about "functional but not quite" that provoke all those (im)purity squabbles that help noone)
Figuring out what to pass a constructor and the actual allocation are different tasks. The anti pattern is conflating them. It's fine to use constructors.
Being able to put an object into a specific well defined state at the time of instantiation is really good.
From syntax perspective, the parent and member constructors (in C++) are in a special list outside function body, it doesn't tease you with it looking like a regular statement inside the constructor body.
C++ not only has constructors, but default constructors and destructors. If an exception gets thrown, the compiler needs to ensure that destructors will get called only for the parent classes and members that were initialized. It does so by having very specific order in which the parent and member constructors and destructors get executed. The use RAII mechanism also means that C++ developers are more likely to make classes where construction/destruction has sideffects which need to happen in specific order.
Unlike Java where due to all the classes are reference types semantics uninitialized members can be left pointing to null, in c++ any nontrivial member will have the constructor called no matter what. If you don't do it explicitly the default constructor will be called. Implicitly or explicitly in both cases the member constructors will be called in the definition order !BEFORE! the constructor body, but after the parent constructor call.
As if things weren't complicated enough, C++ not only supports multiple inheritance it supports the diamond case in both ways: with the shared base class duplicated, and in case of virtual inheritance with shared base class not being duplicated. With all that complexity compiler still needs to ensure that right subset of destructors get executed in case of exception during construction.
Some of these problems Java solves with the help of garbage collector. If you get some members laying around after failed construction in Java, sooner or later it will be cleaned up by GC, and since Java doesn't have destructors it isn't critical when exactly it happens.
The factories will call the constructor, they kind of look similar
The constructors could still be public, so if you don't have the need to do anything fancy, you might still just call one of the simple constructors directly. However, you often see that API designers restrict this approach in favor to a single point of entry into creating objects of that class. That is, you then have to create objects via a factory, even if you could as easily pass already existing values to a simple constructor. But that's a different story.
What's next? Multiple inheritance?!
I'm gonna need a better hypothetical than FilteredSet<E> to assess any benefits. Today, people use factory methods to achieve that result.
Or just use a compound statement. Something like (pseudo-codimously):
class FilteredSet <E> {
public FilteredSet( Predicate<? super E> pred, Collection<? extends E> elems ) {
super( elems.apply( pred ));
this.pred = pred;
}
}
My hunch (believe but cannot prove) is that Java's current constraint serves to encourage composition over inheritance. Which is The Correct Answer™.My limited anecdata wrt Java's "constructors must invoke super()/this() first" is this complaint arises when porting from C++ to Java. Because the original class hierarchy leans heavily on inheritance for organizing behaviors.
This just isn’t true. There are compatible changes and incompatible changes you can make.
Having the soundness of the compiler rely on the assumption that there won’t be any changes just isn’t viable. For example, the thread-safety guarantees provided by final fields would be difficult to ensure.
This statement exactly applies to the case you mentioned: Superclass implementations can change later (you can recompile the superclass without having to recompile the subclass), you can’t assume their current implementation will stay the same.
If you are using "in general" in the sense of "in the general case", you are right. But if you are using it as "typically", then that's not true - it's very common for example for a library to be compiled with Java X while subclassing at least a few standard library classes, and then to keep using the same compiled classes with Java X+1, X+2, and to expect them to just work.
The Java Standard Library actually takes quite a bit of care to make sure modifications to its (non-final) classes are backwards compatible even for subclasses.
Just as you get a runtime exception if you invoke a missing method, you could get a runtime exception for accessing uninitialized state.
Calling overridable methods from a constructor is just a bad idea in general.
Making it illegal is difficult, because there is a lot of code relying on this, and there are valid use cases for passing this to foreign code from the constructor, foreign code which might of course invoke anything on the instance.
As an example of where I have seen this used, your File object could open the file and truncate it in the constructor but this is non-obvious, it would be better to have another method like Init() or perhaps a factory method where calling e.g. File.OpenForWriting() is obvious.
I know not everyone has the same requirements but I can't remember a single time where this restriction was a problem for me in C#, the feature sounds like it will encourage worse behaviour.
Opening a file and truncating it sounds totally unexpected and I never do such things like touching the filesystem or running queries in a constructor.
And on the subject of Java restrictions that get in the way: if you use a local variable in an anonymous inner class or lambda, it needs to be final. This presents a problem for when you call a method that runs your lambda at some point before it returns and you want to modify some variables from there. The ugly workaround is to declare a final single-element array. It satisfies "needs to be final", but its single element can still be assigned to from anywhere since "final" only applies to the object reference to the array itself.
This doesn’t work if it requires multiple arguments to be generated with common code, unless you’re okay with effectively running that common code twice, which in turn is problematic when it has side-effects (like logging). See https://news.ycombinator.com/item?id=34482908 for an alternative way of solving this.
If the former, it's very space-inefficient if you're capturing lots of stuff, as is typical with callback-oriented programming.
If the latter, then a lambda referring to a single variable from some stack frame will keep all other captured variables in that frame alive. Which can be especially surprising if the lambda that is alive captured a single int, and that causes a huge array that referenced by a different local that can no longer be reached to remain allocated.
FWIW C# does the latter, and it works "good enough" in practice. But I've seen more than one mysterious memory leak caused by it, too - I think most of the time they just aren't big enough for people to notice.
Another examples of Java pain points solved in Kotlin which I could remember:
top-level functions (which are actually bundled in some "standard" class name. Pretty much syntax sugar but very nice.
Auto cast. Like you check variable with instanceof and after this check it's automatically casted to the target type. Java has some weird syntax for this but I don't see why this was necessary at all.
More default imports. Why should I qualify java.util. it's pretty much standard. I'm not asking for every java.x package to be auto-imported, but commonly used packages should be.
Class properties instead of getter/setter ceremony. Records are alternative approach but I don't find it useful yet.
Nullability. Of course that's a hard question and adding nullability to existing type system looks like gargantuan task. What could be done is introduce some java.xxx annotation, let module projects to be not-null by default (so you don't need to create package-info.java in every package just to make it not-null) and make all nullability issues to be warnings from javac. Basically it's already exists but not standard and implementation quality varies.
Reified generics. Basically it's a hidden Class<T> clazz argument passed automatically to the function. This is brilliant solution to type erasure problem and it should really have been adopted.
I'd vote for all those features to be implemented in Java similarly to their Kotlin implementations.
public FilteredSet(Predicate<? super E> filter, Collection<? extends E> elems) {
super(checkElems(elems).stream()
.filter(filter)
.collect(Collectors.toList());
this.filter = filter;
}
I think you misunderstood the problem. Assuming you leave the add() method implemented exactly as in the JEP, this will still fail with the same NullPointerException, since super() will still call FilteredSet.add() before this.filter is provided.The only way to avoid it without this JEP is to do something like this:
public FilteredSet(Predicate<? super E> filter, Collection<? extends E> elems) {
super(elems.size()); //initialize super as an empty HashSet, but with the proper capacity
this.filter = filter;
this.addAll(elems); //this will call this.add(), but this.filter will now be set
}
Your description of the contract/purpose of HashSet.add() is also just wrong. The clear intention of the Java standard library designers is that any element added to the set, either by the constructor or later, is added using HashSet.add() - they are explicitly calling that method from the built-in HashSet constructor [0]. Also, it's an obvious design choice, while your proposal is just weird: why would you ever want to duplicate the implementation of adding elements between initialization and later modification?In general, you seem to be arguing for keeping the language more clunky. Why does `7 instanceof int` or `switch(number)` or avoiding a new method definition just to check a variable before calling a constructor hurt the identity of Java in any way? Sure, all of these things could be achieved in various ways before as well, but always in a more clunky way. Many of these JEPs are pure cleanup of bizarre idiosyncrasies of Java - the precise opposite of how C++ is evolving.
[0] https://github.com/openjdk/shenandoah/blob/master/src/java.b... - note that the doc is not very explicit, but it does say that `c` is "the collection whose elements are to be placed into this set", not something like "the collection whose elements will be used to initialize this set".
OOP classes are really a bit of a triumph of engineering, because they offer a lot of nice ways to build out various complex systems (would rather build out GUIs in a Java-based object model than in a Rust-basd one). And the fact that moving the model "forward" in expressivity is so subtle and starts introducing a lot of asterisks is probably proof that it's some local maximum
Java becomes a slightly harder language to learn, because now there is one more thing to explain to a beginner: "Where exactly should super() be put in the constructor? At the start, in the middle, or at the end? What are the benefits and downsides of each, and when does it matter? Is there a convention?"
I think python is a classic example of this - python 2 was pretty simple. Python 3 just gets more and more complex.
Edit- from reading https://stackoverflow.com/questions/25007427/how-are-excepti... it appears my suspicion was correct. try{ is effectively an annotation that has no impact on the sequence of instructions you are authoring. So it's surprising to me that it breaks the this/super first rule.
OTOH it seems to me there should always be a rule to not allow execution paths that never call super nor this, and since try could break that one I'd always be leery of using it in constructors.
> try { } Blocks
> The restriction that super() and this() may not appear inside a try { } block comes from the JVM itself, and is due to how StackMaps are represented. The logic is that when a superclass constructor throws an exception, the new instance on the stack is neither fully uninitialized nor fully initialized, so it should be considered unusable, and therefore such a constructor must never return. However, the JVM doesn't allow the bytecode to discard the unusable instance and throw another exception; instead, it doesn't allow it to exist on the stack at all. The net effect is that constructors can't catch exceptions thrown by superclass initialization, even if rethrown. [emphasis mine]
the parent class however now has to deal with an arbitrary set of operation that the sub class can perform between the initializer block and the constructor method, so it's contract is overall much harder to define, because you can no longer be sure of your internal state in the constructor.
Every expression that you call in super was called before it, obviously.
So you always was writing code like
MyCtor() {
int arg1 = calcArg1();
int arg2 = calcArg2();
super(arg1, arg2);
}
And the fact that you wasn't able to write this code literally was confusing for every beginner.Right now the rule is simple as it should be: you can't read `this`, explicitly or implicitly, before calling `super()`. Now it actually makes sense.
The only somewhat confusing thing about this JEP is that you can initialize fields before calling super. I don't see why this is needed and I'd treat it as code smell. I've read that example and I don't think that it's convincing. This code should be rewritten.
class Base {
public int Count { get; set; } = 0;
public Base(int counter) {
Count = counter;
}
}
class Sub : Base {
public Sub(int number) : base(number) {
System.Console.WriteLine($"This is the sub class constructor with {number} as an argument");
}
public Sub() : this(123) {
System.Console.WriteLine("This is the sub class constructor with no arguments");
}
}
Rather than being a magical function call at a magical position, the call to the super constructor is part of a special bit of syntax that explicitly tells you what kind of call is being done. You can't call any method at that point of the program, you have to call a constructor by its special reference.It also makes the reason why you can do `super(fn1(), fn2())` but not extract those two calls to variables; these calls are considered to be part of a different part of the execution flow than the rest of the constructor. A simple `base()` can be left out, of course; it only matters in cases where you need a specific super constructor.
I would love something akin to this JEP in C#, but I'd want it in a way that makes it very clear that you're executing code prior to the base class' fields being initialized.
Something along this lines of this
public Foo(int bar)
{
prector
{
var baz = bar * bar;
base(baz);
}
ctor
{
// Foo's constructor code
}
}I love null being apart of the type system, but the safe call operator seems like a huge step backwards.
Also I don't see the difference between if(a != null) { a.b(); } and a?.b().
They won't result in NPEs either way and I see the former everywhere in Java as well.
Void func doThing(state : State?) {
var idx = state?.idx ?? -1
//No more ?. Needed
}However, Java has instead received `Optional` (in the shape of `Optional.ofNullable(variable).map(Clazz::getProperty).map(Clazz2::getProperty).orElseGet( () -> 42 );` rather than `variable?.getProperty()?.getProperty() ?: 42`).
I don't know what the JCP EC was thinking going for this approach rather than the much more readable, modern approach, but I can't be too surprised about a language designed by committee refusing such improvements.
There are comments in this very discussion praising this change due to how this stuff really improves their life as a developer.
Anyway, can you pass the link to your JEP with a proposal for "?." ?
Maybe Manifold will add the operator somehow. That seems to be the place modern language features that the Java comities don't like end up.
final Foo foo;
try
{
foo = bar(baz);
}
catch (SomeException ex)
{
foo = SOME_DEFAULT;
}
For some reason, the assignment in the catch clause is seen as a reassignment ("variable foo might already have been assigned"). As a workaround, one can factor it out into a separate method and have “return” there instead of “foo =“. This is particularly inconvenient in constructors when you want to assign to a final field. final Foo foo;
try {
foo = bar(baz);
otherFunction();
} catch (SomeException ex) {
foo = SOME_DEFAULT;
}
that'd truly be invalid - the first assignment could run, otherFunction errors, and foo gets reassigned. Even if SomeException is a checked exception that otherFunction doesn't declare in its throws, it's still possible to get the jvm to throw it, and things do rely on that - see Lombok @SneakyThrows.So the only case that such final field assignment can work is if it's the last evaluated expression in the try case (plus some guaranteed non-throwing arithmetic or boring assignments), which is quite the weird thing to have as specified behavior in a language.
Meaning, when needed you could still always do:
final Foo foo;
try {
Foo tmpFoo = bar(baz);
otherFunction();
foo = tmpFoo;
} catch (SomeException ex) {
foo = SOME_DEFAULT;
}
But at present the "best" (besides introducing a separate method) you can do is final Foo foo;
Foo tmpFoo;
try {
tmpFoo = bar(baz);
otherFunction();
} catch (SomeException ex) {
tmpFoo = SOME_DEFAULT;
}
foo = tmpFoo;
Which has the drawbacks that (a) the compiler doesn't check that tmpFoo is assigned exactly once (because you can't make it final), and (b) tmpFoo stays around for the rest of the scope, with the risk of being misused later.I would understand if an excemption is being made for Errors, because in princible any JVM bytecode instruction could probably throw one, but not for regular Exceptions.
final Foo foo;
try {
foo = bar(baz);
Integer a = null;
int b = a;
} catch ...
(though that's less the assignment, more implicit behavior; point being, you've got quite the restrictive environment after "foo =")But accepting original example should be do-able with pretty simple AST pattern matching.
Were the original to be allowed, it be the first case where adding a System.out.println(); statement into code {}s results in an error. (ok, second case, the first being adding it before a super()/this() in a constructor. But hey look what post we're on!)
import java.io.IOException;
class Test {
int g() throws IOException {
throw new RuntimeException("aaa");
}
int f() throws IOException {
int x;
try {
x = g();
} catch (IOException e) {
x = 3;
}
return x;
}
}
So the compiler definitely believes that x is initialized because we can return it.Not a Java expert so maybe there's some subtlety I'm missing.
I'd really like to be able to define a constructor on sealed interfaces:
sealed interface Path permits WindowsPath, UnixPath {
public Path(String p) {
if (isWindows()) {
return new WindowsPath(p);
}
return new UnixPath(p);
}
}
That way user code to construct an object is always the same: var p = new Path(s);
Instead of: var p = Paths.get(s);1. They don't have a name. That's bad because names are important. It's OK if constructor is primitive and does not do anything but field assignment. But often constructors do something and naming might help.
2. There can't be two constructors with identical types (as a consequence that constructors don't have a name). This might be restricting in some cases.
Factory methods solve both issues.
If anything, I'd prefer to remove constructors completely and use factory methods for anything. Not realistic with Java, of course. But in my code I'm trying to follow this pattern. Constructors are absolutely minimal (usually just initializers for final fields) and any non-trivial code goes into factory method. Of course it causes issues with subclassing, so it's not a silver bullet.
Where perhaps the first wraps, the second copies and the third may wrap and avoid sorting the input. This would all of course depend on some conventions on meaning within the codebase.
Another benefit is being able to name different constructors differently.
If I were to design a new OO language, I would make it so that constructors always work like that in the first place, from the caller perspective.
Is this not what I'm describing? Using one way to construct objects would allow you to change from a concrete class to an interface without breaking client code. I just want the syntax to be the same across interface/abstract/concrete so class construction is uniform and discoverable.
It also doesn’t provide the flexibility in naming.
// you can start with a final concrete class
final class A {}
var a = new A();
// and later move to interfaces or abstract types
sealed interface A permits B, C {
public A() {
...
}
}
var a = new A();
Client code would remain the same with the flexibility to refactor to whatever you want later.Personally I still prefer having the ability to name the constructors. I'd rather have "new" be the default constructor name, with the option of having it be something else (i.e. Path.new(...) vs. Path.somethingElse(...)).
Constructor(A a, B b) { … }
Constructor(C c, D d)
{
// A and B can be derived from C and D, and
// we want to forward to the (A, B) constructor,
// so we use an intermediate object to do the
// conversion and provide A and B:
this(new Intermediate(c, d));
}
private Constructor(Intermediate x)
{
this(x.a, x.b);
}
Of course, with the JEP this will become much simpler.