Nullness markers to enable flattening
mail.openjdk.org
mail.openjdk.org
Well anyway, I diverged from that significantly now; Virgil has value types, tuples, ADTs, and closures and happily unboxes them all.
There's a bit of unnecessary overhead in having to externally declare and import the return type.
public [K key, V value] getElement(int index);
This really comes down to the dynamic vs typed argument, and to me the answer is "minimise the boiler-plate, maximise the guarantees.".But I wouldn't expect that to be the common use case - I'd generally expect that the caller would just be implicitly importing and using the type effectively declared in the method header.
var p: (int, int) = (0, 0);
var q = p.0 + p.1;
type Point(x: int, y: int) #unboxed { }
var p = Point(0, 0);
var q = p.x + p.y;
Both will generate identical machine code with no allocations. It requires no escape analysis or assumption of non-identity (both tuples and data types have structural equality).Right until there’s no more context than “this returns 2 elements” and you end up with 15 different Pair types and more verbosity for no value.
Named tuples are not better than tuples, they’re complementary, both are useful tools.
Not usefully so. If you’re splitting a collection in two, whatever context you add is almost certainly worthless overhead not worth the extra type.
By now I also have my doubts that it will ever come, they should have had better inspiration from Modula-3/Eiffel/Oberon when Java was originally designed.
[1] The two words even become separate, independent scalar values in the compiler IR. Thus the runtime support is minimal, in that it doesn't know anything about multiple-word values: the GC just needs to know whether a word is a reference or not.
That would be very helpful indeed.
Even Limbo had better support for plugins.
I think I've been corrected in the past, where someone pointed out that managed languages don't have to be like that? What are existing approaches that allow for better control than traditional Java offers while retaining a convenient syntax?
Concerning the languages I know -- memory managed, reference only models always come with the possibility of indirection at every turn due to not having the distinction between dereference and member access (not in the syntax and not in the semantics). It's very hard to infer from the code if a dot operation is an injection, mathematically speaking. I believe that is the reason why I've always felt "safer" and more precise in memory unsafe languages.
I also think that some of the newer languages are misguided to hide the distinction by using dot syntax for both.
Most managed languages don’t expose the location of an object, only its semantics. With optimizations (escape analysis), an object can be allocated on the stack and thus only effectively make member accesses. This optimization can be done much more deterministically with value types — if an object looses identity then the runtime/compile time doesn’t have to reason about whether it leaves the function/loop, it is free to copy it in/out of functions as often as needed.
As for preferences, it makes sense in low-level languages, but imo hiding it is the correct choice on a higher level — this is a core benefit of GCs, e.g. your public APIs will be much more maintainable/won’t need refactors because they don’t expose low-level implementation details (where should that argument be stored), only its semantics.
Trying to make this a not-type is doomed to failure. It's a property of expressions, it has rules about how it flows through expressions; that means it's a type, or it's an ad-hoc informally specified implementation of half a type. It's recapitulating the mistake of checked exceptions all over again.
As for checked exceptions, I disagree that is a mistake. They correspond one-to-one with Result types, but are just better integrated in the platform (has proper stack traces, bubble up as default, and auto-unwrap). Java’s implementation has much to be desired (it didn’t have sum types back then), but checked exceptions themselves should get a reevaluation.
People in the design mailing list discussion argued that those methods shouldn't be added because they didn't want Optional to be a monad.
> They correspond one-to-one with Result types
They don't, because throwing them is not a value (there is no way to have a variable that you can substitute for "throw e" and get the same behaviour), and because they're not properly integrated with the language's type system. Throws-ness is a separate parallel universe that behaves almost but not quite the same as expression type. Which is exactly what this is proposing to do with nullness.
While I do like more pure FP languages very much, I don’t think there is all that much value in everything being a value besides some mathematical elegance. Besides, a throwing exception can be turned into a value via a try-catch block, and can be just as easily tested as if it were a value, so where is the downside?
Exceptions are for exceptional situations, result types can still be used for very much expected ones. E.g. a primitive file read operation returning either a byte or EOF is great as a result type. Some specific IOException is better modeled as an exception which you might not be interested in handling at that level.
With all that said, I’m interested in the newer generation of languages with effect types that might bridge this gap.
It's really cumbersome in practice. E.g. when processing results from a database query (say), you probably want to call some callback for every row and then close the transaction and return the result of that callback. But handling exceptions around that becomes a significant pain. Once you've used a language with proper result types you don't want to go back, IME.
void transactionBoundary() {
var list = someLongListOfElems;
var resultList = List.of(); // a sum type for success and error conditions
for (var l : list) {
try {
resultList.add(process(l));
} catch (SpecificException e) {
// add to resulList or errorList or whatever to further process those instances.
// You can just use SpecificException’s fields
…
}
}
}List<String!> will still have nullable methods and fields.
This isn't the first time I've heard this and I'm still not certain I understand.
To me, Optionals are an alternative way to say "no value" over the existing way - null. I like Java's Optional because it means I can stop worrying about null. An Optional that wraps a null doesn't give me that. I still need to do null checks!
Why would I want that?
The other advantage of Optional is that it behaves consistently regardless of what's inside it, which means you can use it in generic code in a way that you can't use null. For example a common problem with null in Java is that if you call map.get(key) and the result is null, you can't tell whether that means the key isn't present in the map, or the key is present but the value is null. Whereas if you're using optionals, you can distinguish between None (key not present) and Some(None) (key is present, but mapped to None).
So Java could have added a getOptional method to Map (with a default impl), and then if getOptional(key) returns None the key isn't present, if it returns Some(null) then the key is present but mapped to null. That would be a legitimate, immediate use case, that would give people a benefit from switching to Optional today, and would start the ball rolling on using optionals, and maybe eventually in a few decades they would be widespread enough that we could start deprecating null in the language.
But instead they made it so you can't put null in Optional from day 1. Which would be great if there was a way to instantly migrate everything, but it makes it harder to start using Optional in an existing codebase and reduces the short-term benefits, for the sake of a future that (because of that very lack of short-term benefits) will never come.
There are already lots of techniques, from annotations to Optional.ofNullable but none matches the simplicity of != null.
They can also go the route of `effectively final` and make the compiler deduce nullness instead of marking it explicitly.
Although perhaps they could resolve this with magic imports?
import java.lang.magic.optional;
public class Foo{
public static void main(String...args){
Foo? f = ...;
if(f){
f.doThings();
}
f.doThings(); //compile time error
}
}
The idea being that importing the magic class enables the cleaner syntax.It would be like a variant of inviting Dracula across the doorposts, only that this one may or may not actually be a cool guy and not a vampire, but in any case he will sure bring all his friends, and their friends and so on forever.
So `Point p = null` won’t type check, you would have to mark it as `Point?` where Point is a value (primitive) class. There were also points that it should work with the usual type inference, e.g. `final var p = methodThatReturnsNullablePoint()` would be inferred to `Point?`.
String! myString = “Definitely not null”;
doesn’t look too bad.Edit: it’s worth saying that a lot of careful thought around compatibility goes into features like this. That means the core libraries should be able to transition to non-billable types where this is already a run time requirement, and existing compiled jars must continue to work.
So now we are back at warnings if non nullable references are enabled on the compiler settings, and eventually we can turn them into warnings as errors.
The biggest issue is the large ecosystem of existing libraries.
`MyClass! myvar` is not supposed to accept nulls. Assigning null to `myvar` would throw exception, comparing `myvar` with null is pointless and probably should emit a warning. Or maybe even error.
`MyClass? myvar` is supposed to accept nulls. Dereferencing this value without checking for null is either warning or error.
`MyClass myvar` is supposed to be unspecified with regards to nulls. So basically current code. You can do whatever you want but you should refactor it to `MyClass!` or `MyClass?` some day.
Annotations are, unfortunately, all over the board. There's a dozen different `NotNull` annotations out there. Having it in the language will make it the last solution to this problem.
You can't do deduction at compile time because you'll end up with different method signatures if someone calls a method with a null vs not. (not to mention the problem with someone calling these methods with reflection)
You can do deduction at runtime (and they likely will), but then you have to deal with the problem of unwinding an optimization when the hypothesis is proven wrong. That can be really costly when the optimization is dealing with memory layout. (Imagine an array of value types where you deduced they should be non-null, yet some unexpected path inserts a null. In that case you have to reallocate the entire array and it's elements in order to handle that).
It's more than just syntactic sugar.
In retrospect it would be very nice to have a language level ability to define something as optional, and require existence checking for such values;
But it wasn't so obvious at the time.
Personally, I like the Rust approach. It seems a bit more flexible and can evolve easier over time.
What is baz? Is it Option<Bar>? Is it Result<Bar>? Is it some user defined enum<Bar>? Who knows! You have to find that by looking up the foo definition.
For kotlin, the answer is simple. "baz is a nullable Bar"
Rust did this because interacting with the enum directly was cumbersome.
The concept of "no value" is so integral to day to day programming that elevating it into the type system with "nullable types" makes more sense to me vs using generics trickery. Even if it's slightly less "pure".
I do use it a lot with Result though. But if you have any question about the return type you could[1] just look at the return type for the current function you're in.
[1] I think there's a way to define custom types on nightly, and I wouldn't be surprised if it let you map custom types on Result/Option, but I don't actually know.
Syntax sugar that works with a normal library type is much nicer than dedicated syntax for a special-case builtin, IME.
> What is baz? Is it Option<Bar>? Is it Result<Bar>? Is it some user defined enum<Bar>? Who knows! You have to find that by looking up the foo definition.
Sure, or if you have a decent IDE you just mouseover it. But that's no different than any other method. `let baz = foo.add(bar)` doesn't tell you what type foo or bar is, and I don't think I've ever seen anyone argue that it should (e.g. by requiring method names to be globally unique).
> The concept of "no value" is so integral to day to day programming that elevating it into the type system with "nullable types" makes more sense to me vs using generics trickery. Even if it's slightly less "pure".
I've found this isn't really true. Once you don't have language-level support nudging you to use it all the time, wanting to have a possibly-absent value is actually pretty rare. (E.g. a lot of the time you want to include a "reason" for why it's absent, so you want an Either/Result-like type - but if you're using Kotlin you end up using a nullable type because you're lazy and the language makes that easier. That's bad for long-term maintainability IME, especially because if you want to switch the nullable type for a result you have to change all your code - unlike Rust where you can switch fairly easily because ?. works the same way for both types).
So I think we largely agree.
(I left example syntax in another comment.)
Clojure has a powerful meta-data system which could be used for checking for errors if Clojure didn't depend on the host platform for that. Also, it seems, Common Lisp also has some ideas about how to handle things more or less in-line with other code without having a special system that you can't see most of the time.
Btw. some of this can go very deep in the technology stack, there were computers (https://en.wikipedia.org/wiki/Ternary_computer) with ternary logic (https://en.wikipedia.org/wiki/Three-valued_logic). Such a computer could improve many aspects of computing, e.g. density and therefore efficiency. It could improve reliability (things could be more explicit). You could divide by 3 precisely and efficiently. It could improve our understanding of logic by making this for most of us alien concept to a more practical and widespread tool.
Most of the time this isn't a very useful type to have, but it has the big advantage of being mechanically obvious. For example, consider a hashmap of optional types, something like `Map<String, Option<Player>>`. If we write a get method for the map, it should return an Option to indicate whether the value was present or not. But what happens if the value is present, but it is explicitly Null?
The advantage of having an explicit option type is basically that it's easier to compose generics without having to understand what the values might be, which is usually what you want when using generics. That said, most of the time, if you've got an option of options of something, you're just going to flatten that type down anyway.
E: Thinking about it, the other advantage is that it's easier to create a separate namespace for "methods that should exist when T might be null but are meaningless the rest of the time". For example, Rust's `Option` type has methods to map the internal value into another value, unwrap the value and panic if it was null, swap the value with a different one, etc. In languages which use null | T, the equivalent is usually to use Elvis operators and similar (obj?.field), but that requires more special casing.
Also, the Map::get method that can return null is the problem here (even though your example is great and thanks for that, I didn’t think of it), something like ‘Option<Player?>’ should allow to differentiate between those cases.
But I do remember reading that a bottom type does make typing rules harder (e.g. scala 3’s explicit nulls feature which is basically “T is non-nullable, write it as T | Null” is not sound)
var foo: (T | None) | None = None
Which None do you mean here?
This is just basic logic: (X or Y) or Y <=> X or Y
Untagged unions cannot model that. Contrast it with tagged unions:
var foo: Option<Option<T>> = Some(None)
var bar: Option<Option<T>> = None
Now you can reason about where the None means.
Surely, if you want to distinguish between 3 states you need more data, hence my recommendation of Option<T?>
What you say about `Option<Player?>` is a good point though. If we want to distinguish between the different results, we need an explicit Option type. But now we've got Option _and_ we've got nullable types. Which should we use? They're both doing the same thing (i.e. marking where a type may be present but might not be), so why do we need both?
In practice, my impression that nullable types are really good for integrating with languages that already have unchecked nulls in then, either for historical reasons (like Java) or because they're dynamic languages (like Python or Javascript). It's a way of acknowledging the null value in the type system without demanding that all the code that been interacting with nulls be rewritten.
However, if you were going to write your own language from scratch, it's difficult to see why you would allow nullables to exist when the Option type does pretty much everything that nullables can, but with more clarity for cases of "nested nullability". You can also still add syntax sugar for it (Rust is going down this route, for example), but you don't have to special case nullability to the same extent.
They certainly don't see the amortised cost across the decades long operation of a highly successful industry.
So it comes across as a far deeper cut than the stubbed toe that it actually is.
That hour you spent chasing an NPE? Yeah, it happened. (really not that often though) The time saved by having a single, universally accepted way of implementing optionality, even it's inconveniently always on? That's invisible because we did not live through the alternative.
I guess you can take that up with Turning Award winner C. A. R. Hoare FRS FREng. I'm going to go ahead and keep citing it despite being a epic underestimate.
> But it wasn't so obvious at the time.
It was, just not to the hackers that invented the languages plagued by these mistakes.
Sure it was. That was how it worked before. Hoare created null because of pressure from other programmers to write shitty code that looked like it worked even when it didn't.
How so? Null checks are mostly “free” on the JVM, they trap on the zero page and get converted to an exception via a signal handler.
I remember working on a codebase where everything had null check pyramids like
if (foo != null) {
if (foo.getBar() != null) {
Bar bar = foo.getBar();
if (bar.getBaz() != null) {
for (Quux q : bar.getBaz()) {
if (q.getFoobar() != null) { ... }
}
}
}
but that was a long time a go in a code base that operated by passing humongous mega-objects with hundreds of fields (running into the limit of how many parameters could be in the constructor was a constant headache), it solved a lot of its problems by printing a stacktrace and shrugging, and objects would be half-constructed and methods would return null in all sorts of scenarios.In comparison, the entire codebase for marginalia search has about 65 nullchecks in total. A decent chunk of them are in dealing with older APIs that sometimes idiomatically use nulls to communicate things. java.util.Map, BufferedReader and the Servlet api in particular, the rest in dealing with the inherently noisy nature of crawl data. So while there are a few examples where there are nulls in the data, but they are almost always contained to the module that produced them and don't cross interface boundaries. It's very far removed from some pinnacle of code quality, but just gets a few engineering principles right that that other code base didn't.
Personally, I lean towards the "use nulls sparingly and be very explicit where null values are possible", use immutable objects etc. This reduces the unnecessary null-check noise and makes things more robust. In java, I even prefer to add @Nullable to fields to try to switch the default mindset to "assume non-null unless @Nullable annotation exists" - though another developer made a note on a code review that "that is not how Java is designed and @Nullable is just noise".
Having nullability as part of the type makes everything more explicit, so I'm a huge fan. Beyond just the perf and safety benefits, it's a win for domain modeling.