Overhead of Returning Optional Values in Java and Rust
pkolaczk.github.io
pkolaczk.github.io
Sadly, the two benchmarks are not equivalent, the Rust version loop over ints, the Java version loop over longs.
Looping over longs is notoriously slow in Java, because there is no vectorisation and the JIT insert insert a GC checks because looping over longs may take a long of time.
Also one of the codes uses Optional<Long> instead of OptionalLong, so the code does a double boxing long -> Long -> OptionalLong.
I thought longs were 64 bit signed integers in Java, what's the difference with a u64 in Rust? Is it the signedness that affects performance?
> OptionalLong
Thanks for the info! I didn't know there were Optional primitive types in Java!
Because the function get_int() takes a u64, Rust will conclude that the variable i must be a u64 so as to fit in that parameter.
It seems to me that "Java is bad at 64-bit integers" is also useful information but not pertinent to this benchmark, unless you assume it's trying to compare Java to Rust, at which of course Java will be far slower - but I don't think that's the point of the benchmark at all.
You really need first class support for optionals, a la TypeScript and Kotlin, to get a benefit IMO.
- it’s verbose because handling optional values safely is complex
I mean, no shit. But the problem is that tons of existing Java APIs do use nulls, so you can't just pretend they don't exist, and importantly a major point of the optional construct is to get compile-time guarantees, and the Optional class doesn't give you that.
> it’s verbose because handling optional values safely is complex
No, it's not. Kotlin has optionals built into the language, and it's safer and much less verbose.
Their recent decision to axe typeclasses in favor of multiple function receivers, which solves a small subset of problems that HKTs do in Scala and Haskell, make me look for greener pastures whenever I need them.
That's why you're supposed to use Optional<Optional<X>>. ;-)
The great thing about Rust is that `Option` is a normal, generic type. You could very easily implement your own equivalent type.
So it can optimize the memory overhead of Option<> away completely.
I've found I agree with the principle that Optionals only make sense for return values (and if the method returns null, that should be treated as a bug).
I've seen Optionals used for member variables, and they might as well just be null. (And are highlighted as such by IDEs like IntelliJ).
Java On its own it doesn’t support this however there are ways of enforcing this in a Java codebase. Eg: Errorprone and NullAway comes to mind here.
Does it change the benchmarks? Not if you let the inner function get inlined: https://rust.godbolt.org/z/Ejc5c7sE8
Allocating anything like a “Long” or “Optional” on the heap seems like an extremely bad idea. Even if allocation and GC is almost free, the pointer chasing might not be.
Here’s a fascinating read on the current state of HotSpot escape analysis:
https://gist.github.com/JohnTortugo/c2607821202634a6509ec3c3...
2021 was the first year I deploy anything related to .NET Core in production, and it was still .NET Core 3.1.
Will this velocity no one is going to their .NET repos and rewrite the libraries with the required annotations to keep the compiler happy.
If you want null safety in Java today, just make use of @Nullable, a good IDE with static analysis, and deploy PMD, FindBugs or SonarQube into the CI/CD pipeline.
If you want to improve nulls handling in Java, you should use @NotNull annotation and corresponding linters.
Is there any analog for a const pointer? I want to point into or just beyond a const array, intrusively using null as a None value.
There is `ptr::NonNull` but this always converts to a *mut T, and one is just supposed to "be careful" to not use the mutable features if your pointee isn't actually mutable.
If you look at the implementation of slice::Iter, it stores a (cursor, end) pointer pair, but there's no way to say "this is a non-null const pointer" so it casts the cursor pointer to *mut.
Do you know what optimizations are disabled? Are they disabled automatically, or only if certain parameters/functions are used (e.g., BlackHole)?
"JMH provides a [snip] foundation for writing and running benchmarks whose results are not erroneous due to unwanted virtual machine optimizations"
unsolicited personal opinion: they're almost always erroneous for other reasons
i'm not sure which optimizations are currently disabled
To be fair, now that I think about it I suppose some deoptimizations make sense (e.g., preventing replacement of loops).
Thanks for pointing me to that article!
> The most ugly and error-prone solution turned out to be the fastest: primitive types and magic values.
That's often the case in "older" languages. ES6 iterators on arrays are usually twice as slow as a simple "for" loop. Which usually means that by targetting ES5 or lower, you get a free optimization by Babel/tsc!
In the NonZeroU64 case Rust can literally do the same trick with magic values ... in the machine code, the magic value is zero, a NonZeroU64 can't be zero and so Rust will squeeze None into that value of the register.
The programmer and their colleagues both reading and writing the code don't need to understand how this trick works, don't need to be aware of magic values, and if some day the code needs a full u64 and a NonZeroU64 isn't good enough, they just change it and sure, it might be a little slower because the optimisation wasn't available, but it still works as expected because magic values were just an optimisation not part of the program logic.
HotSpot has a reputation of being close to a Sufficiently Smart Compiler, so it’s interesting to see where its limits are, compared to lower-level language constructs.
For example, here's a post about Discord switching from Go to Rust with great success https://discord.com/blog/why-discord-is-switching-from-go-to..., and here's another about how CockroachDB optimizes Go https://www.cockroachlabs.com/blog/how-to-optimize-garbage-c.... In the same vein, there's Shopify AOTing Ruby using Sorbet and LLVM https://sorbet.org/blog/2021/07/30/open-sourcing-sorbet-comp....
All of that to say that I'm not surprised that Rust has better defaults compared to Java, as Rust is supposed to be faster, and option types are a code Rust language feature, while they are a relatively recent addition to Java.
Beating Java with optional values is nice, but meaningless.
Speaking of, one thing that bothers me about nearly every Java benchmark ever, is nobody checks to see what was _actually_ compiled to native code. To really compare Apples and Oranges, was the code being executed in interpreted mode? What about compiling with Graal? These questions are important for any long-lived program (which most Java programs are).
Here, the warmup iteration of "2" and limit of "5" test runs doesn't really accomplish anything. Depending on the JDK the default compilation threshold is 1.5k or 10k IIRC.
Is this still true in the world of CI/CD? If you're following "best practices" and doing multiple deployments a day, how much will your performance suffer? I don't have any answers to that, though I suspect that it may be one of the reasons why Go is used so much for "cloud stuff".
1500 or 10k in production for hot methods should only take a few mins in most production code I'd hope. Workload dependent, of course!
And, you can always tune the JDK for your workload. An optimizing compiler is a very powerful tool
It's been more of a surprise to me to see people defending null as somehow preferable style in these comments.
...
sum = 0
opt = getOptional(n)
sum = sum + opt.orElse(0);
...
Also maybe use something like OptionalInt (but for long).Every NullPointerExceptipn my code ever caused was good, because it helped find an error state that had no acceptable default non error state alternative.
> Every NullPointerExceptipn my code ever caused was good, because it helped find an error state that had no acceptable default non error state alternative.
Yeah, and with Optional you would have found it at compile-time.
so we put on our big-dev pants and build all sorts of structures around handling errors to make it more sound/ergonomic.
but at the end of the day:
o its still very difficult to respond in a semantically meaningful way to errors as code
o they still almost always just get dumped into a log or ignored, or just result in a panic
o even if we have managed to surface them - they still aren't really that actionable by the end user
so we haven't helped the situation much, but some of these answer like catch/throw have really unpleasant consequences - we may have made it worseScala's ZIO is one of the better examples if you are interested to look into an alternative.
I kind of think that sentence is contrary to the point of Optional, and if you start from that mindset then Optionals do end up no better than nulls.
Instead you enforce good hygiene where you unwrap early and pass non-Optionals for anything that you "know" is non-null. If you call a function which returns optional, and you're find yourself thinking "nah, I _know_ this is definietly non-empty" then you're incorrect; the API explicitly is telling you that it definitely can be empty.
If you just do "Map.get(x)" and you get null because your thinking was wrong, then the NPE will pop up potentially 10 layers later or an hour later, when the null-value was tried to be used in your program.
On the other hand, with "Map.get(x).orElseThrow()" you will have an exception thrown immediately! Even better, you should customize that as "Map.get(x).orElseThrow(NoSuchElementException(x))" so that you know the value that was unexpectedly not in the list.
That will make debugging much easier and will potentially fail a test case while a "Map.get(x)" might not cause the test to fail.
case class Address(street: String, city: String, state, String) case class User(id: Int, name: String, address: Option[Address])
def filterCAUsers(users: List[User]): List[User] = users.filter(_.address.exists(_.state != "CA"))
If you just want to provide a hint to the caller at compile time about whether you intend for the possibility of returning null, you could use @Nullable/@NotNull annotations to achieve that with no overhead.
E.g. you might want to return the streetnumber of a user's address. When a NPE is thrown, was it because there is a bug somewhere in the code or because the user's address has no street number?
With Optional, it is clear: there was no street number.
Isn’t that the point of Optional? It just moves that identification of those error states to compile-time rather than at runtime.
The only sane approach is removing all nulls as early as they come and declaring/initializing all member variable in the c-tors.
The ideal state is that anything that isn’t Optional is guaranteed to be non-null. So when you go ahead and remove all nulls as early as possible, you can statically guarantee that remains true forever-more.
But if all your libraries just go ahead and drop this knowledge anyways, and when they return an object without meaning non-nullable, then you’re back to where you started.
Whereas in an ML, I know every library is already adhering to this, so something being not-optional actually means something. Bolting it into a language ad-hoc, you go from everything can be null to this weird tristate — a type might be nullable, or might not be nullable and if it’s in an optional, it definitely is nullable… so your best bet is just checking for null every time anyways
Since I started using "safe" nulls with Ceylon, nearly a decade ago, later with Kotlin and TypeScript, I've been firmly in the camp of supporting nulls in the language instead of `Optional` or `Maybe`. Rust should've used them IMO... using `?` for error propagation instead is such a weird choice.
Example (from [1]):
// Assume
// fn halves_if_even(i: i32) -> Option<i32>
fn do_the_thing(i: i32) -> Option<i32> {
let i = halves_if_even(i)?;
// use `i`
}
Besides, the `?` is syntax sugar for the `try!` macro as explained in the `try!` macro docs [2], and has nothing to do with sum types, at least not directly.Ceylon would be a closer case where nullable types are represented as sum types and not a special-case in the language (`String | Null` is the same as `String?`)... but Ceylon, as Rust, special-cased the `?` operator for "something"... in Ceylon, they are used for null-checks, and in Rust, for propagating errors via the `try!` macro, which is what my previous comment was about.
[1] https://stackoverflow.com/questions/42917566/what-is-this-qu...
> Ceylon would be a closer case where nullable types are represented as sum types and not a special-case in the language (`String | Null` is the same as `String?`).
Aren't those union types? My understanding is that they are different from sum types, as you usually have to declare the sum types beforehand, and sum types are collections of values, while union types are collections of types, and are often declared "on the spot". At least that's how they work in Typescript, I don't know much about Ceylon.
To go back to option types, people usually like them because there are well-known patterns to program with them (like map). Of course you can get the same safety with static flow analysis (Typescript does this), but I think part of it is ergonomics, influenced by functional programming. Though if there is an overhead like in Java, I'm not sure about their value.
Ceylon:
alias StringOrInt = String | Integer;
StringOrInt myValue = "foo";
StringOrInt otherValue = 10;
Rust: enum StringOrInt {
Str(String),
Int(usize),
}
let my_value: StringOrInt = StringOrInt::Str("foo");
let other_value: StringOrInt = StringOrInt::Int(10);
In Ceylon, you're right you don't need to name a "union", which IMO makes it much nicer, as I was able to use adhoc unions everywhere, e.g. inside streaming operations: [ for x in foos if (x.something()) "foo" else 10 ]
The above "automatically" has type `[String | Integer]`. In Rust (and Kotlin) I would need to name a type just for this.Another disadvantage of Rust's approach is that the type variants are not types themselves. That's really annoying sometimes, e.g. you want to have a method that only handles one of the variants, but you can't (you need to pass the values the variant wraps separately, but that may be undesirable semantically as that would allow something that has no instance of the enum to also use the function).
As an interesting aside, Rust has what seem to be intersection types for trait bounds (fn toto<T: Trait1 + Trait2)(a: T)).
I'm still not sure about the tradeoffs between union/intersection types and sum/product types. I think the difference is that "set types" are structural, while algebraic types are nominal, which comes with the usual tradeoffs, but I'm not sure.
It is like Google uses their pseudo Java dialect when selling Kotlin features, instead of comparing it with proper Java.
Option types are 8 years older than null (ML: 1973, null: 1965). There's nothing clever about it. It's simple to understand, and is an explicit wrapper around an object, while a pointer is an "implicit" pointer around an object. Option<Obj> is clear, you know that it can be something or nothing. Plain Obj is deceiving.
> Every NullPointerExceptipn my code ever caused was good, because it helped find an error state that had no acceptable default non error state alternative.
That's the point of option types and pattern matching, it can force you to handle both cases.