Java Is Unsound: The Industry Perspective
dev.to
dev.to
I was at the conference where this was presented. The paper is terrifically written, and it's good that someone is investigating this.
That being said, it's just not that big of a deal. In practice, this will never be a problem for anyone writing Java.
The Java wildcard system is insane and no-one is trying to build type systems based on these ideas, not even Scala and the new batch of JVM languages (Kotlin, Ceylon). They have their own problems, but this isn't it.
public interface Helper<T extends Number> {
}
public static <T extends Number> Number nice(Helper<T> helper, T value) { // (1)
return value;
}
public static <SUB_T> Number naughty(Helper<? super SUB_T> helper, SUB_T value) { // (2)
return nice(helper, value); // (3)
}
public static Number evil(String value) {
return naughty(null, value); // (4)
}
From what I understand, the following is happening:1: Nothing unusual here. Just defining a type parameter with a constraint and referencing another type which happens to have the same constraint.
2: I think this is the actual bug (?) - SUB_T is permitted as the subtype of a constrained type, yet itself is not subject to the constraint. This compiles.
3: The compiler derives T from pattern matching (this is the sole reason why the helper is needed)
T is inferred as the set of hypothetical types that are both subtypes of Number and supertypes of SUB_T.
4: Because SUB_T is unconstrained, we can happily instantiate SUB_T such that the set of hypothetical types is empty (e.g. set it to String).
This does mean there is no way we could create a helper object as no type would satisfy the constraints - but we don't actually need a helper object, so we can simply pass null.
In short, this bug lets the compiler infer a type hierarchy of Number -> T -> SUB_T without ensuring that Number -> SUB_T holds. Usually, that wouldn't get you far as there is no type that could substitute for T - except that you don't need a type because the compiler will accept null even for "impossible" types.
It's interesting to note that the class cast exception occurs in (3), not in (1). From what I know, that is because Java's erasure takes constraints into account - the arguments of nice() become (Helper, Number) after erasure, not (Helper, Object).
People like Paul Philips (yes yes, I know he's ranty and rubs people wrong and is kinda intense, but he makes good points) show that you can build a sane language on top of the JVM, but Scala makes a lot of the same concessions and refuses to fix a lot of implementation bugs in favor of backwards comparability and enterprise support.
I wrote Kotlin off originally as a late into the game pull by Jetbrains, but now I really want to take another look at it and see what type of library and JVM support is has.
Lately bigger players have started to take notice:
https://spring.io/blog/2017/01/04/introducing-kotlin-support...
https://blog.gradle.org/kotlin-meets-gradle
I'm still not sure where it will end up, but things look pretty good right now.
But wrt Type Erasure, Kotlin doesn't completely get rid of it; but it does offer some ability to retain Types in some generic situations.
You probably would never write this, and your colleagues would never write this. That reasoning is good to apply in many situations. I in fact do this all the time; it’s a huge part of my research agenda. But you have to be careful where you apply it. Soundness is a security guarantee, not a usability concern. What matters is if someone can write this, because then that someone can get around the security measures that people have placed their confidence in and do whatever they would like to do. In the case of soundness, it’s the malicious programmer you should be worried, not just you and your friends.
I really hate this attitude. Programming languages should be designed to be safe and strict by default, but with easy outs. I utterly reject the notion that we should design languages first and foremost to protect us from people can't program that well or aren't careful when they write code.
Do we always have to appeal to the lowest common denominator? Empower the developers that care rather than trying to turn this into an unskilled profession.
C# has Nullable<T> (aka T?)[0]. I'm not a Java expert, but they appear to be the same (represent a value or a null value, allow you to inspect for the presence of a value or null).
ex:
int? x = null;
int y = x ?? 0;
int z = x.HasValue ? x.Value : 0;
if(x.HasValue && x.Value == 42){}
C# definitely doesn't have checked exceptions.edit:
For reference types, as compared to value types, c# has the null coalescing operator
List<string> y = someList() ?? new List<string>;
and the null conditional operator int? length = customers?.Length; // null if customers is null
which are defintely not Optional, but are what I use.[0] - https://msdn.microsoft.com/en-us/library/1t3y8s4s.aspx
Checked Exceptions was a design choice, http://www.artima.com/intv/handcuffs.html
What is really necessary is an equivalent for reference types.
struct Maybe<T> where T : class {
private bool _isInitialized;
private T _data;
public Maybe(T data) {
_isInitialized = (data != null);
_data = data;
}
// ...etc
}no - Java's Optional applies to reference types.
Your quote also omitted checked exceptions, which is another reason why it's easier than me to write robust code in Java than in C#.
public static TDest IfNotNull<TSource, TDest>(this TSource obj, Func<TSource, TDest> func)
{
return obj != null ? func(obj) : default(TDest);
}
Or something similar (writing c# code on the iPad before sleeping is not exactly simple), inspired by some blog post ages ago.
Just adding these three line of codes in a static class in the solution will enable you to write something like this: var id = order.IfNotNull(o => o.Customer).IfNotNull(c => c.Id) ?? "INVALID"
Still not as good as the null check operator but comparable with the Java optional (arguably better from my point of view) String id = order.map(o -> o.Customer).map(c -> c.Id).orElse("INVALID");
The Java code is actually shorter and works straight out of the box, so I don't know why you're calling Optionals "verbose".I think that was the point he was making, that it's always been possible, and now is simple.
string? id = "Name";
While you can do stuff like this: int x = 0;
int? y = x;
Again, F# has nullable types and options, and the fact that no one in F# uses nullable types should tell you something.According the documentation, Optional is only valid with value types[0]:
"This is a value-based class; use of identity-sensitive operations (including reference equality (==), identity hash code, or synchronization) on instances of Optional may have unpredictable results and should be avoided."
Again, I have no experience, so if this is wrong, then additional corrective references would be awesome.
[0] - https://docs.oracle.com/javase/8/docs/api/java/util/Optional...
http://fsharpforfunandprofit.com/posts/the-option-type/
Scroll down to "Option vs. Null vs. Nullable" for an explanation of the differences.
Now, if a programming language has error aware return types - using algebraic data types or multiple dispatch - that's a better solution than both. But in the absence of that, give me checked exceptions over unchecked ones any day.
You absolutely can:
Microsoft (R) F# Interactive version 14.0.23413.0
Copyright (c) Microsoft Corporation. All Rights Reserved.
For help type #help;;
> let s:string = null ;;
val s : string = nullI don't think the author has a beef with easy outs, but with accidental outs. Specifically with regards to the JVM, too, there are some practical security ramifications, with regards to running plugins.
It returns a ClassCastException because underneath it all, in the Java bytecode the Java compiler stores generics as casts. Therefore it will complain because it can't cast. If they had baked generics in at the beginning of Java then it wouldn't have compiled.
It would be the equivalent of writing String one = 1;
I hope professors today teach the 1.4 vs 1.5+ equivalents and explain what happens in the underlying type system.
A lot of these problems with the JVM are exacerbated in the Scala world. With such a heavy dependence on type matching in Scala, things like
a match { case List[MyCustomType] =>{}; case List[SomeOtherType] => ...
become impossible since the underlying types are removed at compile time.Yes, the Scala docs do explain that type aliases are not really distinct types, so this sort of bug could have been expected in advance. I came to Scala with no prior Java experience so I got to discover the ways the JVM leaks through on the job. I thought that I could reuse my previous knowledge about types in program design in obvious ways and just get the obviously (to me, anyway) desired behavior. Mea culpa!
example:
object Example {
type IdA = Long
type IdB = Long
def usesA(i: IdA) = "it's an A"
case class IdC(i: Long) extends AnyVal
case class IdD(i: Long) extends AnyVal
def usesC(i: IdC) = "it's a C"
val a: IdA = 1L
val b: IdB = 2L
val c = IdC(3L)
val d = IdD(4L)
usesA(a)
usesA(b)
usesC(c)
// these won't compile
// usesC(d)
// usesC(a)
}Right - the article already says that doesn't it?
> If they had baked generics in at the beginning of Java then it wouldn't have compiled
Why wouldn't it have compiled if the JVM had generics? It's an error in the design of the type system of the language isn't it? The type system is the same if we have reified generics or erased generics isn't it? And that's where the failure is - not at the JVM level.
I'm not even sure how you would be able to write this in Scala without violating rule number 1 of Scala, which is "never write the word null in your code".
Look for the "Thanks, NULL!" section near the bottom.
So, right now, nobody can show an exploit. But security researchers often manage to find exploits given this much to start with.
Whenever you instantiate a generic instance, you need to provide a concrete type--be it a known class name, or a named type variable. This applies both to instantiation via new and to the inference of generic method arguments (you can't say Unsound.<?, ?>coerce(), for example). Since named type variables come from outer generic instances that need to be instantiated themselves, this provides the inductive proof.
The first fly in the ointment is wildcard types. Wildcards are only legal as the type of a variable. They also don't work like you expect--each variable gets a different "capture" of wildcard types. Effectively, you use the capture conversion process to build up a complete set of constraints that a wildcard type must satisfy, and then you have to prove that it can be satisfied.
The actual problem comes when we are asked to provide proof that there exists a type that satisfies the constraints. Normally, we would derive the value from something that forces an instantiation of a type. Instead, the example passes in literally a null proof: the null value, or I have no proof that such a type exists. Even then, it's not so bad, except we borrowed the "proof" provided by the null instantiation in a different context to prove that we could go from an Integer to a String.
Put another way, the flow of the program works like this. The Constrain class is used to carry around in the compiler the notion the subclass constraint, and to provide the thing where wildcards get used. The upcast method is used to borrow the proof-of-existence from Constrain to actually do the cast. (The compiler doesn't complain here because it's all in terms of type variables--if there's a problem, it will be caught at the calling site).
Since the coerce method declares no bounds on the type variables, there's no reason for line 15 to miscompile--all of the constraints between T and U have to work under the assumption that there is no axioms about their relationship. Line 10 compiles cleanly because all we're doing is stating that, to use this variable (store anything in it), there needs to exist a type between T and U. The use of null means we defer it to later. Line 11 has no problems whatsoever. In Line 12, we use the fact that ∃x: T ≤ x ≤ U to capture that x and so cast from type T to x to U without any problems.
Fixing this is not easy. You somehow have to modify the type system to reflect the fact that a null value doesn't prove the existential qualifier, and it's not clear to me that there's a clean way to retrofit that into the way type checking is currently done in Java.
[1] http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/p017...
I'm not saying that C++'s type system is great. I'm just saying that Java's problem mystifies me.
Most problems described in article are not for backward compatibility. But because java at runtime (JVM) is dynamically typed (Java compiler is statically typed). You can do much nastier tricks with reflection.
Scala is not much better, I wrote some examples: http://www.mapdb.org/blog/scala_has_weakly_typed_syntax/
>Is it safe for Foo<String> to be a valid type?
Wrong question. Is it useful for Foo<String> to be a valid type? The answer is probably no. So why allow it? The below example does not come close to explaining why this should be a language feature.
First, it calls unsound the fact that an exception gets thrown in the shown code. There is no data security issue, no possible JVM escape, no type system failure. You can get the same cast exception with straight-forward code. The claim of unsoundness is merely based on the fact that the problem is obfuscated. Using unsound here is misleading.
Second, the whole point relies on implicit type relations that are not checked at compile-time when nullptr is involved. Certainly, it's a weakness in the specification that can cause suprising run-time exceptions, but it's not a hole in the type system.
Third, it makes claim that the JVM was saved by its early choice of backward compatibility. This is a bad habit that I've seen elsewhere: making claim about unspecified, unknown alternate histories. In this case, had generic not erased type, then assignment from nullptr would not work. Surely, the nullptr would have acquired typed alternative (nullptr<T>: nullptr<String>, nullptr<Integer> ...) which would have letthe compiler type-check properly and catch the problem at compile-time.
"Here’s our “true” unsoundness example that has been making the rounds most recently. It assigns an Integer instance to a String variable. According to the Java specification, this program is well-typed, and according to the Java specification it should execute without exception (unlike the aforementioned backdoors) and actually assign the Integer instance to the String variable, violating the intended behavior of well-typed programs. Thus it illustrates that Java’s type system is unsound."
That is, the author's unsoundness claim is independent of the class cast exception.
On your second point, you may be in violent agreement with the author. What is the difference between "a weakness in the specification" and being "unsound"? The author's definition is that this is an unintentional hole in the type system, unlike the intentional ones he outlined.
There are other examples (e.g. see for Java 8 http://www.oracle.com/technetwork/java/javase/8-compatibilit...) where they've broken actual source backwards compatibility.
interface IFoo<T> { Number foo(T t); }
class Foo<T extends Number> implements IFoo<T> {
public Number foo(T t) { return t; }
}
Number bar(IFoo<String> foos) { return foos.foo("Nan"); }
This implies that something bad is going on. But the class in the middle is a red herring. It is not used in this example, nor can it be used as an argument to the `bar` method. Try instantiating a Foo<String> and it wont work. Try instantiating a Foo<Integer> and then passing that to bar(), which requires an IFoo<String>. Both are compile time errors.By all means claim, correctly, that support of unchecked code and raw types means that the type system is not "sound", but the incorrect examples undermine the claim that this is a problem. (I don't personally think that its a problem, and the fact that someone claiming otherwise can only provide examples that fail to prove that its a problem isn't terribly convincing. But then I grew up with 6502 ASM, which had a fairly limited type system).
Like you note:
> Try instantiating a Foo<Integer> and then passing that to bar(), which requires an IFoo<String>.
That's what Ross is trying to get you to see. He's saying that Java does not have to make the declaration of IFoo<T> an error because it can rely on the fact that your later instantiation of it with <String> will be an error instead.
As long as the bad code gets flagged as an error somehow before it gets used, you're safe.
The actual unsoundness bug is the "class Unsound" example in Josh Bloch's tweet. What it shows was that you can't safely rely on "we'll flag an error when you create an instance of this type, so we don't need to flag an error in its declaration." Because you can assign null to be an instance of the type and then you sidestep the later error.
class Foo<T extends Number> implements IFoo<T> {}
Does not implement an unconstrained T. It implements "T extends Number". When the T appears in an implements, its not declaring a new type. You can't do, for example
class Foo implements IFoo<T> {}
java: cannot find symbol T
The mistake is believing that the T's, in the example above, are in any way related. It would be more appropriate to write: interface Foo<TUnconstrained> {}
class Foo<TExtendsNumber extends Number> implements IFoo<TExtendsNumber> {...}
And then it becomes clear that the foos in bar will never be a Foo, because Foo does not implement IFoo<String> and never will. Yes you can nuke it unchecked code. No, I don't have a problem with that. But the example given is not a useful example.And this example is in response to the previous one where he asks "Is this code safe? .. your gut instinct would say this is unsafe". No. My gut instinct is that it would not compile. And indeed:
java: type argument java.lang.String is not within bounds of type-variable TI don't know what you're disagreeing with.
> And then it becomes clear that the foos in bar will never be a Foo, because Foo does not implement IFoo<String> and never will.
The statement "Foo does not implement IFoo<String>" is not meaningful. "Foo" is not a type, it's a type constructor. To refer to a type, you need to say Foo of what. And it is the case that Foo<String> implements IFoo<String>.
Consider:
interface IFoo<TUnconstrained> {
Number foo(TUnconstrained t);
}
void useFoo(IFoo<String> foo) {
Number n = foo.foo("not a number");
}
This code compiles fine. Add this: class Foo<TExtendsNumber extends Number> implements IFoo<TExtendsNumber> {
public Number foo(TExtendsNumber t) { return t; }
}
It still compiles fine. This is exactly what the author is saying. The reason it is fine for the compiler to permit the above code is because it is relying on the fact that you will not be able to instantiate Foo<String>.This is one half of the puzzle of the actual unsoundness example. The other half is that you can then sidestep that restriction using a wildcard to define a type that hypothetically meets the required bound even though no such real type exists.
That would still be safe, since you've never be able to initialize that variable since no value of that type is possible. Except that Java lets you initialize it with null.
David Byrne repeating "Same as it ever was" comes to mind ... ( https://www.youtube.com/watch?v=I1wg1DNHbNU )
Meanwhile, javascript is winning.
From a practical perspective, static type systems have proven useful for tooling: to help us with code completion and refactoring. Beyond that, they appear mainly to hurt developer productivity.