Type inference has usability problems
austinhenley.com
austinhenley.com
The author is talking specifically about inferring the type of a variable from the type of the LHS of its initialization. Opinions differ on how much you should use this (he points out Go, and Go is nothing if not opinionated) but there are clear benefits to having it as an option.
Suppose you have
foo(super_long_name(), x, y, z)
and you want to refactor it into var bar = super_long_name()
foo(bar, x, y, z)
This is only a stylistic change, but if you have to write down variable types, you need a piece of information for the second that you don't need in the first. If the first was not unclear (just long), why should the second be?The type name may also be difficult to write down. It might also be unenlightening. Knowing the type of a variable is not always helpful during normal reading (eg. the types of iterators in C++ or Rust).
Also in generic code (eg. templates or macros), you might just not know the type of some expression. I believe this was a major argument in favor of adding auto to C++.
Not only that, but there are some types in Rust that you _can't_ even specify; closures specifically don't have concrete types that can be specified in a Rust program.
I don't think that's an improvement. Languages like Java and C++ don't work that way, and they're the ones being decried for having to spend too much time writing out redundant types.
I think the problem is that people don't treat types as documentation. List them explicitly when it clarifies. When it doesn't, leave them out. A lack of documentation is better than useless documentation. Treat types the exact same way.
Your conclusion is the same as mine. Use them when appropriate. When is appropriate? We need some data on that.
I know a lot of teams that conduct code reviews outside of the IDE, so your suggestion wouldn’t help those cases.
(This is obviously a very personal preference and depends on whether it's a code base you're already familiar with, etc. etc.)
ML and Haskell, for example, have basically had (local and hear-global) type inference since the 1980s.
[1] As of, say, 2009 or so where even local type inference was relatively rare in mainstream languages.
Err.. I think it's the other way around.
and yet as soon as you introduce auto you've got the mob bringing out the pitchforks :-)
Maybe it's because I've spent so much time in dynamically typed languages, but I see var/auto as absolutely essential to making static typing palatable. I'm a big fan of static typing, but I would be much less so if I had to do it without inference/deduction.
List<String> names = new List<String>();
Why not let people write just: List<String> names = new List();
Why did they even need to invent the diamond operator?I quite like type inference when it's clear from context what the type will be, e.g.:
var xs = new List<Int>();
for (var x in xs) { ... }
I'm pretty happy for it when the type "doesn't matter" because you don't do anything with the value except pass it elsewhere. // Presumably refactored from processSomething(x, y, z, getSomething()) because the line got too long.
var something = getSomething();
processSomething(x, y, z, something);
And sometimes the type can be sufficiently unwieldy to name, most famously in C++ taking advantage of templates, but also can show up in other languages such as Haskell that encourage this behavior.Finally there's this comment:
> A weaker argument I have heard is that you have less code to type with type inference!
I'd say that I agree with it as stated, but I'd hope the people making that argument are actually trying to say "it's less code to read/be distracted by". IMO omitting the "less important" type information draws the reader's attention to the ones you did write, which are presumably the more interesting parts of the code.
Having said all that, I think my brain is a bit quirky because even in languages like Python I automatically/internally infer all the types based on the names of variables and roughly how/where they're used as I read through. Maybe it's an experience thing (or maybe I don't work with bad/differently-thinking developers), but I find that my first or second guess is right 99% of the time.
// Presumably refactored from processSomething(x, y, z, getSomething()) because the line got too long.
var something = getSomething();
processSomething(x, y, z, something);
Sigh, people will add extra lines, refactor and risk introducing bugs simply to avoid a text line longer than a punch card from 1965.I use it as an indicator that I am either nested too deeply or my functions have too many parameters. Or just in general my code is too complex / hard to read after the fact.
Sometimes things call to break the soft limit that is fine, it's only a signal of potential problems, not an indicator of definite problems.
And fwiw, if you can't refactor like that without concern it's a language design failure. A pretty big one actually.
First, if "var" was such a huge problem then language like python would be totally unusable! The entire language doesn't have explicit static type definitions.
LHS type inference is a tool. Like literally anything else you can write code that makes things less maintainable using this tool. The languages that support LHS type inference all have style suggestions that you use it when the type is either not especially relevant or is clear from the expression RHS. People don't complain about type names when they are useful. People complain about type names when they are extraneous noise. That is when LHS type inference is useful.
So, no, it's not ridiculous. Overstated, possibly. But I am less interested in Go than ten minutes ago.
Yes, but there's _plenty_ of use for static type validations even if you don't use type annotations.
I certainly feel that way.
(Disclaimer: I've not used it, I just think it's a reasonable idea.)
Why do you think so?
The hardest code bases I've worked with have been very explicit with type info and that's done nothing for me. The examples here feel very contrived; very little of the code in a 100k LoC project is gonna use types like Maps and ints, especially if it's OOP. Knowing that something is an AbstractIterator does very little to help me if I don't know WTF an AbstractIterator is. And complex inheritance chains absolutely destroy my ability to find the method I'm looking for on these types.
Coming from functional programming, it really doesn't feel so different looking up a function vs. looking up a class.
Regardless, every language has some degree of inference, otherwise every function call would be:
let result : int = (f : int -> int)(x : int + y : int, 2 : int)
If your code looks like this:
f(g(x, y), a(b(c)))
then inference isn't really doing much to hurt you. And what I've found is that for any large-scale project you see a lot of type of code.
Why do you think that? Map is a super common Java type, which is what the example with the most Map in it is written in. Hard to see where 'complex inheritance chains' come into it, as well.
Complex inheritance chains come into this because it makes it hard to track down the code you're looking for. Like let's say we have Cat : Feline & Pet, Feline : Mammal, Mammal : Animal, Pet : Property, Animal : Alive. Now when I do cat.eat(x) it's impossible to predict where eat is defined. Knowing that it's a Cat didn't make it any easier to understand the implementation of eat.
Err... do you mean to say type inference isn't doing much to help you?
var gxy = g(x, y)
var abc = a(b(c))
f(gxy, abc)var a = f(x) var b = g(y, a)
Isn't any different from:
g(y, f(x))
A lot of real world code I encounter doesn't bother with temporary values for everything and giving them explicit types, which is essentially the equivalent of just using inference here. Thus in this case inference isn't doing anything negative because the code would look like that regardless.
Also I've noticed in experimenting with Rust lately that use of type inference in code examples slows down the learning process. Often times I have to look up a few method signatures to figure out exactly what's happening in a given code block.
That said, I can't imagine working in modern languages without type inference altogether. I tend to agree with the last argument that the author doubts: if I repeat redundant type information several times in a scenario where it should be obvious, that makes my code less clear.
If you write C++ STL code or C# LINQ code or any other generics/template heavy code you probably will appreciate type inference as a huge improvement.
When I'm writing hundreds of lines of frame layout code, I have to explicitly write `let transientValue:CGFloat = ...` hundreds of times or else the compiler quits on me saying that it was unable to infer the required type before timing out. The error makes sense if you think about what the compiler is doing (solving a system of equations with all of the number like built in types) but frankly it's unnecessary to begin with.
Compiler performance aside, it also makes working on code a lot heavier since the editor needs to be able to infer the types (something the full blown compiler struggles to do given all cores and tens of seconds!) before it can give sane autocomplete suggestions. I'm not sure if a better IDE than Xcode could handle this better but I've found the tools have forced me to be explicit just so I can get my work down more quickly.
In Swift's case it seems like they've tried to add very flexible type inference (leaving off lots of stuff) to a very flexible type system (with inheritance, overloading, protocols, etc.). This is something that academics have been aware for decades causes problems with type checking performance. Similar problems have caused the designers of Scala 3 to go back to the drawing board to redesign the type system from the ground up, with a clearer formal semantics.
I do know that the Swift team are working really hard to make on demand editing/compiling better, but I don't know how far that work has come. Hopefully that improves things over time!
But type inference in Rust and in Haskell can get pretty wacky, with the type determined by code half a page away. The built-in limits in C++ seem well-advised; anyway it is very rare to hear anyone asking for more, or complaining about code over-depending on it.
It is very common, in C++, to drop in a modifier, like `auto* p = f()` for both clarity and insurance. Usually that is just the right amount of type information.
Hence the general guideline (in Haskell at least) that all top-level definitions should have a type ascription. You can also add ascriptions locally if you have a large function that you think would benefit (readability-wise) from it.
(I still very much appreciate Scala for what it is. No real experience outside of toys with Rust.)
The latter is something that isn't in any languages I would consider production-ready (although it seems tantalizingly close), but the former is already quite useful depending on how much type inference your language is capable of. If paired with the ability to search for code by type, it makes exploring APIs extremely friendly. Even without that ability, I've found the ability for the compiler to tell me what type belongs in a given empty chunk of code to be extraordinarily helpful for those times that I'm programming in a top-down fashion (scaffolding first, then fill in the concrete implementation).
What languages have you seen it in?
Some other examples include:
- Haskell: https://hackage.haskell.org/package/ghc-justdoit
- Scala: https://github.com/TypeChecked/alphabet-soup
- Agda: https://youtu.be/3U3lV5VPmOU?t=3159
I was under the impression that justDoIt comes from Scala, but I can't find a reference for that.
This is the worst possible argument against a language feature. Hideous code is exactly the kind of code that is most difficult to read, and so exactly the kind of code where you want the most help from your language.
Step 1 : Take a convoluted example code in another language/framework that you want to shit on, this example should break as many good coding practices as possible.
Step 2 : Show a simple example that adheres to good coding practices in your favorite language/framework.
// Old school
String string = new String();
over // C++ style
auto string = new String();
// Rust style
let string = String::new();
I don't see a single advantage to repeating the type, especially in the case where it's already apparent what the type will be. String s = doSomethingOpaque();A lot of people have preferences on this front. Somehow, these preferences get promoted to best practices if not moral license. Just, which approach is claimed to be better really depends on who you ask, and then mainly on whichever language that person first cut their teeth on. And nobody's going to be able to cite any well-designed study that demonstrates their preference is actually superior, because none exist.
auto string = std::string{};
(old school being just std::string string;
)NSString * string = [[NSString alloc] initWithString:@"myString"];
Why not:
String string = new ();I've never really struggled with this in C# where nearly everything is implicitly typed. Never been a problem in non editor code views ( like github ) either. One of the major benefits, not mentioned, is when you change the type of something it doesn't cause needless code churn.
There are only a few things where I've seen it be a problem is for new people to programming where even remembering all the basic types is still a significant cognitive load and templated types look like unholy magic. They'd struggle to know what var sum = 1 + 2 + 3 would result in and really want to know.
The other, in C#, is where you end up with a IEnumerable instead of a List. This can result in some not so nice side effects
I also make the point that often I’m reading code on GitHub or outside an IDE, which does not have such features.
Inferred types feels like an invention for people who don’t like IDEs, that somehow forces everyone _else_ to use an IDE.
Dictionary<string, KeyValuePair<int, string>> contrivedExample = new Dictionary<string, KeyValuePair<int, string>>();
And DictionaryOfStringsWithKeyValuePairs contrivedExample = new DictionaryOfStringsWithKeyValuePairs();
And var contrivedExample = new Dictionary<string, KeyValuePair<int, string>>();
I'd definitely prefer the latter.-- Edit: rephrased to be less presumptuous
Dictionary<string, KeyValuePair<int, string>> contrivedExample = new Dictionary<>();
Which seems ok to me? And we don't enter the slippery slope of allowing variables without types, and hoping people just use it where appropriate. for (var x : someMap.entrySet()) {
String key = x.getKey();
Entity value = x.getValue();
...
}
Here, the reader just needs look at the next two lines to understand the type of x. In the article, the key and value types were long and complicated, and they asked us to refactor. But in the above example, the types are very simple, yet still it's better than for (Map.Entry<String, Entity> x : someMap.entrySet()) {
String key = x.getKey();
Entity value = x.getValue();
...
}
I feel in this example, the type in front of x doesn't help me, as a reader.The benefit of type inference is that there is less to write, which makes changes quicker to make.
Type inference does not preclude adding extra annotations if they would help.
A good ide will show you type annotations at all times, if you choose. No small actions are required.
I would say that there actually is a problem with global type inference. If two types do not match (e.g. usage and implementation of a function) then most systems do a poor job of guessing which part (or both!) is broken.
At some point, if a module/pass/service is large enough, it becomes worth investing in a local naming convention that carries the relevant type information, and enforce that convention through code reviews. Module-specific hungarian notations, if you will.
I'm working on an inference pass in a compiler, and it's up to 2000+ lines of F# manipulating tuples/dimensions/tables/chains/indices, and identifiers referencing these.
1. each variable in the module can have one of these ten types, not immediately obvious from their names. While I can hover the variable and have the IDE tell you the type, it happens so often that it feels like hunting for an invisible cow.
2. a typical method will contain several variables related to the same concept. When inferring a "group by" statement, I will have an "origin tuple", from which I infer the "origin chain", from which I extract the current "origin table", and construct an "origin index", from which I discard duplicates to obtain an "unique index", which defines a new "unique table" for which it is the primary "unique dimension", and there is a corresponding foreign "origin dimension" on the "origin table". Some of these will also need to be indirected through identifiers.
In the end, there's a "hungarian.md" file in the module that explains that a `vec` suffix is a vector, `vecs` is a vector tuple, `vid` is a vector identifier, and so on.
More modern languages (Rust, Swift) have also gone and say "well, omitting types from arguments and return values of function signatures is a bad idea, so let's not do that."
For example:
var thing = getThing();
processThing(thing);
If you rename the type that getThing() returns, for example from "Thing" to "TheThing", without type inference the above code has to change (the type of "thing" has changed); with type inference no characters in the source code are changed.If somebody else makes a change to those lines, e.g. writes the whole block in an "if" statement indenting it, without type inference you now have a merge conflict you have to resolve. With type inference, merging works without manual conflict resolution.
This is a dumb thing that every programmer just accepts. The tool should do this work for you.
Thing thing = new Thing();
from my cold dead hands!Not once have I wished for type inference in Java for instance. Just learn to use the damn IDE and its code completion feature.