The Unreasonable Effectiveness of Dynamic Typing for Practical Programs [video]
infoq.com
infoq.com
Fear not, my fellow Ruby/Python/Lisp/Perlers - the pendulum will swing back, and all that will be discussed is how much of a pain the current static typing systems are and how type hinting is the real future.
He admits that the papers he cites are too few to make large decisions and that the analysis he does by looking at how many issues on GitHub are actually caused by type system errors (1-5%) but he's trying to steer away from dogmatism in favor of empirical evidence.
I don't think we have that same effect in software but I think there is an related one: when you don't have the protection of a static type system, you may become more careful.
It's an interesting theory. On the other hand, the amount of time and money lost to null pointer bugs over the years suggests that we still aren't careful enough.
You could use that reasoning to argue we should be writing in assembly. I want all the automated checks I can get to get as close to bug free as possible. If you have to be careful about everything, you're going to exhaust yourself and miss even more bugs. It's much better when bugs are limited to smaller areas and you know where to be extra careful.
Having to be careful because there's no type system checking your work is a useful skill to learn in some contexts, but I think the current state of the software industry is a pretty good argument that human programmers will never be careful enough. So, we should do what we can to accommodate human fallibility and use machines to check our work early and often.
A static type system also requires continuous attention. The question is, does it free attention faster than it consumes it?
I'm not sure what you mean by this, or how it constitutes an increase over the demands of a runtime or weak type system.
> how it constitutes an increase over the demands of a runtime or weak type system
It's the difference between "I understand X", and "I have to figure out how to describe X in my language". And the difference between "my understanding changed", and "I have to figure out all of the places in my program where I need to update the code because of my new understanding".
I would not call that "continuous attention", because most of the time you're simply writing down a decision you've already made, and would have had to make even in a dynamic language.
In contrast, I think "continuous attention" is having to remember the current kinds of data being held in different variables in a long data-flow.
> And if your needs for that variable or class or function change, you need to explain to the compiler the differences.
Yes, but that's because the compiler is remembering your "checklist" for you. In a dynamic language you'd need to make all the same decisions, but you'd have to waste more brainpower (or introduce more bugs) by trying to manage it all yourself.
I'll be honest, I've never wasted much brainpower trying to remember the difference between covariant and contravariant, or whether I want Try[Future[String]] or Future[Try[String]], or anything like that, when working in a dynamic language. You just don't do shit like that.
Part of my job is maintaining a PHP/JS codebase that's 10+ years old, and I wish I had days where it was all about covariance, rather than the stupid stuff.
The presented studies can hardly be used as a basis for making his arguments. The study where a student invests a new language that nobody else could possibly know is especially bad. Of course the learning curve is going to be longer for an invented statically typed language! You cannot assess productivity without factoring out the initial learning time. If anything, that study only shows that dynamically typed languages are easier to get into.
However, funding the development is sometimes harder than funding the maintenance: You need the system to operate in the first place in order to get the benefits/money/revenue/etc.
Once the system is in place and brings benefits, it is easier to fund further investment on the system, be it maintenance, extension,etc.
Try reading your own code written two weeks ago. How about a month ago? Three month ago?
Let's say your code is making some calls and gets back an object. Can you list all the important methods on that object from the top of your head? With a statically typed language IDE can always do that for you. If your team member added some more methods, they will be listed too. No need to waste time reading the function source to figure out the return type, and then waste more reading the class's source.
In fact i did. Code of 8 months ago (december 2017). Easy to read, because i used named parameters extensively and added comments at every function Signature. Language Python (dynamic).
Don't assume that all programmers can't understand their past code. All my peers (we have over 12 years doing this professionally) would do the same. It is part of our job to read code written even years before by us. You never know when a past customer (internal or external) will call you to want an enhancement or upgrade.
> With a statically typed language IDE can always do that for you.
So? Not everybody needs to use a crutch to walk. How about having to patch a live code on a running server using only a plain text editor? This is not fiction, this is real. (Servers almost never have any kind of IDE or even fancy text editor installed).
> And you will have to go back and change your code many times while using your initial funding.
I am speaking about production software for using within an enterprise context. You have a defined budget and you need to ship a working product. For us , "maintenance" phase starts after the product has gone live in production. For us, if we deploy a working product, it is easy to get additional funding for maintenance and improvement. If we don't, we don't get subsequent funding for anything. Business will value more a slightly buggy program that is ready in 3 months than a perfect program that only is deployed after 8.
This is not a sound argument to advocate dynamic languages. There are good reasons industry best practices include things like code reviews, testing, and staging environments. But I suppose, if you are going to ignore all of these, then the lack of type safety is the least of your concerns.
> There are good reasons industry best practices include things like code reviews, testing, and staging environments. But I suppose, if you are going to ignore all of these, then the lack of type safety is the least of your concerns.
You are doing a strawman fallacy. I never ever spoke about having or not having code reviews, and separate testing and staging environments. All of these are also necessary, and what in the world does this has to do with static typing? Testing and staging environments are required and doable using any programming language.
That's really an issue with class-based programming, where you have lots of ad-hoc methods, you have identified one of the complexity it brings. You need the bloaty IDE to make that problem bearable to code in, because as you said, who is going to remember all that crap?, but it doesn't make the problem/complexity of the system go away.
From my personal experience, the completeness and quality of the tests has much more impact on maintainability than the static-ness of your language. But to be fair, I haven't seen a study proving this either.
[0] http://haskell.cs.yale.edu/wp-content/uploads/2011/03/Haskel...
My own experiences learning Haskell after working Python for more than 10 years suggests this is quite plausible.
Static type declarations can help program understanding a lot without no effect on performance. But on the other dynamic assertions of properties of objects can be more flexible, you can program your own type-checking algorithms and define your own types and what it means to be a member of them.
However advanced type systems like the one in Elm, now that's something else. Especially when refactoring and almost-writes-it-for-quality error messages from the compiler.
Just FYI: If one wants his/her "programming language" to catch typing errors/violations/mismatches, one needs to use a language with strong typing, regardless of if the language itself supports static or dynamic typing.
Dynamic typing relies on being able to infer the types; this can be often achieved at compile time but in any case at runtime, if the platform is strongly typed, any type error will be caught. Then, after being caught, what could be done about it depends on your platform -- some platforms will die with an error; some others will die but allow you to inspect what happened; then others allow you to patch (correct) the error and continue.
Static typing requests you to explicitly declare types. This helps the compiler do as much checking as possible at compile-time. This has the added side benefit to increase the speed of the compile code, since less checking is needed at runtime.
Some systems are dynamically typed but weakly-typed. Examples are PHP and ES/Javascript.
Some systems are statically typed but fairly weakly-typed. The Kerninghan&Ritchie C language is an example.
Some systems are dynamically typed and strongly typed. Python is one of them, although it also has what is so-called "duck" typing.
Some systems support dynamic and (optional) static typing, and are strongly typed. Common Lisp is an example, as well as in Julia. In both cases (CL and Julia), static typing helps for increasing performance, and for achieving multiple dispatch (i.e. for Object-Oriented programming). CL is able to do compile-time checking as well, if the implementation supports it.
Some systems have fairly complete strong typing which relies on extensive static typing facilities. Haskell is an example.
Do you identify that problem in this specific video?
"Strong typing" is itself loaded. There is a definition of "strong typing" which means "no conversions occur between types without programmer requests". For instance, under this notion, a language in which we cannot use an integer expression in a floating-point context without putting in an integer-to-float coercion is more strongly typed than one in which the conversion is inferred and takes place implicitly.
"Strong typing" has been used in reference to the facility, as in the Pascal family, for declaring type alias names for identical types (such as integers) such that the aliases are considered distinct types. The C typedef mechanism is weakly typed because a mode_t and pid_t, both perhaps just aliases for int, are considered to be indistiguishable from int and compatible. C++ has more strongly typed enum types than C, because a location of some enum type in C++ cannot be assigned a value of an integer type, or of a different enum type, without a cast.
Sometimes "strong typing" means that there are no holes in the type system: no datum is being misused if there are no diagnostics about type. (This is related to the other senses of "strong typing" because if an implicit conversion is allowed in a program which is free of type diagnostics, that conversion could be unsafe and hide a software defect. The type system knows very well that here there is an expression of real type and the context requires a integer value and silently generates the conversion; but what if it is out of range? In, C, the behavior of an out of range floating to integer conversion, and vice versa, is undefined.)
> Static typing requests you to explicitly declare types.
Not in anything that can be called state-of-the-art static typing, except at module boundaries (where the compiler cannot always infer parameter or return types since some of the information is in another module that it doesn't know about).
Types can be inferred implicitly from basic terms in the program syntax tree whose type information is known (like literal operands, standard functions and operators and such) and then from there by inference (which can also go in the other direction, like from known function parameters to unknown terms nested inside the function). https://en.wikipedia.org/wiki/Hindley%E2%80%93Milner_type_sy...
I have never heard this. I do, however, use "sound" and "soundness" for virtually rhe same thing. The reason why is because both of these terms are worthwhile as distinct, meaningul, well-defined terms.
It also doesn't make sense: type correctness is a binary calculation. Strong is a comparative, which isn't super useful when there are only two values.
A concern Smallshire spesifically addresses in the talk, pointing out the correct distinctions.
Normally it's hard to tell when somebody hasn't read before commenting, but given that a non-trivial portion of this talk is explicitly laying out the different facets of type systems and doing compare-and-contrast analysis, it's really obvious that your reply was written without watching.
From the HN guidelines: "Please don't insinuate that someone hasn't read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that.".
It's very probable and totally OK that the parent didn't watch the video. Or that they watched it but skipped around most parts. It's a 50+ minute video, and it's not required to watch it in full for participating in this HN discussion.
People often get into the discussion inspired only by the title of the submission (which serves as a topic for the thread) and not the content in the linked article or video. That's just as well -- more often their contributions, even without having seen the article, can be way more informative and intelligence than what's in the article. And yes, sometimes they can even repeat what's already said there.
I'm letting you know your mistake so that you can avoid being downvoted next time.
My post was about "discussions" and targeted to the comments that usually come whenever static typing vs dynamic is discussed.
Not about the video.
https://news.ycombinator.com/item?id=15083003
As chriswarbo said:
> One reason for this is the tight coupling in developers' minds between what I'd call "types" and "representations" (these terms are very overloaded, so others may use them in different ways).
> Just because, say, a function name and a string are represented in memory the same way, that doesn't mean they are the same type; in particular there are many strings which aren't function names, and there are many operations (e.g. append) which make sense for strings but not for function names.
This is a well-known problem, called 'stringly typing': The language is enforcing the type system it knows about, but it only knows the representation of values, not the semantics, not what it is valid to do to them at a high level. Therefore, you can defeat the type system entirely by just passing everything as strings.
I brought up the idea of having layers of representation:
> Length is a type of value, whether it's expressed in inches or centimeters or light-seconds is a representation, and whether it's in ints or floats or strings is another layer to the representation. You can add inches to centimeters with the right conversion, much like you can add numerical values represented ints and strings with the right conversions. The conversions just have to be at the right layer of representation.
This, to my mind, means that autoconversion is orthogonal to strong typing, in that it preserves the type system if it's done in a way which converts between different representations of values with the same (or compatible) semantics. Of course, if your type system is focused entirely or almost entirely on representations, autoconversion defeats it totally.
Finally, the little historical note:
> Your ideas sound a lot like the original Hungarian notation, BTW: If your language-level types are representations (as in, your type system says int and float, as opposed to semantic notions like pixels-from-edge or alpha-percentage) you can encode the real type information in variable names. People mutilated this to encoding language-level type information in variable names, which is utterly pointless and potentially harmful.
You are confusing explicit typing with static typing. "Static" means "compile-time checking". "Dynamic" means "defer checks to run-time". Some language implementations can disable run-time type checking if optional type annotations are added, such as in Typed Racket.
In other words, all the following combinations are possible:
- Implicit + dynamic + strong: Python
- Implicit + dynamic + weak: JavaScript
- Implicit + static + strong: Standard ML without type annotations
- Explicit + dynamic + strong: Elixir with type annotations
- Explicit + static + strong: Java
- Explicit + static + weak: C
I struggle to come up with examples for these... - Implicit + static + weak
- Explicit + dynamic + weak
... but I'm sure they exist, perhaps as historical languages from when type systems were less sophisticated. After all, "strong" and "weak" are entirely relative, constantly in flux, and exist on a continuum. In a couple decades, anything lacking dependent types might perhaps be considered "weakly typed".It is not really about checking, it's about allocation/binding. Static typing, for binding a variable, requires declaring the type of the data, this can be explicit (as in C++) or inferred from another value where the type was declared (as in C++ with 'auto' or as in Haskell).
Dynamic typing does not require a type declaration. You can bind an unassigned variable and then assign a value later, without having to declare the variable type.
What you're talking about seems to be more related to how type systems are implemented, i.e. that dynamically-typed languages are "unityped", such as the one and only "static type" of Python being the PyObject. Thus, when you "bind an unassigned variable and then assign a value later", you're actually implicitly declaring the type to be PyObject.
But it is deceptive to define type systems purely in terms of this implementation detail, because then it would appear as if dynamically-typed languages are untyped (i.e. "everything is one type" is the same as saying "there are no types"). Since we know languages like Python do some type-checking at run-time despite everything being a PyObject, it is more useful, I think, to think of these languages in terms of the practical result rather than in terms of the implementation.
But if that's not what you're saying, I apologize.
tl;dr:
A bunch of stuff about typing that you probably already know and for which there are better resources.
He cites a couple studies showing that static typing doesn't improve productivity.
His own analysis of python projects on Github shows that only 3% of bugs are caused by type errors.
He acknowledges that the evidence is not good and thinks that the industry should invest in doing better studies.
And it doesn't even come close to stopping there...there are hundreds of examples of errors that are not Type Errors in your dynamic language that are type errors in some other language. NullPointerException? That's a type error that is entirely prevented by many type systems. If you accidentally multiply your time and distance instead of divide to get your velocity, that would be a Type Error in F# using their units of measure types. If you accidentally pop an empty stack, that might be an EmptyStackException even in many statically typed languages, but would absolutely be a Type Error in a language with dependent types like ATS or Idris.
In the end, dynamic typing is just a way to make runtimes slow and building correct programs slower, in exchange for some entirely subjective aesthetic or ergonomic quality. And sure, most type systems in practical languages aren't "complete" enough to eliminate the need for tests. But don't tell me that because you don't eliminate tests that you don't need types. That's bullshit. Statically enforced types eliminate errors, and every check that you defer to runtime is an error waiting to happen. If you're willing to make that tradeoff, by all means do so...but don't lie to me.
Static types prove the absence of some errors.
Any type snob will tell you that even in supposedly good type systems, without dependent types, all lists are dynamic! Yet most of us that swear by static types don't use dependent types (much). There might be a limit where "more types" don't pay off their added cost relative to the thinking and typing needed and the security gained. Where that limit is depends on the application. Formally verified programs aren't used for many simple web backends - for good reason.
Also, type systems aren't equal. There are decent dynamic ones and terrible ones. Same goes for static type systems.
Of course, you can easily do things in a dynamic language that are observably correct, but which are very difficult to prove correct without advanced logic (i.e., an expressive type system). And more advanced type systems typically lose some nice ergonomic properties of simpler type systems, such as clear error messages, type inference, principal types, or decidability. So I can definitely see how someone moving for example from Java to Python would find the latter very liberating at first.
On the other hand, you can argue that people shouldn’t be writing code that’s so difficult to prove things about in the first place. :)
To my mind, the benefit of static types is that they allow proving properties of your code, but the cost is that they may force you to do so; the benefit of dynamic types is that they don’t force you, but the cost is that they don’t allow it.
On balance I think static types are therefore preferable, because you can always fall back to dynamic types (strings/hashes/variants/sums) in a static language, but you can’t “fall forward” to static types in a dynamic language.
Saying they don't allow it is too strong (they allow it, but it might be harder).
That testing can be done before runtime in some languages, e.g., with Perl’s BEGIN/UNITCHECK/CHECK/INIT blocks. But it’s not really the same.
In a referentially transparent, statically typed language, if you give me a function of type ∀a. a → a, then I know for a fact that if this function halts, it must be the identity function. The only thing it can do is return its input.
The same function in a dynamically typed language could return any value of any type, because its result’s type can depend on both the type and the runtime value of its input, or not depend on its input at all, even if it’s referentially transparent.
Even if you supply that type annotation and check it at runtime (“this function returns a value of the same type as its input”), it would only be checking a single code path, and it could still give me any value. That is, a valid implementation would be “return 3”—I wouldn’t get a type error if I only ever called it on integers, but it would still be non-parametric.
(defun foo (x)
(declare (integer x))
(round (/ x)))
(describe #'foo)
Derived type: (function (integer) (values integer rational &optional))
When given an integer, the function returns an integer,
as well as an additional rational value. ROUND's
specification says the type of the remainder is a REAL,
but since I call it with integer, it is more precisely a
RATIONAL (a subtype of REAL). (foo 0)
=> arithmetic error DIVISION-BY-ZERO signalled
Note, the meaning of the derived type is: if the code returns a value, then its type is .... The same applies in OCaml, where you can raise exceptions.Let's restrict the input domain to non-zero integers:
(defun foo (x)
(declare (type (or (integer 1 *)
(integer * -1)) x))
(round (/ x)))
(describe #'foo)
Derived type:
(function ((or (integer 1) (integer * -1)))
(values (integer -1 1) (rational -2 2) &optional))
The primary return value is an integer between -1 and 1; the remainder is a rational between -2 and 2 (ranges are inclusive).I consider gradual typing (e.g. Typed Racket, Hack) an interesting case of static typing where you’re trying to provide good interop with your dynamically typed surroundings, typically while migrating from dynamic to static types. Not all dynamically typed systems that allow type annotations have “gradual typing” in that sense, because many of them don’t give you enough power to get to “fully annotated” code (where the compiler is finally free to erase types completely).
No, (∀a. a → a) is not easily expressible (unless you write your own DSL).
You'd work around not being able to express (∀a. a → a) by focusing on the specific use scenario in the optimized code where that function is being called, and the concrete types that are involved. The identity function fits this type. If we know that x is fixnum, we can do this: (the fixnum (identity x)). Using the, we an assert types of individual forms.
Which is not to say that an implementation cannot do type inference for (identity x) and represent it type as (∀a. a → a) ; it's just that type inference is a separate thing from the ANSI CL declaration system.
Good example.
Really?
Straight from my common lisp (SBCL) command line:
CL-USER> (defun myfunction (x) (atanh x))
MYFUNCTION
CL-USER> (describe #'myfunction)
#<FUNCTION MYFUNCTION>
[compiled function]
Lambda-list: (X)
Derived type: (FUNCTION (T)
(VALUES
(OR (SINGLE-FLOAT -1.0 1.0) (DOUBLE-FLOAT -1.0d0 1.0d0)
(COMPLEX SINGLE-FLOAT) (COMPLEX DOUBLE-FLOAT))
&OPTIONAL))
Source form:
(SB-INT:NAMED-LAMBDA MYFUNCTION
(X)
(BLOCK MYFUNCTION (ATANH X)))
; No value
CL-USER>
What happened here?
I defined a function, i never declared which type does the function use, but the compiler can infer all the types.If i supply a string to "myfunction", the compiler will complain. Typing is strong, but it's dynamic.
What will atanh return?
CL-USER> (describe #'atanh)
#<FUNCTION ATANH>
[compiled function]
Lambda-list: (NUMBER)
Declared type: (FUNCTION (NUMBER)
(VALUES
(OR SINGLE-FLOAT DOUBLE-FLOAT (COMPLEX SINGLE-FLOAT)
(COMPLEX DOUBLE-FLOAT))
&OPTIONAL))
Derived type: (FUNCTION (T) *)
Documentation:
Return the hyperbolic arc tangent of NUMBER.
Known attributes: foldable, flushable, unsafely-flushable, movable, recursive
Source file: SYS:SRC;CODE;IRRAT.LISP
; No value
CL-USER>
The wonders of CL... not only i get all the possible input types, but also the return types. And the documentation, and the source code file(!)Your fallacy is saying: there is a whole class of errors that can avoided if only we use a different type system, so there is no point in using that type system which defers (some or all) checks at runtime.
I can write code that manipulates values which represent types in the language, at runtime. If the types happen to be known or declared during compilation, fine enough, they can be optimized away. Otherwise, they will be checked later (which transparently invokes dynamic dispatch if needed), and I can even use them once they are known to compile the code myself, on a closure; there, along with other values which are then known, they can be taken into account to perform more aggressive optimizations. Why don't all language provide this?
As for bugs or errors, if you cannot lower risks (with static guarantees) then you have to improve on fault tolerance. You take other approaches to ensure the system behaves well, even if some parts can fail (e.g. Erlang).
To be honest, I really don't understand what the benefit of dynamic types are supposed to be. People say it's easier and less hassle to get something up and running because you don't need type annotations but this isn't an issue with decent type inference and all you're doing is replacing compile time checks with runtime crashes.
Do you have any data to back this up or have links to papers that try and do a better job of counting issues that statically typed languages catch?
When you're re-coding all the time, the upfront cost isn't amortized over future runs as much as it used to be. Libraries are, but the massive real-world crowd-testing (i.e. usage) catches them...
BTW I really like units of measure types, you get dimensional analysis almost for free. You can do it in e.g. java but it's verbose (i.e. it's java). Classes, with arithmetic methods (add, sub, mult etc), You need an ezplicit class for each unit combination (though you can do a little better with generics). And get compile-time type errors when you assign, compare or operate incorrectly.
Ok, so you are writing an algorithm for a fast approximation of ATANH using IEEE floating point standard.
In type: IEEE floating point number Out type: the same.
Tell me how the magical type checking of the excellent programming language you propose will ensure that my calculations are correct.