The advantages of static typing, simply stated
pchiusano.github.io
pchiusano.github.io
- Lets you create typed collections so that if you insert the wrong kind of value, you get an error immediately, albeit only when code runs, not before;
- Has abstract types that serve the same function as Haskell's type classes, and allow you to inherit a huge amount of functionality for a very brief type definition;
- Serve as documentation of APIs;
- Allow performance similar to C/C++.
Another advantage that wasn't listed is that in languages that don't allow type annotations, you end up writing a lot of boilerplate type-checking code in libraries to give better error messages.
The point is that you don't need a static language to get many of the benefits of types, you just need your language that allows you to express type-based properties. One of the way to do that is to have a set of formal rules for deriving the type of every expression in a program, but that's not actually necessary.
> For example, most of these advantages apply to Julia as well...
and then you listed it as an advantage.
It enables a whole class of highly effective program manipulations that are just unavailable to a non-statically-checked language.
* Changing the collection's type and not knowing until runtime -- when your program crashes -- that it's being used wrong.
* Changing the collection's type and the program refusing to compile until all usage is corrected.
Because of this, I don't agree with the most modifier you're attempting to apply. The difference to me is confidence. In the former, how confident am I that I didn't miss a use-case on some branch? Considering that the program will crash or otherwise fail to operate, that seems like something I want to be pretty confident about.
It's very useful to have something telling you that everything looks OK. That something can be testing, but why rely on writing good tests when we have formal proof systems available?
In D for instance you do have an auto type, yet it's still a type. An auto function that returns real or BigInt will trigger a compiler error. You need to convert the real to BigInt before any compilation can occur.
The advantage of static typing is obvious to me; the more my computer understands of my code, the less work I have to do. I can offload my memory and instead think about things that are more interesting, not be looking up Python obscurities on Stackoverflow.
But static typing is not perfect, and there is still more the computer could be doing if it had more understanding. That is where we should be focusing, that is where the real problem is.
I could say "python is better at understanding what the programmer wants to do". But I think you'd disagree with me.
The compiler is doing lexical analysis of the program being fed into it. That's it.
In case there are readers not familiar with the current state of things, you can pick solid languages today whose type systems support a LOT. It is completely fair (and good) to point out that this isn't a panacea, but it can significantly help.
There's nothing about dynamic typing which precludes this.
There are also things which it is not necessary for your computer to understand but the rigidities of an excessively strict type system will demand it be told those things anyway.
You know that they update the language regularly, and have a games sig? https://groups.google.com/a/isocpp.org/forum/#!forum/sg14
It could really use profiles. Strict mode or something, where the code should only use N "good practice" features.
So that's pretty much been that way since the Java plugin for Netscape Navigator. That started the plugin battle which morphed into the mess we're in today.
> Look at the game industry in 2016, where C++ is thrown around as the the cause and solution to all life's problems.
Software written in C/C++ often have LUA/Python/Perl plugins to add scriptability.
> The advantage of static typing is obvious to me; the more my computer understands of my code, the less work I have to do.
Nope. Computers don't understand code, they just processes it and point out where your errors are, there's very little understanding. Type checking is really just rattling off a checklist to validate the type being passed around is the type that's expected.
For example,
Like puzzle pieces with shapes that we can observe fit together, we can think of types as specifying a grammar for programs that ‘make sense’.
All programming languages have a grammar for programs that make sense (it's specified by the parser). Algebraic data types give you a grammar for values that make sense. Separately, a type checker assigns types to expressions and checks that they're consistent. You can have algebraic data types without a type checker (e.g. Racket's 2htdp/abstraction), and you can have a type checker without algebraic data types.
You can also have constraints on values that are too complex to be easily checked at compile time (e.g. clojure.spec) or that can be checked at compile time but only incompletely (e.g. Erlang's -type and Dialyzer).
Anyway, I don't mean to be critical of the article: it does a good job covering the high-level tradeoffs. I just think people are too quick to generalize their experience with particular languages to entire classes of language features.
In a statically typed language (the stronger the better), aided by a good IDE, I don't need to be constantly executing my code against data during development; My editor is constantly validating my code, and when it stops complaining, my code will work. And months later when I or someone else uses that code in another part of the system, they won't need to execute that code to see how it behaves, as the types themselves provide documentation and as-you-code feedback.
You mean your code will compile and run. Whether it behaves as desired is completely unknown without testing.
[Edit] To be fair, this is not only, and even perhaps not so much, due to the static typing per se but also due to the mental discipline the particular programming language may require from the programmer even to write code that can be successfully compiled.
I'll bet my Python code has a similar probability of working correctly on the first run, conditioned on the complexity of the task.
'course, it probably wouldn't work fast, and it was 50/50 whether I'd introduced space leaks, but...
Indeed, difference in complexity is the key here.
Having developed professionally in both Python and Haskell, the probability for me was significantly higher in Haskell.
For the past few years I've been coding almost exclusively in dynamically typed languages - I did some Python (dynamically typed) but mostly a lot of JavaScript/Node.js (dynamically typed). I understand all the pros and cons, but for me personally, I am much more productive with dynamically typed languages than statically typed ones.
It was a long time ago, but I still remember clearly when I switched from AS2 to AS3 - I was writing games for Flash at the time; I did feel an improvement and I really liked the additional structure which types brought to my code. There was a certain satisfaction that came with defining fixed classes and interfaces and making use of polymorphism and various formal 'design patterns'. It gave me extra 'confidence' in my code.
In retrospect, after having spent years praising statically-typed languages, and then later switching back to dynamically typed languages, I think a lot of the benefits that I felt during my static typing phase came down to one simple fact:
"Statically typed languages force you think more before you do things" - This was really valuable early in my career when I had a tendency to rush things. However, now that I more fully appreciate how complex programming is (and how easy it is to break stuff), I am always very careful (regardless of the language).
Static typing for me has become a tedious process through which I no longer derive much value - Though it was really useful at a specific point in my career.
That said, I think there is some stuff (anything to do with low-level hardware/systems and optimizations) where statically typed languages cannot be avoided.
Also I wouldn't say that people who like statically typed languages are inexperienced - I know some very experienced engineers who are just addicted to that extra feeling of 'confidence' and structure which statically typed languages give you.
Cant't help thinking that that's probably something like 90% of code out there...
The more pieces of the project that you didn't write or otherwise have low knowledge of their inner workings, the more dangerous modifying code becomes. It's very useful to have something telling you that everything looks OK. That something can be testing, but why rely on writing good tests when we have formal proof systems available?
This is a very important point that I've rarely seen talked about explicitly though I think all programmers develop a tacit sense of it.
On one hand, refactoring code written in a statically typed languages is less error-prone - But on the other hand, such refactorings tend to affect more code than those of dynamically-typed code.
If your business requirements change often and refactorings are common, it can be a pain to keep having to rethink your code structure.
Dynamically-typed code is often easier to extend and modify. I find that with statically typed code, if you start messing with a small part of your code, sometimes you have to rethink your entire class hierarchy. With dynamic languages, your code can handle quite a few changes before it gets to a point were you need to rethink the overall structure.
The author asserts that static typing allows the compiler to answer this question, but this only allows the compiler to spot type errors in advance. There are many other kinds of errors that are completely invisible to the compiler.
In a dynamically typed language, if I don't spot the error from reading the code, I must wait until runtime/testing to discover the error. This is also true for a statically typed language for every kind of error except type errors. Personally, type errors haven't been the kind of errors that haunt my dreams. I guess that's why I'm not enthusiastic about static typing.
Many people assume that 'types' are simply primitives like Int and String, and that a type checker just makes sure you don't pass an Int to a function expecting String. However, it is possible to express far more powerful statements about your data using a good type system.
For example, you can express the idea of non-emptiness of a container, as mentioned in the article. Then you know that, say, taking the max element of a non-empty container is guaranteed to give you an element, whereas with a possibly-empty container you might not have any element at all, causing a null, or exception, or at least requiring an Optional type.
You can express safety properties such as a sanitized string vs. unsanitized. You can have a Sanitized type that can only be created by calling a sanitize function - which carefully escapes/handles any invalid characters - and then functions that might, say, pass a value into an SQL instruction can be typed to only take Sanitized strings. Now the representation in memory of Strings and Sanitized strings is identical, but by using different types and a certain set of allowed functions on those types, you can encode the invariant that a string cannot be inserted into an SQL query until it has been sanitized. Now your type checker can catch SQL insertion vulnerabilities for you. How's that for a type error?
First, when most people talk about static typing, they're talking about the near-useless version -- just types like Int and String. I think we agree there, so I won't mention it further.
Second, a dynamically typed language like Python has more typing information than some folks first assume. Python's AttributeError is quite similar to a TypeError. In fact, with old-style classes (v2.1 and earlier), many errors that are now TypeErrors were AttributeErrors. Calling len() on an inappropriate object would raise "AttributeError: no __len__". In many cases where folks talk about wanting a static type system, they really just want interfaces.
The Sanitized string example is a good counter-point because the interface needs to be near-identical to a regular string. I'm not certain a more complex memory representation (caused by defining a different class) would cause noticeable inefficiency. We're probably not doing vectorized operations on strings.
This brings me to my third point, that Python 3 has a similar split between two types: bytes and str. The memory representation is slightly different, bytes vs unicode, but the interfaces are nearly identical. Two differences would be decode vs encode and that getting an element from bytes (annoyingly) gives an int. The distinction between the two types is enforced mostly inside builtin functions, implemented in C. This was a big deal, causing backwards incompatibility, many flamewars, and we're still resolving it, though I think it's clear to most people now that Python 3 is the future.
Is it possible that the Python 2/3 split could have been avoided if we had a static type system? Perhaps, if we had multiple dispatch, the function signatures could have remained the same, avoiding backwards incompatibility... I'm just speculating here. My guess is no, getting rigorous about unicode would cause incompatibility regardless of the type system. I'll get back to the main topic now.
> Now your type checker can catch SQL insertion vulnerabilities for you.
This sounds useful, but a good interface solves the problem just as well. I'm a Pythonista (if you haven't noticed), so my example is PEP 249 that specifies a DB API for all database wrapper implementers to follow. It states that it's the wrapper dev's responsibility to implement a sanitizing string interpolation for the cursor's execute method.
My conclusion is that designing a good interface is important whether you have dynamic or static typing. Static typing errs on the side of safety, dynamic typing errs on the side of flexibility. Both can mimic the other. Arguing that one is better is like saying linear regression is better/worse than k-nearest-neighbors.
I don't agree. Who is "most people"? Certainly not PL designers and not most of what I've seen here in HN. More importantly, it's also not what the article under discussion is saying, either.
> Static typing errs on the side of safety, dynamic typing errs on the side of flexibility. Both can mimic the other. Arguing that one is better is like saying linear regression is better/worse than k-nearest-neighbors.
In my experience, this isn't true. Modern statically typed languages have all the convenience of dynamically typed ones, such as REPLs and elegance, plus the safety of early warnings and the guidance that static types give you while writing your code (if you've ever written code like this, you'll know the feeling of working with building blocks that "fit" with each other). So you can have your cake and eat it, too.
Also in my experience, not having experience with these languages is what leads some people to think their type systems can only state trivial things such as "this is a String". They can do more. They can say things such as "this expression/function doesn't write to disk as a hidden side effect", which is useful!
Like a dynamic language with optional type hints? As I said, both techniques can mimic each other, with the corresponding tradeoffs. As you use more generics in a statically typed language, you're sacrificing safety. As you use more type hints in a dynamically typed language, you're increasing syntax clutter and decreasing flexibility.
Actually, in languages like Haskell, the more generic your type, the more "safe" you can expect it to be.
As an example, consider a function that gives you the first element of the tuple you pass to it.
The most generic type of this function is
fst :: (a,b) -> a
However, it can also have the type fst1 :: (Int, Int) -> Int
Now, you can be sure of the behaviour of fst immediately by looking at its type, but that doesn't hold for fst1Pretty much the only definition of fst that the compiler will accept is
fst (x,y) = x
However, the compiler will accept all the following definitions of fst1 fst1 (x,y) = x+y
fst1 (x,y) = x*y
fst1 (x,y) = 2^x
fst1 (x,y) = 7
...Like the other commenter says, generics actually increase safety: there are fewer assumptions (and therefore, incorrect assumptions, aka bugs) you can make when your functions are generic. Also, modern statically typed languages do not increase clutter by much, and can be very elegant and brief.
The point of strongly-typed systems is that you can represent your constraints as types. This takes extra thinking and work, but gives you almost almost unlimited expressive power (ref: agda).
Simple example: meters and feet as different numerical types. When you multiply them, you get a silly unit (foot-meters) that doesn't fit with whatever you wanted (meters^2), and thusly fails compilation.
I also wonder when you would make that distinction in the lifecycle of your application. I suspect not until you first encounter the bug of accidentally mixing units. If so, we'd be solving the problem at the same time, just using different techniques.
https://docs.microsoft.com/en-us/dotnet/articles/fsharp/lang...
I think if you're used to a language like Java or C++ you may not see what types can really buy you, but that's because most languages have bad type systems.
Types are far more rigorous - you can ensure, with types, that your code behaves in a certain way given any input. Yes, this extends to "design" bugs ie: not just crashes - you can write a type that ensures that an API can only be used correctly, for example. You can encode logic like "Don't allow unauthenticated users to access this content" into your type system - and you no longer need tests.
Of course, type systems don't prevent you from writing tests. They actually make it easier, you can generate test cases based on types, as an example.
I also have never met someone who claimed to never need tests because of a powerful type system.
Of course having great test coverage helps alleviate this, but it's very rare where a large project has 100% test coverage.
For example: a function that takes an integer between 1 and 100. That's a straightforward constraint! Elixir is an example of a language that lets you express some of those kinds of constraints, by using guard clauses. Along with its powerful pattern matching, you end up with very nice code.
What are other languages that are known for these kinds of constraints?
Consider the following: I have a "rating" type, which is a whole number from 0 to 5. If you try rendering a view with a number outside of that range, that's a bug! Which means you probably made a mistake somewhere.
I want software to help me catch bugs. If something can't be confirmed with a linter or compiler, getting a good error message during runtime is also fine. Once I've established that I expect some value, I don't want to write extra checks to confirm that my expectations are met.
Regarding your rating example, you could encode that restriction in a type (or class in OO) and have a guarantee that once you have an instance of that type you no longer need to check for it's validity. Rendering a Rating would never fail at runtime, you'd get a type error first.
I think this is not generally true and not a point against static typing, but rather against statically typed languages with poor support for generics. Java stands out as the poorest implementation of generics I have ever seen.
> For instance, a generic serialization library can be written in a dynamic language, without anything fancy, but providing the same thing in a static language requires more machinery, and is sometimes more complicated to use.
As a counterexample, have a look at NimYAML (my work):
http://flyx.github.io/NimYAML/
The examples there show how easy it is to provide the user with a generic interface for serialization in a statically typed language. The implementation does not differ much from what you'd do in a dynamic language: Provide a pair of serialization/deserialization handlers for each of [simple types (string, int, float, enums), array/sequence types, tuple/object/struct types, dict/map types, pointer/reference types].* Introduce new generic types in a new namespace. .NET did this with `System.Collections` and `System.Collections.Generic`.
* Default unspecified generics to `Object`, including usage of generic types in bytecode tagged with pre-1.5 versions.
I don't see how this can be true. Wouldn't there be less to specify if the programmer didn't have to specify types at all?
"A large class of errors are caught, earlier in the development process, closer to the location where they are introduced."
I would re-state this as:
"A large class of errors are introduced, which otherwise would not exist."
Consider dealing with JSON in Java. Every element, however deeply nested, needs to be cast, and miscasting leads to endless errors, elsewhere in the code. Given JSON whose structure changes (because you draw from an API which leaves out fields if they don't have data for that field) your only option is to cast to Object, and then you have to guess your way forward, figuring out what the Object might be.
Consider the Salesforce clone of Java, Apex, which I have had to work in this month.
The if() statements here are the same one's that I would have to write in Ruby or Python or PHP, but meanwhile I've had to do a bunch of other, useless work:
public Object deserializeJson(String sandi_data) {
System.debug(sandi_data);
Object objResponse = JSON.deserializeUntyped(sandi_data);
if (objResponse instanceof Map<String, Object>) {
Map<String, Object> mapResponse = (Map<String, Object>)objResponse;
List<Object> dataList = (List<Object>)mapResponse.get('data');
if(dataList == null) {
String err = 'The Sandi API field for data was null';
System.debug(err);
ApexPages.Message msgErr = new ApexPages.Message(ApexPages.Severity.ERROR, err);
ApexPages.addmessage(msgErr);
return null;
} else if (dataList.isEmpty()) {
String err = 'The Sandi API field for data was empty';
System.debug(err);
ApexPages.Message msgErr = new ApexPages.Message(ApexPages.Severity.ERROR, err);
ApexPages.addmessage(msgErr);
return null;
} else {
System.debug('dataList:');
System.debug(dataList);
return dataList;
}
}
return sandi_data;
}
And then, downstream of this: List<Object> dataList = (List<Object>)deserializeJson(sandi_data);
for(Integer i=0; i < dataList.size(); i++) {
Map<String, Object> dataMap = (Map<String, Object>)dataList[i];
System.debug('dataMap:');
System.debug(dataMap);
String response = fetchCompany(dataMap);
SearchResult__c profile = saveProfileResult(response);
cr.add(profile);
}
I'm leaving out the code that is downstream of this function, but it is full of more of the same: guessing at fields, guessing at how they should be cast, using if() to guard against null or empty. Tons of unnecessary bloat. Lots of easy errors to make.Again, some of the if() statements need to be made in Ruby or Python or PHP, but the rest of it is just pure bloat. Verbose, unneeded and unhelpful.
In a dynamic language I could simply work with a deeply nested data structure of maps and lists, and I'd handle the casting at the very end of the process. In a dynamic language, I could treat everything as a string till the very end, and then cast to integers or dates or floats or strings as needed. In a dynamic language, I could write the code faster, with less errors, and with less code.
Static typing does not live up to its promises.
[ Edit to add ]
We have no control over the API that we draw from. We are drawing from the API of a different company. I wish they didn't use JSON. If they have to use JSON, I wish they at least enforced a consistent schema. But they don't. And that is why static type checking fails: because the real world is chaotic, and when you have to interact with that real world, you are often forced to do so dynamically, because of the mistakes that other companies have made. The real world is dynamic.
The idea that you can know an external API perfectly is a fantasy. The real world is messy. The real world does not always conform to a strict schema.
The notion that An External API Is Reliable is as stupid as the notion The Network Is Reliable:
https://blog.fogcreek.com/eight-fallacies-of-distributed-com...
[[ Further edit to add ]]
the_af wrote:
"Using dynamic typing will just hide the problems under the rug, and they will explode in your face later on. Static typing just made those problems explicit."
What I wrote was:
"In a dynamic language I could simply work with a deeply nested data structure of maps and lists, and I'd handle the casting at the very end of the process"
I'll simplify this: there are 3 times when we can enforce a schema:
1.) when the API call returns with a string
2.) on every line, scattered through dozens of functions
3.) at the end, when I have the data that I want
In my original comment, I advocated for #3. Here are the reasons I don't like the first 2 options:
#1 - the external API is bloated, so writing a schema for the whole thing would be difficult to justify in terms of business. We only need a tiny slice of the data.
#2 - having casting discovery information scattered through dozens of functions makes the code brittle and refactoring difficult.
With Ruby or Python or PHP or any dynamic language I have the option of #3: grab the data, cast everything as a string, grab the tiny sliver of data I actually need, and then enforce the schema on that tiny sliver. This is the data that I can cast to integers, floats, dates, etc -- whatever is actually needed.
In static-type languages such as Java, I'm forced to go with either #1 or #2, and they are both bad options.
About this, from tigershark:
"it is only the usage of an awful JSON library in a not so nice language"
Bad JSON is part of the real world. If your static-type language can not handle bad JSON, then it can not handle the real world. That is my point: static-type checking is too academic, too pure, for the real world.
As to "not so nice language", you are engaging in the No True Scotsman fallacy, which goes like this: no True statically typed language would be this bad! But following the No True Scotsman illogic, the rest of your unstated assumptions amount to: It's only the statically typed languages that most programmers actually use that are this bad! But somewhere there is a statically-typed language of such unbelievable purity, it overcomes all of these problems!
Furthermore every dynamic language I've used has had tooling spring up, e.g. in the form of linters, that includes at least some basic checking for type mismatch where possible. I'm not so sure static typing leads to problems. It isn't problem free--no programming language or paradigm is--but I do think the lack of it leads to some problems that are easily avoided without too much extra overhead.
Now, instead of fixing problem by enforcing some schema you propose to use dynamically typed programming language which won't solve the problem but instead will make it spread thorough the whole codebase.
At least static typing limits the damage to the code that directly deals with parsing JSON.
Although I'd agree that I find development with dynamically typed languages a bit more chaotic than static or strongly typed ones. And not necessarily in a fun, good way, either.
In Java JSON is horrible for serialization because of Java behaviors (your best option is Object). That doesn't mean JSON is horrible for serialization, in general. Fields do have types (int, string, object, array) making it slightly better than arbitrary serialization.
As it happens, I agree that JSON is probably a long-term mistake, although it is definitely an improvement over XML, the previous player in that space.
There are less awful ways to deserialize JSON in other languages. In C# with JSON.Net, you have options ranging from very dynamic-style handling, all the way to fully-typed deserialization into POCOs, for example.
[1] I have no love for Salesforce; it's a mess. Their weird custom dialect of SQL for a REST API query language is also buckets of fun.
If Java were conceived of today - I think it might look a little different with respect to this.
I'd like to point out that there are many areas wherein fluid typing might help a little bit, especially on the UI.
Have to make a class for 'every little thing' gets cumbersome.
I wish there was a 'lighter typing' opportunity in some cases (schema-ish JSON).
But I agree that in the long run, static typing is really the way to go for the most part.
It's a common challenge for Elm beginners.
It does the job but it's not nearly as nice as javascript or python. It's about double the code in Java and it doesn't read nicely.
I think it depends on what you're doing.
EDIT: also, java is bad at dealing with json if the api could return data in multiple different schemas and you don't know which one it is until you parse it but I've only run into this once and it's just bad API design.
Developing/experimenting/scripting is where it's painful.
The problem you're pointing out is not a problem with statically typed languages. It's a problem with bad languages.
type Simple = JsonProvider<""" { "name":"John", "age":94 } """>
let simple = Simple.Parse(""" { "name":"Tomas", "age":4 } """)
simple.Age
simple.Name
and most importantly you don't have nulls.But this is hard with any typing system. Using dynamic typing will just hide the problems under the rug, and they will explode in your face later on. Static typing just made those problems explicit.
There are many benefits to static typing, beyond the external boundaries of the system, where it definitely lives up to its promises.
All the research from expiercal studies mostly all show that static vs dynamic languages all correlated with a Expertise reversal effect. This is for beginners not experience programmers. That at the start of the studies for the first 8 to 9 months novices have a negative correlation with the speed of development work with a statically typed language. After 8 to 9 months they gain a singificant boost from static type language.
Source to backup my statement and resource if you're interested @ Susan_hall
Functional Geekery Episode 55 – Andreas Stefik https://www.functionalgeekery.com/episode-55-andreas-stefik/...
Secondly another study that was published for Polymorphism in Python from the website (neverworkintheory)
Polymorphism in Python http://neverworkintheory.org/2016/06/13/polymorphism-in-pyth...
Gives the qoute: `Our findings show that the receiver in 97.4% of all call-sites in the average program can be described by a single static type using a conservative nominal type system using single inheritance. If we add parametric polymorphism to the type system, we increase the typeability to 97.9% of all call-sites for the average program.`
I could qoute other resources to backup my statements though if you want to have a skype coversation I'm up for a talk.
data DataMessage = DataMessage { company :: Text
}
deriving (Show)
$(deriveJSON defaultOptions ''DataMessage)
data DataList = DataList { _data :: [DataMessage]
}
deriving (Show)
$(deriveJSON defaultOptions { fieldLabelModifier = drop 1 } ''DataList)
Which will fail properly when the data is malformed and in the real-world is even shorter, because you can set your data-type to protocol naming conventions and share them throughout.Then there isn't much bloat. You wrote your types and generated your way to parse them.
You could test with:
main = do
print (encode (DataMessage "some-company"))
print (decode "{\"company\":\"some-company\"}" :: Maybe DataMessage)
print (encode (DataList [DataMessage "c1", DataMessage "c2", DataMessage "c3"]))
print (decode "{\"data\":[{\"company\":\"c1\"},{\"company\":\"c2\"}]}" :: Maybe DataList)
https://gist.github.com/b30f6f09a737dcc980e052b0f3d2a39eBut these aren't needless errors. If something is not of the type that you expect, there should be an error, and static typing makes this more explicit. The verbosity is more to do with that particular implementation, not anything inherent to static type systems.