Dynamic types have the potential to be more than "no static types"
buttondown.email
buttondown.email
The key problem has been that the resulting programs are fantastically hard to understand. One of the things I think Dijkstra got right in "Goto Considered Harmful" is that humans can reason about programs better when there's a strong correspondence between the textual structure of the code and its execution structure at runtime.
Runtime metaprogramming deliberately upends that.
I think it's probably totally fine to use those kinds of features as debugging aids when hacking on a program and figuring out where stuff is going wrong. That aligns with many of the use cases Hillel lists here. But you probably don't want to commit code like that to the repo and require everyone else maintaining the program to understand it.
For better or worse, we seem to lean strongly towards languages and language features that are used not just during the development process but also make sense in committed, maintained code.
Dynamically typed languages lean in the other direction. They can be very powerful and productive while you're in the middle of working on a program, but the resulting artifact is harder for others to understand and maintain.
Dynamically typed languages to me feel like unbaked plasticine. You can have a lot of fun molding it and you can get it into the shape you want much quicker than you could carving wood or stone. But if a dozen people are collaborating on a sculpture over months or years, it's impossible to build on top of someone else's plasticine part without it shushing down and falling apart.
Or course, the holy grail is baked plasticine: a language that is dynamically typed and moldable while you're figuring stuff out and then can be baked with static types for the parts that are solid and need to be reliable. I've yet to see an optionally or gradually typed language pull that off well. The boundary between them just gets too weird.
https://dl.acm.org/doi/10.1145/358027.358043
It was a joke because of course it was obviously worse than GOTO. But much of modern programming seems to consist precisely of COME FROM style constructs. A common term is "inversion of control", which can also mean "You don't have control".
As with metaprogramming, it's powerful in that it can take control of your system without your direction, but dangerous in that it's hard to predict when it will. I find that the hardest thing to debug is when things don't happen, when it's out of your control to make them happen. There's no place to put a debugger on "never happened".
Dynamic types, similarly, are hard to debug because they fail by not happening. Errors aren't caught, but there's no instant at which it wasn't caught because it was "any time".
1978 actually: http://www.modell.com/Magery/SPharmful.html
But that was a followup from a much older paper, from 1973: https://web.archive.org/web/20180716171336/http://www.fortra...
And to make you feel really old, "Go To Statement Considered Harmful" dates back to 1968. The context was the debate on structured programming which mostly happened in the 60s and early 70s (though it has flareups unending, they tend to be limited in scope).
Back in the Slashdot days I told people inversion of control means "in Soviet Russia, API calls you!"
This resonates very strongly with me. I've been doing a whole bunch of side/toy projects in Rust for months, and I can attest that even compile-time metaprogramming can make programs "fantastically hard to understand". Runtime metaprogramming makes that problem an order of magnitude worse.
Don't get me wrong, I'm not against metaprogramming, either compile-time or runtime. But it's way too easy to get overly enthusiastic when doing it.
The IDEs haven't quite embraced this concept properly yet, though: they usually have distinct "editing", and "running" states, but what's really needed is "editing", "building", and "running", with debugging available in the latter two.
When I briefly worked on a Ruby on Rails app in 2015, my first ticket was to investigate a crash that had been reported. I quickly reproduced the issue and got a stacktrace that ended in a function call to "is_readable", so I said: "No problem, I'm just going to grep for that identifier to find the definition." Cue two hours of utter frustration until I learn that you can use `.source_location` on a function to get a file and line where it is defined, which pointed me to a pile of metaprogramming held together by strings and class_eval. To no one's surprise, I fully support the statement quoted above.
But I had previously had experience with OO languages - Java and C++, and they often had similar challenges, though not quite as pronounced, along with problems of their own.
Eventually, I came to realize statically typed functional languages were the panacea. Sum types offered better composition than OO, and they had a good trade off between power/succintness (coming from higher order functions) and program understanding. In particular, the problems from other languages: what code is this, what type is this, where did this nil come from, were solved.
What's weird about ts with js files allowed?
But it took possibly the best living language designer on Earth to do it, the type system is (deliberately) unsound and will never be able to be used for optimization, and the type system is just incredibly complex. Literally a Turing-complete language running in the compiler.
And that's the least bad that anyone's been able to come up with. I applaud the TypeScript folks for doing it, but I think it's a worse user experience than most other modern statically typed languages unless you happen to be sitting on a pile of JS code already.
It's like plasticine with wires through it. Sure, the wires help but it would have been better to not start with plasticine in the first place.
http://journal.stuffwithstuff.com/2010/11/26/the-biology-of-...
Professor Sussman mentions this in this talk: https://youtu.be/HB5TrK7A4pI
Opinions on Coalton in Common Lisp?
There are perfectly valid use cases for these things, which do not make it harder to understand the code. Properly done, they might even make the code easier to understand. One example is writing a "timing function" (a macro). WHy a macro? Well, because you cannot call (time expr ...) when time is a normal function, because of order of evaluation. You could put the expressions into lambda and reduce readability that way, or you could make time a macro and get rid of the lambda wrapping of the expressions.
It happened thirty years ago:
https://en.wikipedia.org/wiki/The_Art_of_the_Metaobject_Prot...
But the thing that really drives me nuts about this never-ending debate is that static and dynamic are not a dichotomy. You don't have to choose one or the other. You can have both.
The real choices are:
1. Do you want the language to require the coder to provide type information (e.g. Java), or permit the coder to provide type information (e.g. Common Lisp) or forbid the coder from providing type information (e.g. early versions of Python, Javascript).
2. How much work do you want the compiler to do with regards to finding potential problems at compile-time, and what do you want it to do when it finds one? Options include: do nothing (e.g. Javascript), warn if it finds a problem (e.g. C), or forbid you to run the program if it finds one (Java, Haskell).
Personally these choices have always seemed like no-brainers to me: allow but don't require type information, and warn when the compiler finds a problem but allow the program to run anyway. What possible advantage is there to be gained from any other choice?
It's not pretty but it can be done. :P
That’s why my Python APIs never fatally crash, they just log an exception and move onto the next request.
One of the most annoying things moving from Ruby / Python into a language with static types that are checked at compile time is dealing with JSON from an external API.
I know, I know. The API could send me back a lone unicode snowman in the body of the response even if the Content-Type is set to JSON. I know that is theoretically possible.
But practically, the endpoint is going to send me back a moustache bracket blob. The "id" element is going to be an integer. The "first name" element is going to be a string.
Sometimes a disable-able warning is all we want.
With code gen you can usually just print out your JSON directly from the source and import that as a type into your source code.
This is why apis are moving towards specifications that trigger multiple generators.
JSON typing is possible without generation, but more cumbersome.
interface TypedJSON {
string?: string;
number?: number;
array?: TypedJSON[];
boolean?: boolean;
object?: {[key: string]: TypedJSON}
}{foo: {string: "my string"}, bar: {number: 4}}
Probably makes sense to make the "string" key "s" for minimal payload size impact, or this could also be done by converting standard JSON to this typed representation at runtime.
And within the bounds of pragmatism, it's nice that it actually checks this stuff and throws exceptions. In Ruby if the ID field is the wrong type, the exception won't happen at time of JSON.parse(...) but instead much later when you go to use the parsed result. Probably leading to more confusion, and definitely making it harder to find that it was specifically an issue with the JSON and not with your code.
Because e.g. in Rust what you're talking about is:
#[derive(Deserialize)]
struct Whatever {
id: usize,
first_name: String
}
// ...
let thing: Whatever = get(some_url)?.json()?;
You encode your assumptions as a structure, then you do your call, ask for the conversion, get notified that that can fail (and you can get an error if you but ask), and then you get your value knowing your assumptions have been confirmed.It's definitely less "throw at the wall" than just
thing = get(some_url).json()
but it's a lot easier to manipulate, inspect, and debug when it invariably goes to shit because some of your assumptions were wrong (turns out `first_name` can be both missing and null, for different reasons).For a 5 lines throwaway script the latter is fine, but if it's going to run more than a pair of time or need more than few dozen minutes (to write or run) I've always been happier with the former, the cycle of "edit, make a typo, breaks at runtime, fix that, an other typo, breaks again, run for an hour, blow up because of a typo in the reporting so you get jack shit so you get to run it again".
That doesn't make much sense - either the API has a well-defined schema, and you define your static types to match (which there are usually tools to do for you automatically, based on an OpenXML/Swagger schema, if there isn't already an SDK provided in your language of choice), or you use a type designed for holding more arbitrary data structures (e.g. a simple string->object Dictionary, or some sort of dedicated JsonObject or JsonArray type).
There are issues around JSON and static typing, e.g. should you treat JSON with extra fields as an error, date/time handling, how to represent missing values etc. (and whether missing is semantically different to null or even undefined), but I would hardly see it as "one of the most annoying things".
Edit: the other issue I forgot is JSON with field names that aren't valid identifiers, e.g. starting with numerals or containing spaces or periods etc. But again, not insurmountable. I'd actually be in favour of more restrictive variation of JSON that was better geared towards mapping onto static types in common languages.
1. The program is in fact correct, but the typechecker can't prove the program is correct and so reports it as a failure (for contrast other typecheckers only report errors if it can prove that there is an error)
2. The type checker is correct in that the program is flawed in some way, but you are the middle of development and it would be more productive if you could observe the behavior of the program before fixing the types. Since your CI pipeline will fail on type errors it's not like you're going to be able to merge your changes before fixing them.
3. The program is incorrect but correct for all expected end user inputs.
I somewhat agree with your point #2 which is why Haskell provides a deferred type checking mode for prototyping purposes[1]. I think this should be more common among typed languages, but this is not a very typical scenario either.
[1] https://downloads.haskell.org/ghc/latest/docs/users_guide/ex...
* An integer that is sum of two primes.
* A string that is created between the hours of 2 pm - 4 pm on a Tuesday.
* A dictionary who's fields comprise a valid tax return.
What do you mean by future proving? If the new version of the code doesn't ship by 2 pm today, the company goes bankrupt and there is no future for the code.
Any language with types can enforce any proposition you can reliably check, which you run as part of the type's constructor. There are multiple ways to do this, for instance in pseudo-ML:
> * An integer that is sum of two primes.
type SumOfPrimes
fn sumOfPrimes : Prime -> Prime -> SumOfPrimes
type Prime
fn primeOfInt : Int -> Maybe Prime
> * A string that is created between the hours of 2 pm - 4 pm on a Tuesday. type Tuesday2To4String
fn whichString : Clock -> String -> Either String Tuesday2To4String
> * A dictionary who's fields comprise a valid tax return. type ValidTaxReturn
fn checkTaxReturn : Dictionary String String -> Maybe ValidTaxReturn
> What do you mean by future proving? If the new version of the code doesn't ship by 2 pm today, the company goes bankrupt and there is no future for the code.Shipping broken code will not save the company from bankruptcy. It's also a myth that dynamically typed languages will get your code out faster.
Nevertheless let's move on.
It's not a myth, if you are using dynamic typing you will get three times the number of features out the door as someone who uses static typing.
It's just a fact.
1 dynamic typed programmer = 3 static typed programmers in productivity and output.
That is why people use dynamic typing.
From a commerical perspective, static typing does not make sense in most cases. You need something more than a typical business CRUD app to justify it. Usually you only need static typing for performance reasons. Like for a video game or image processing.
In practice, the shipped code doesn't have noticably more bugs. That just something which people who have never used dynamic typing properly like to believe.
If you ever read an article by Eve online where they say Python is their secret weapon or watched what Discord did by implementing everything in Python first. You will understand.
See the signature of the constructor function I provided.
> or that a tax return is valid given the thousands of rules involved.
Tax software does it every year. Clearly the government has an effective procedure to decide the validity of a tax return, even if that procedure is executed by a human.
> It's not a myth, if you are using dynamic typing you will get three times the number of features out the door as someone who uses static typing. It's just a fact.
That's just laughably false. There is zero empirical evidence supporting this claim.
> In practice, the shipped code doesn't have noticably more bugs. That just something which people who have never used dynamic typing properly like to believe.
No True Scotsman fallacy.
"That's just laughably false. There is zero empirical evidence supporting this claim."
There is a reason all the startups are using Python. Just ask Eve Online or Discord. Dynamic typing is great for getting products to market fast.
Neither of them have problems with too many bugs. What you are saying doesn't even make sense, since you would just increase your QA budget if it increased the rate of bugs. So how in earth could the end user see more bugs?
Dynamic typing has an issue but that issue is code performance. Eve Online gets slow downs during massive battles and Discord rewrote parts of their interface in Rust to save on server costs.
Show the evidence that "all the startups" are using Python.
https://games.greggman.com/game/dynamic-typing-static-typing...
Because all the studies say that dynamic typing reduces development time without increasing the number of bugs.
This is literally false, which you can clearly see if you'd bothered to actually check the references cited in that article that links to this review:
https://danluu.com/empirical-pl/
Numerous studies demonstrate that static typing had little effect on development time, if the typed language was expressive, but did have noticeable effects on the quality of results and the maintainability of the system. The more advanced and expressive the type, as with Haskell and Scala, the better the outcome, but all studies done so far have flaws.
As for the specific data discussed in that article, here's what that review had to say:
The speaker used data from Github to determine that approximately 2.7% of Python bugs are type errors. Python's TypeError, AttributeError, and NameError were classified as type errors. The speaker rounded 2.7% down to 2% and claimed that 2% of errors were type related. The speaker mentioned that on a commercial codebase he worked with, 1% of errors were type related, but that could be rounded down from anything less than 2%. The speaker mentioned looking at the equivalent errors in Ruby, Clojure, and other dynamic languages, but didn't present any data on those other languages.
This data might be good but it's impossible to tell because there isn't enough information about the methodology. Something this has going for is that the number is in the right ballpark, compared to the made up number we got when compared the bug rate from Code Complete to the number of bugs found by Farrer. Possibly interesting, but thin.
In other words, this "evidence" you cited is a clearly biased anecdote, nothing more. I also recommend reading the HN comments also linked, where former Python programmers describe why scaling dynamic typing to large codebases is problematic.In the end, my originally stated position on this stands: you have literally no empirical evidence to justify the claims you've made in this thread.
With that in mind, your other claims like the 1:3 ratio really need very solid sources to be taken seriously.
On Monday morning, I wrote the dynamic typed version of the program. Then until Thursday afternoon wrote the statically typed version of the program. Then the rest of the week was writing tests that compare the output of both versions of the program.
After 2 years of doing it, I think I would know. If that sounds weird to you, it's called Prototyping, shocking that someone might prototype something in a scripting language designed for prototyping, I know.
The two classic examples are Eve Online with their "Python is our secret weapon". They used Python to out develop their competition. The other example is Discord who used Python to get to market quickly. Then they used Rust to reduce their operating costs by increasing the code's performance.
Python adding type annotations is an unfortunate problem associated with too many people coming in from statically typed languages. The problem with success and all that.
Statically typed Python codebases tend to be awful because the developers don't realise that they need to break the codebase down into smaller programs.
JavaScript is not a very good language due to the way it was designed and built and pretty much anything is an improvement to it.
Sometimes I don't care which political party the barista voted for, I just want to get my coffee.
People and programs are never perfect. Type checked program is not guaranteed to be correct, it just has stronger constraints.
You're thinking like a single individual. Almost all programming that happens in the real world is in teams, and most of that probably within companies.
Teams and especially companies often benefit from strict standards that keep everyone on the same page and make it harder to screw things up.
Sure, but those standards don't have to be enforced by the language design. They can be enforced, for example, by policy, e.g. "Don't push anything that produces warnings to the master branch". That still allows you to, for example, test parts of an uncompleted program, which can vastly improve productivity. Not being able to test-run a partially completed program can increase your development cycle by orders of magnitude.
All that being said, your point is a valid one. It's a tradeoff and depending on what you're used to, you'll feel the pain from one approach or the other more strongly.
Yes, policy debates are amazing, so fun, so productive. And nothing funnier than having to discover completely new and incompatible policy when changing company, or trying to change policy.
That's one of the few good things I have to say about Go, they considered policy debates and went "fuck that". I think they went too far and wrong in lots of ways, but on that front I can't fault the heart, even if I hate the language.
In a TypeScript codebase you would not let CI pass if the TypeScript build fails. Locally you can still run the output. There's not much room for debate and I've never seen a TypeScript project that both has CI and allowed failing builds to be merged. I've also never heard of someone proposing that PRs failing to build should be allowed to be merged.
Who says there's going to be a debate? The commit hooks and/or CI pipeline should prevent you from merging code that doesn't pass typecheck. It's still enforced by tooling, just at a different place in the process.
The question is really, before you commit/push your code, and you just want to run the code, do you really need to be prevented from running the code just because it didn't pass typecheck? I'm inclined to say no.
If I wanted something that could kinda work, I’d just use a scripting language?
Which is usually either no-one, or an already busy/overloaded team lead or (even worse) manager.
Every software engineering organization I’ve been at either wanted, or put in place, automated policy enforcement because otherwise it either never happened (and hence there were constant problems), or everyone hated doing it because it felt like a waste of their time (with it’s own constant source of problems).
it’s best for everyone if a problem is identified and called out as close in time to when it’s created as possible, and as clearly as possible. Ideally before it can be part of any product, or be depended on by other work that would need to be refactored.
Which automated tools can do cheaply. If it is in the compiler, even better!
Doing it manually delays the feedback cycle, AND burns people decision energy, which is almost always the thing everyone is in the shortest supply of anyway.
Sure. But it does not follow that it is beneficial for your language to forbid you from testing some code you wrote over here until you fix a completely unrelated problem over there, indeed until you fix every single problem that the compiler can find.
> Doing it manually delays the feedback cycle
I'm not talking about "doing it manually", I'm talking about the compiler issuing warnings instead of errors.
If the policy is ‘you can generate “warnings” all your want locally, but you can’t check them in’ then sure.
The issue of course is folks want to check things in as soon as they ‘work’ in their basic use case, even if it means bypassing a bypassable alert.
So it means enforcing it at check-in time, in a non bypassable way (usually), which just means folks have a lot of ‘hidden’ cleanup to do before submitting, while they have their code ‘working’, which is always a drag.
Even in static languages world, a lot of projects allow dev/nightly builds to be unstable. Which is literally what you're talking about: checking in problems for everyone else.
But checking in policy is a VCS/CI domain, not a language domain.
CI/CD, based on CTO edict? Have them (or their startup equivalent) make the edict at the start of the company's lifecycle for best results.
Don't make it a human procedure, make it a "you're physically unable to deploy until tool X doesn't throw a warning". Anything else will, as you say, fail.
I laughed out loud when I read this.
Funny thing is, this would work where I'm at now, Google, because Google actually has an enormous number of presubmit checks that require all KINDS of things out of your code, it's practically crazy how many checks it runs. Since there's already so much setup for these checks, yeah adding one more would be fine (and it already does this for type annotations for dynamically typed languages like Javascript).
But most places aren't Google, and it's way more realistic to prevent this kind of problem by just having a language that'll enforce it automatically, then you never have to think about it, or have a policy debate, or always make sure during code review to look for that thing, etc.
If I'm putting out some small open source library where I might get an occasional pull request, I'd much rather have the language automatically enforce correct typing than have to remember to look for it myself every time (and then possibly miss something if I screw up).
I have never worked on a project that has actually followed such advice.
This sounds insane to me. Do you really think needing to specify types up front saves you that much time?
Multiple orders of magnitude means you're talking a 100x speedup at minimum.
For example, making unused function arguments a fatal error would be an easy place to increase iteration time.
Then the following arguments are moot, as CICD will run it in the “sane default” setting.
The GP asked what’s the point of languages like this. The two massive advantages in my mind are that compiled programs are usually correct, and compiled programs run really fast. I often have the experience in rust of writing hundreds of lines of new code, and my program works perfectly the first time I run it. That’s a delight.
It has to do with documentation, intent, and clarity. Also clarity of error messages, compilation errors with generic functions under GTI are as bad as C++ template errors.
I wish it were common [0] for source code to show a combination of programmer-specified and compiler-deduced details.
Seems like we could get the best of both worlds.
[0] I know some IDEs show e.g. type information as mouse-over text. I'm talking about something that integrates more cleanly into the actual source-code text. So, e.g., it would show up when diff'ing the source code.
This is absolutely false. Compiled programs might be correct with respect to type safety, but that is very different from being correct. There are all kinds of ways that programs can fail besides type errors.
In fact, this widespread false belief is IMHO one of the strongest arguments against using enforced-compile-time-strong-typing because it lulls you into a false sense of security. I have a long list of war stories where this led to problems that ranged from annoying to very nearly catastrophic.
Yes, you still need tests. But testing is never perfect. Bugs still slip through. I like not needing to be as paranoid in rust as I am in other languages.
And as GHC actually has a compilation mode where it only warns about compilation errors: https://gitlab.haskell.org/ghc/ghc/-/wikis/defer-errors-to-r...
So they're 0/2, on the subject they're espousing upon, through a language they mentioned.
Your compiler giving you a type check error is saying that the program is meaningless. What is the advantage of running a program that is known to be wrong?
The design of the type system will define what is meaningful/meaningless per type check rule (eg a missing branch which can be a warning vs an ambiguous type inference that means the program cannot be compiled). The question is really whether you want to design a type system that can be checked statically, which to me the no-brainer answer is "yes."
I don't want to have to run the program to know if it type checks. I don't want to have to run it to know what types are in the program, or the signature of functions. I want my tool to tell me when my code is provably wrong, and to prevent provably wrong code from being merged into a codebase without executing it.
Well one bit is you might not be looking at the meaningless part.
Sometimes when doing large refactorings or updates never getting to run, try, or look at anything is discouraging. It's a great feeling when it compiles and everything works (or near enough), but spending two days cycling between compiling and fixing compilation errors is not as enjoyable as e.g. seeing the unit test report's green bar tick up.
Maybe compilers should provide more enjoyable feedback, not just better error messages (though I fully support that movement).
I can't think of many worse experiences than my past self trying to refactor Python 2 code. Compared to _any_ type system it's just not even close.
What's the difference between seeing a bunch of test failures and fixing them one by one until the test suite is all green, and seeing a bunch of compiler errors and fixing them one by one until the compiler runs clean?
Speaking of which, that WOULD be a cool IDE feature. I always liked the feeling of my junit bar going all green after a big change.
For example, you might not be pushing buttons that trigger the meaningless part, but the other parts of the system might still be functional.
> What's the difference between seeing a bunch of test failures and fixing them one by one until the test suite is all green, and seeing a bunch of compiler errors and fixing them one by one until the compiler runs clean?
The difference is in one of the cases a highly qualified professional has freedom to decide _when_ to do the fixes.
Like updating data structures should never result in days of work before you can run simple tests. Create a new build artifact that only pulls in the new data structures and the software modules you're working on (and their tests). If you need to pull in the universe to run any unit tests, they aren't unit tests.
The TypeScript compiler will absolutely give you errors for well-formed programs that are guaranteed* do the right thing with respect to what the author intended and that anyone else can see is correct but that the TypeScript compiler itself doesn't reason about correctly. In instances where this happens, it's very useful to be able to ignore them—just like you might ignore the protests of a misinformed team member who insists that such-and-such won't work, even though it provably will. (Although the alternative of just not using the TypeScript compiler can be a worthwhile decision as well.)
* not to be confused with undefined behavior in C where it might just happen to work with one known compiler/platform combo but break horribly for other valid targets
For example, TypeScript explicitly declares as a design goal that they don't use the type system to generate different code:
https://github.com/Microsoft/TypeScript/wiki/TypeScript-Desi...
Meanwhile in Dart, you can write:
double d = 1;
In Dart, doubles and integers are different types with different representations. The above line of code works because if we see an integer literal in a context where a floating-point number is expected, we implicitly treat it as a double literal. Note that there's no runtime conversion here. We can do this because Dart does have a sound type system and is happy to use it for language features.That's also why Dart (and other statically typed languages) support extension methods while TypeScript does not:
https://github.com/microsoft/TypeScript/issues/9#issuecommen...
If the language gives users the ability to type their code, it's really nice to be able to give them as much value in return for their effort. That means building language features on top of them and using types to generate smaller, faster code.
But languages with unsound type systems or that treat programs with type errors as well formed can't really do that.
I've been meaning to write up something comparing dynamic typing + SLIME to -fdefer-type-errors and ghcid...
If you're the only one writing the entire stack of code that gets bundled into the delivered program, you're right, no real advantage from other choices.
But as that's rarely the case, there is a lot of advantage from having a forced type system. Otherwise your code might be great but you're pulling in libraries and code from many other teams who chose to omit type info and/or ignore all warnings so you still have the same problem.
This is actually a really great argument. Having well-typed dependencies is a real benefit. But I would say you don't need to force typing on people to achieve that. In my experience with gradually typed languages (TypeScript and Erlang) you tend to see decent adoption of typing in prominent libraries. There are natural incentives (i.e. crashing software) that drive people away from depending upon libraries that don't have strong type guarantees.
Substantially better performance, which still matters in some domains.
If you occupy any space on the spectrum besides fully static typing, you must represent and check your types at runtime. This means that every object (even Int and Bool!) needs to have a header that tells the runtime what type it is, which can quite literally double your RAM requirements.
If a language is already including object headers anyway (like Java does), you're correct that there's not an enormous barrier to embracing gradual typing [0]. But there are a lot of places where we really do still care about RAM usage, and erasing types at runtime is an obvious place to free up a lot of RAM very quickly.
[0] Indeed, you can sort of do this in Java already by just typing everything as Object and coming back later to correct it with actual types, though please don't actually do this.
If the language requires you to give at so much type information that the compiler can infer all the types, I would say that language, from the compiler’s view, would have fully static typing, so it wouldn’t have to do any checks at runtime, but from the programmer’s viewpoint, it could look as if it didn’t have fully static typing.
The modern auto to deduce a type in C++ makes a small step towards that, but I think it could be taken further. The language could deduce types both in the direction “if you pass foo to a function taking an integer, it must be an integer” and “if you only pass integers to bar, bar must take integers”.
I expect it will be hard to specify a powerful algorithm for filling in the blanks that feels natural to programmers, though. If the compiler frequently complains about missing type information where the programmer doesn’t understand why it complains, the language won’t be popular.
However, better type inference does not mean that a language is no longer statically typed, it just moves a lot of the burden of assigning types from the programmer to the compiler. You still don't have as much flexibility at runtime, and you still can't run your code until the compiler can prove that it it correctly typed. The only thing you get is less syntax, not dynamic semantics.
This is the problem in a nutshell. Folks in the optional typing and gradual typing communities have been trying to pull this off for decades and so far no one has been able to come up with a language that:
1. Lets you ignore types completely and write dynamically typed code in some parts of the program.
2. Gives you the expected performance of completely typed code in parts of the program that are typed.
3. Is sufficiently pleasant to use that it gets adoption.
There are languages that give you either 1 or 2 along with 3. There may be some language out there that gives you 1 and 2 but if so, it doesn't give you 3 or I would have heard of it.
I suspect it's intractable.
I am not aware of any static languages that check the types at runtime (if such languages exist, please mention them in the comments); however. they require a static analysis to happen at compile time, and refuse to compile if the analysis breaks at any point. Python, on the other hand, allows to be compiled and executed without ensuring the type safety first; however it can still be statically analyzed, even without explicit type annotations, using an optional tool such as mypy (which can infer types quite correctly in many cases).
For example, Java will throw an exception at runtime if an object is cast (using .asInstanceOf) to an incompatible class. Python throws if you add an integer to a string. Both of those languages are strongly typed.
On the other hand, C has no runtime checks whatsoever. If you add a number to an array the program will happily move the array pointer and maybe overflow, but even that won't cause a runtime failure every time.
I expect this must be fairly common among IDEs with incremental compilers. You can run valid sections & gradually progress the complete correctness. No need for a dynamic language.
With static typed language, reliability and refactoring are easy.
I'm scripting some small scripts in Python this week but every time I come back to it, I am reminded how not actually very good it is.
I was with you until your last sentence. The world of programming is a lot bigger than your tiny corner of it.
I could be working in a safety-critical situation, where preventing errors is more important than developer velocity.
Or, I could just have that one annoying co-worker who checks in stuff that "works for him", but always crashes when others try to use it.
Or...
Per #2, why not a language that facilitates "lint" style inspectors to give a list of warnings, as in "do you really wanna do that?" This reduces the burden of the compiler. Meta-markers can switch certain warnings off so we don't clutter the results for intentional spot looseness.
Liberties are rarely created and destroyed. Instead, they tend to be trasferred from one place to another. So an equally valid way to phrase this question is from the compiler's perspective:
1. Do you want the compiler to be able to rely on the presence of type information (e.g. Java), or sometimes have access to type information (e.g. Common Lisp) or be forbidden from using type information (e.g. early versions of Python, Javascript).
In that equivalent formulation, the value proposition of static types is clearer: When the language specifies that everything has a computable static type, then the compiler can rely on types for useful language features (extension methods, overloading, static dispatch, etc.) and to generate smaller, faster code.
Dynamic types give the programmer some liberties (they can not worry about types) but takes some others away (they can't reach certain performance goals or use type-dependent features). Static types do the opposite. Pick your poison.
A pseudocode example is something like:
trait ToString for T {
fn toString(self: T) -> String;
}
impl ToString for number {
fn toString(self: number) -> String {
// implementation here
}
}
impl ToString for Point {
fn toString(self: Point) -> String {
return f"({self.x.toString()}, {self.y.toString()})";
}
}
fn printString<T: ToString>(t: T) {
print(t.toString());
}
Then when it generates code, it uses the inferred type information to generate different code: let p = Point(1.0, 2.0);
printString(p);
// translated into
printString(implToStringForPoint, p);
let n = 2.0;
printString(n);
// translated into
printString(implToStringForNumber, n);
And then type inference becomes implementation inference, which is really powerfulAlso, I was explicitly asking besides performance.
I would argue that something like "dynamic" is preferable in a sense that it gives you an escape hatch with minimal syntactic overhead for those cases where dynamicity is really that beneficial, yet defaults to a safer option that's usually the best choice.
I guess the part where it will also pick the correct overload at runtime depending on the actual type of the arguments is a bit like multiple dispatch, especially if the argument itself is "dynamic". I'd say that the difference is that the list of possible candidates is closed once the class is defined, so even if you declare a new type, you can't customize the behavior when passing it to an object of an existing class.
It's in some sense the "most dynamic" language, because you start one instance and repeatedly mutate it. And people mainly use introspection as the main tool for finding definitions/help/etc. instead of static documentation, so you don't have a half-hearted REPL.
I'm more of a static types person myself, but CL would be what I compare them to.
I'll quote the author:
> a dynamically typed language is one where types are runtime values and manipulable like all other values. It’s a short hop from there to thinking of the whole runtime environment in the same way, where everything is a runtime construct.
If you have a decent compile-time metaprogramming system paired with a statically typed/compiled language, (almost) the same statement holds true. It's not as easy to, for example, instrument a production instance with brand new code at runtime, but the mental model while programming is very similar.
> The overall picture is of going one step further than homoiconicity: instead of code as data we have programs as data.
This is the heart of what metaprogramming is all about and, to my mind, really has nothing to do with static vs dynamic types.
The author finishes with 'crushing disappointment'. I somewhat share this sentiment, though I have great hope as well; there are some incredibly talented developers working on the compile-time version of these problems. Maybe one day we'll have good languages. I guess the only thing we can do is continue writing new ones until we do ;)
A very good example to look at is the `comptime` in Zig, it could be enlightening and the opposite of 'crushing disappointment'.
Not really? I mean it is more difficult in that you have to specify all sort of things which are left unsaid in the article's snippet, but fundamentally you just need the ability to define or override the act of "calling" an object, aka
void operator () ()
or impl FnOnce<...> for <...>
Something not all statically typed languages have, but not all dynamically typed languages either (doesn't exist in javascript for instance, you can staple shit on a function object or extend Function but AFAIK the language has no callability protocol).> I can imagine all sorts of uses for this kind of trick. Here’s some I’d like to see:
Most (though not all) of these are already better and more generally solved via introspection tooling.
> I guess the closest to the hyperprogramming model is Pharo? The dynamic code inspector is a big deal, plus there’s this fun talk.
The author should take a look at Self. It has similarities with Smalltalk systems, but is an other window into an... interesting world.
So static typing vs. dynamic typing isn’t about what you can or cannot do, it’s about having a mechanism to verify the soundness of your reasoning vs. having no such mechanism.
You need unit tests to confirm the correctness of your program. And unit tests check most of the types for free. So you don't need static typing after all.
Great for readability when I'm wondering what type this parameter is supposed to be and what I can do with it.
You can obviously have documentation of types without static types, but static typing means you have ever-present compiler-verified documentation of types. Not always enough, but it's a lot more helpful than just having nothing.
Having exhaustive unit tests doesn't necessarily help with this.
Lisp macros allow us to specify multiple versions of code at compile time. And Lisp conditions allow us to redefine or discard runtime contexts. But this time-based manipulation is simpler than what you seemingly want. The easiest way I can think of is to offer an API to track the evolution of types, but I'm guessing the performance hit would be heavy.
I really, really don't mind if people have preference for static typing. I do too - not as strong as the 2022 zeitgeist to be sure but I still have it.
What I dislike is the inference that people who prefer dynamic typing are ignorant, that they just haven't discovered hindley-milner and need to be converted.
A lot of incredibly smart, talented programmers chose (and choose) to do work with dynamically typed languages. Don't assume it was because they knew less than you.
After ~30 years of professional software development in pretty much every major language created, I prefer dynamic languages to statically types ones for a whole host of reasons, none of which amount to me being a n00b or uneducated, which is frequently implied by the static types die-hard folks.
When you develop in this style, you tend not to write type errors because you are constantly validating your program inputs and outputs, which the tight feedback loop encourages. Most statically typed languages have slow compilers. It is not uncommon for an incremental compile to take at least 1s, which is 50-100x slower than what I'm used to. Indeed, I find it distracting when compilation is that slow so with such a language I will still use a monitor script for iteration but I will manually save when I want validation. This means that I am not only getting slower feedback, I am getting less of it, which ultimately makes me less confident in my program. If I were making lots of type errors, maybe it would be worth it to have a type checker, but I'm generally not.
I think it's mostly about time to market and reduced total LOC.
I find that in strongly typed languages, frequently up to 30% or more of the total lines of code are dedicated to nothing but type satisfaction.
Sometimes this formalism is warranted, depending on the problem domain (banking, air craft systems, etc) but most of the time I find that the type systems are merely mental abstractions that are created to assist the developer in modeling the solution.
There's value in that.
As I've gained more experience, trivialities like the compiler catching junior programmer level silly type mistakes has become less valuable. Type systems don't address logic errors.
The flexibility of duck typing with loose contracts reduces total time to market for applications where that's a primary constraint.
This happens to be true in a vast class of problem domains, where having something quickly is more important than performance or formal correctness.
Since I tend to live in the startup space where TTM and MVPs are critical to business success, I'm inclined towards languages that support this.
I typically don't have to deal with large team coordination issues or extremely complex interdependence, and my systems are normally fairly basic - crud with business rules and a fancy UI.
In this case, things like type errors present themselves immediately in the UI and are easy to catch during development, reducing the value of compile time error detection.
They are different development experiences, one isn't better than the other. If you want to ship something to market tomorrow use Python with dynamic typing. If you want something to run fast and be error free use Rust with static typing.
Both are good options.
TL;DR:
I work with a lot of very smart, very talented software developers who thrive with Python's dynamic qualities. They have little problem reasoning about our software's runtime behavior.
Unfortunately, I'm a bit less intelligent than them. I have a much harder time understanding the code's design and behavior due to that dynamism. I would really benefit from the crutch of strong static typing, and the more obvious relationship between static source code and runtime behavior.
If that same software were written in e.g. idiomatic Python without the use of typing hints, I often need to read a lot more comments, and look at the particular contexts in which that function or class is used, in order to understand its intended purpose and supported use-cases. Things like monkey-patching make it even harder for me.
So for someone accustomed to understanding a piece of software by reading the source code (i.e., human, static analysis), a big software system written in highly dynamic Python has a really long learning curve.
The only place monkey patching is typically used is in places it has no effect on code meaning such as GEvent.
Are you implying type hints triple the size of code? This is very far away from my experience, no idea what kind of monster types you're writing.
The only places you need hints are really function parameters and return types - your IDE can infer the rest.
It's not the type hints that are responsible. It's the supporting code for allowing the usage of static type checking. Abstract base classes, interface, generics, that stuff which just doesn't exist in the dynamically typed world.
I’ve also seen monkey patching done in crazy ways even on small projects. Languages like Ruby pretty much encourage it.
I can’t imagine what it would look like to edit a Python or Ruby codebase that was a decade old and had 1000 people working on it. But I’ve done that in Java fairly easily because the types help. A little bit of “public static final” here and there is a small price to pay for more maintainable code in the long run.
If you ever find yourself asking what's actually in this variable, and the code is too dynamic for the IDE to tell, it's going to be too hard for a lot of developers to tell, and maybe even the author themselves in a few years.
If the type hints are too complex and involve multi line annotations with 500 unioned types, then we should remember: the programmer does still have to remember what type that parameter can be, even if it's not annotated. Leaving out the annotation doesn't simplify the program, it just moves some information from from explicit to implicit.
Calling static typing a crutch is like calling a passenger airliner's takeoff checklist a crutch. I mean maybe it is from one perspective...
I’ve worked on million+ line Perl codebases, and even bigger FAANG codebases in each of those languages, and untyped large code bases turn into impenetrable balls of mud much quicker than typed ones.
And myself and others were doing runtime meta programming in Perl long before it was cool or anyone we knew was calling it that.
But never underestimate the ability for a software engineer to create code in any language that will perplex and mystify them a year later.
And I’ve not met a codebase in any language yet that couldn’t be turned into a ball of mud pretty quickly.
The problem is that infinite malleability is cool, but a lot of tasks do not benefit from it. Static typing, on the other hand, keeps you from making a lot of dumb mistakes. Static typing lets you spend time on your logic, whereas with dynamic typed languages like JS or Python, you spend time on logic AND on typos, misremembering function arguments and or positions, etc. The thing is, most routines can only really accept one or two types, and only very rarely does one want to treat data as code.
A word processor, or a sorting algorithm, or banking library have no need to modify their code; in fact, doing so is likely to break something. The sorting algorithm would benefit from polymorphism, but the others would not. You cannot render an image the same way as text. And the banking library had better be adding integers, not strings or even floats. However, all programs will benefit from the compiler saying "the function doesn't accept this type" or "this variable doesn't actually exist".
This guy should find a job maintaining a complex Python server, he'll learn why people like static types. Nothing like digging through a bunch of code to find, "ah this looks like the function I need, what does it take? Oh 'connection'. What's that?" Static typing at least forces a minimum of documentation. Or you make a change, server runs but still not fixed. Then you notice it's not even doing what it used to do, but no errors or anything. After tearing your hair out, you discover you swapped two letters in a variable, which caused a syntax error exception loading the module, which was silently discarded by the dynamic loading code.
This guy wrote several books on formal methods. Not only does he know why people like static types, he works to encourage people to adopt more advanced forms of static analysis.
I find such statements incredibly daunting. Like the person is acting like a child for no other reason then to be contrarian. I'd hate to work with such a person.
We contain multitudes.
But the moment you try to use the full arsenal of dynamic reflection, metaprogramming, etc. then the runtime checks and gadgetry needed to analyze all of that explodes in complexity and even if it's possible to make an analysis framework capable of analyzing it, the results might be very difficult to interpret.
So as long as you try to stick to first-order constructs as much as possible, you can get quite far.
I don't really understand it yet and have barely used it but this seems like it's supported by expression trees / Roslyn / code gen stuff in .NET/C#?
Is there a difference I'm missing here, e.g. writing some lambda and having the code generate code to run it against NoSQL or SQL or some other storage at runtime which is what I think things like LINQ2SQL do under the hood is basically equivalent right?
This type of programming seems like it can be supported by both types of types?
And that difficulty scales up fast. Consider the Python function decorator. It was quickly discovered that the decorators need to preserve the metadata for the function's name, etc., or the decorated functions broke in places that looked at that metadata. But if you have a "function name" parameter that the decorators can modify, now there's a mismatch between the fact that the function claims to be "the original GetUser function" but it's actually the decorated function. So now you ought to have "the visible function name" and "the real function name", so for instance, the former works with metaprogramming correctly but the latter allows for correct debugging output on crashes.
But what if you decorate a decorator and then want to do metaprogramming on it? Now you're got the apparent function name, the real name of the function beneath it, and the real name of the function beneath that.
And this is just one dimension of problem. There are many others when you get to metaprogramming like this.
The core problem is really this style of metaprogramming, in my opinion. It almost inevitably slips into becoming Deception-Based Programming, in which increasing swathes of the program are trying to figure out what lies to tell to other chunks of the program to get the desired effect. Dynamic typing makes this harder because it makes it very difficult to get a manifest that lists all the behaviors this particular chunk of code or code object has, but even if you were doing this in a static language that didn't make that hard, a language would still have a problem exposing all the correct properties and assisting programmers in getting them completely correct.
It's fun at first, and very powerful at first. But deception-based programming is one of the quintessential ways a code base fails to scale and becomes an incomprehensible tangled mass, code bases that use heuristics to figure out what lies to tell to other bits of the programming to cover up the effects of the lies the other bits of the code are telling. This leads to your programming becoming an embodiment of the phrase "What a tangled web we weave, when first we practice to deceive".
That said, before someone pops up with "but we could do that without deception"... I agree! Or at least, I agree that it should be explored better. I would be interested in a language that somehow enables such metaprogramming, but with a philosophy that says that it will stay "honest" throughout, and basically deliberately eschew and exclude this sort of "transparency". I don't think it's necessarily intrinsic to the style. I just think that it's sooooo tempting that if you don't build a language that excludes it from the get go that the temptation to just shim this little behavior in just here is too tempting, plus the languages tend to end up requiring it, be it accidentally or deliberately, because they afford that behavior and the language ends up being built around how it is at least supposed to be possible. I don't know entirely what this would look like.
Edit: I say I don't know what this would look like, but I have had a couple of ideas. One I've talked about before is to have a clean separation between an "initialization" phase in which metaprogramming is legal, and a "sealing" phase in which it is all done and the metaprogramming becomes invalid/illegal. I wrote this for the runtime not to have to be super flexible all the time, but it works for programmers using the system to know that things won't change out from under them.
The other is some convenient and powerful way to dump things out in their final state post-manipulation. I've gotten some good mileage out of this in my own code that is using a lot of composition. It is very helpful to be able to take the top level of the composed object and say "Tell me what exactly you are made out of". Even if I'm not doing metaprogramming based on the response, it's a huge debugging help. I see traces and fragments of this idea here and there, but no system that has it coherently integrated in a way I'd like to see.
It's the programming version of being penny wise and pound foolish. This style of metaprogramming reduces a few lines of boilerplate code at the expensive of making a system as a whole extremely difficult to understand or maintain.
Having restrictions like types and no meta programming is less “fun” but the new person on the team will get up to speed more quickly.
The one job where in C# actually they used a tonne of meta programming I could barely get anything done. Because of course the documentation is non existent and the codebase so big would take you years to comprehend, or trying to get some help from never-free team members.
tldr KISS!
Some implementation details need to be visible to debuggers and for analysis, but hidden in "regular" code so that you can substitute a different implementation without invalidating constraints.
For example, functions normally have names for debugging purposes, but the names can't be relied on in "regular" code or you wouldn't be able to rename them without breaking things. It may be better if the names aren't visible at runtime at all.
The reason Python function decorators have these issues is due to a lack of structure; a lot of metadata is just there for anybody to use, with no constraints.
It's also possible to do this by convention - for example, many things in Go's reflection package will break abstraction boundaries if you use them the wrong way. It's less of an issue if you have strong language conventions against doing that. In many scripting languages, the conventions for when you can use metadata are weaker.
What I found was basically no IDEs supported runtime generation of code for things like autocomplete. Jupyter does but it’s a limiting environment.
It made me realize this is a mostly unexplored space with little tooling
https://learn.microsoft.com/en-us/dotnet/csharp/roslyn-sdk/s...
For example, there is an SQL provider that gives you generated types and auto-complete but it does this without running a source generator that emits files. It just happens inside the compiler.
But I don’t have to time to write all that. So I have actually been writing helper scripts to generate object code and function signatures via ast libs.
This is how I originally “learned programming” in the 90s. The industry zeitgeist shifted to “churn out code” and all that wizardry was lost. I have no idea how to prove it but I sometimes think it was intentional to capture worker agency.
It’s been fun reconnecting to it but the ecosystem of helper tooling … well there is none. Tons of blogs with basics about for loops though.
In app code bases, there are always multiple duplicates of the same type, no one truly respects reusability partially because it’s not always practical.
With dynamic typing, you can be completely transparent and have multiple “Car” classes. It doesn’t matter because types don’t matter. To us they both represent a car, no need for converters, helpers etc. Car1 definition has some method you need for Car2? No problem, just tack it on in that scope! It’s not polluting the definition of Car2, it's not going anywhere.
The author seems to want to use dynamic typing in a restricted manner, to help in testing and debugging and understanding a program, particularly within a REPL. Further down, he seems to be saying that the lack of empirical evidence on how this works in practice (on account of it rarely, if ever, being used in available tooling) seems to suggest it is difficult.
use Moose;
sub foo {
...
}
around foo => sub {
say "Called with args: @_";
}We don't spend most of our time writing generic libraries.
We don't write generic code, nor do we want to. We write code that takes specific inputs, and generates specific outputs.
So I want to use statically typed languages, maybe with a template system at MOST.
If it makes my life more difficult when I write a library, I accept that.
Source: LISP programmer for 6 years. Learned the above the hard way.
My favorite "dynamic" type that is stupid useful has to be numbers in Common Lisp. (I think most all lisps, actually?) Having it widen to be as big of a number as you need it to be, while you are using it, is stupid useful.
Always idly thought that could work in a statically typed lang and was a feature well worth copying (how many languages REPLs are useless as calculators due to floating point errors!)
Haskell has infinite-precision numbers.
You probably could just call the data statically typed, but with the idea that the concrete type that you are actually dealing with could be any of a few types. Think a function in Java that takes in a Number, but obviously has to have code to work with any actual numeric types. Even more fun, the code could do basic bounds checking to make sure that the result you are working with fits, such that the code starts with an Integer it is working with, but auto moves to Long, etc.
And that is where I wish dynamic typing focused more. It isn't that the type of things is not knowable. Clearly they can be known. And it is nice to have parameters constrained to be within a range of expected types. The actual type, though, is dependent on runtime data.
I was not disappointed. This article is complex and interesting and deserves better.
As a Java programmer, you can guess where I lean. Now excuse me while this autowired annotation Dependency Injection causes a new self inflicted level of hell.
I like static typing, but his this is just a silly argument. If the program behaves correctly according to spec, how can the types be wrong?
The only ways I know of to get rid of (some of) these bugs are to: - use as much static typing as possible; - wherever static typing is not sufficient, sprinkle this with a paranoid dose of assertions.
If I read what you write correctly, that's way beyond "testing for the correct behavior", no?
Sounds more like a problem with tight coupling than a problem with types per se. A function should not change behavior due to an "unrelated change", regardless of static or dynamic typing.
``` function foo(config: {hasYakShaver: boolean}) { if (config.hasYakShaver) { // ... } else { // ... } }
foo({hasChuckYeager: false}); ```
This will accidentally do what I want, but for bad reasons. And one day, quite possibly, a change that appears unrelated will break this.
And that's me being nice. If I start using reflection to do complicated stuff, or if I'm using C++ and the reason for which my code works is memory corruption... well, when I need to figure out what suddenly broke the code, I'm going to be up for serious a moment of loneliness.
If you have unit tests (or other types of automated testing) you have a much bigger chance of catching logical errors and it will also catch trivial error like misspelling an identifier.
> Testing for the correct behavior should be enough.
Also, in my experience, automated tests tend to catch many bugs, static typing also catches many bugs and defensive programming catches many bugs, too. There is an intersection between all three, but it's by no mean 100%. So I use all three of these.
And of course, let's not discount monitoring, fuzzing, etc.
Yeah but with dynamic typing there's a whole bunch of behaviors you need to think about. You pass an int into a method that expects a string, what's the expected behavior? Or maybe something more nefarious, like a method that expects a Python dictionary but is passed in a pandas Dataframe, which has similar syntax with totally different meaning; what's the expected behavior?
Much easier to just specify what types you're expecting to work with and let a compiler or type checker deal with it.
Okay, but that doesn't answer the question. We're in dynamic typing world; my method accepts any data type by design. What's the expected behavior when it doesn't receive the expected type? Do I add assertions and fail with an assertion error? Let the runtime fail for me? Silently convert the type? In any case, I need to add a test for that.
If your answer is "humans shouldn't make a mistake when calling the method," that's arguing for static typing, not against, because that's exactly what static typing prevents: a human making a mistake. Turns out static typing makes for a really good baseline level of documentation, as well.
> And some mechanism that uses types like method overload should be transparent from a behavior perspective.
Your method might work today for both a dict and a dataframe because it's only using simple accessors. But what if someone wants to go back and change it and they use something that's only on dictionaries, like the * operator? Your code would fine, your unit test cases probably all use dictionaries so they're fine, and it's only when you deploy to production and find out that someone's code elsewhere was incorrectly passing in a Dataframe that just happened to work, no longer does.
Static typing has a solution for behavior-based types too, they're called interfaces. Even better in languages like Go and Python (with Protocols) that use structural type checking for interfaces, rather than explicit interface implementations.
When I’m talking about testing behavior, it’s about testing that the code does what it’s supposed to. Not what the arguments is. So you test that ‘trim’ works on all strings, including empty. What static typing helps in checking assumptions when programming, but every bet is off when executing. Especially when interacting with the outside world.
If someone changes the code, it’s on them to notify everyone about breaking changes. If the documentation says that the method accept dict, and you call it with dataframes, it’s on you to check that it will continue to work with dataframes, not on me to ensure that someone is calling it with dataframes if I use specific code related to dict.
This is equivalent to re-implementing static types and a compiler with human processes. This isn't scalable.
Probably an exception, depending on the language. E.g. if the method expects the argument to support the method "substring()" it would be expected to throw an exception if the passed object does not have such a method.
If you pass an object which support the expected interface but with a completely different semantics, then you probably get a weird unexpected behavior. But that would also be the case in a statically typed language.
Static typing is something you use for performance reasons primarily.
If you want to catch errors with static typing you need Haskell or Rust levels of support for it. C# / Java / C++ don't make the grade.
That is not how to decrease the complexity of reasoning about a Python program.
What you should do instead is break the program up into smaller microservices.
Errors that would be caught by static type checking are exceptionally rare when you do commerical development in dynamically typed languages.
You right now are complaining about a problem that only exists with inside your own head.
> Errors that would be caught by static type checking are exceptionally rare when you do commerical development in dynamically typed languages.
Those kinds of errors are not rare at all.
Dynamically typed languages tend to be simpler and easier to use. This means on the whole they have less bugs than their static typed counter parts.
The real issue with dynamically typed languages is that their performance sucks. (That and according to this thread people not having a clue on how they are properly used. You can't just use the same development techniques as you use in a statically typed language. It's different.)
And yet your comments demonstrate the opposite.
> Dynamically typed languages tend to be simpler and easier to use.
Very debatable.
> This means on the whole they have less bugs than their static typed counter parts.
Not at all.
> The real issue with dynamically typed languages is that their performance sucks
Every time you say something like this it just makes it obvious you have no idea what you're talking about. You can write almost completely typeless, highly performant, c and assembly code, while high level languages with advanced type systems are generally not the most performant.
You seen to be confused about interpreted/jit/compiled languages and static/dynamic typing - which are not the same thing.
The key points are as follows:
* Development in Dynamically Typed Languages is faster
* Dynamically typed languages use less lines of code than statically typed languages by a significant margin
* The level of bugs is the same between Dynamically typed and Statically typed code.
"You can write almost completely typeless, highly performant, c and assembly code,"
Yes, I'm sure the LOAD instructions take pictures of cats as operands via their typeless instruction sets... What in earth are you on???
"completely typeless, highly performant, c"
well once someone does that let me know. The only thing near is Javascript and that is only faster for microbenchmarks not real programs.
From one of the comments:
https://dev.to/aussieguy/the-non-broken-promise-of-static-ty...
> The article covers a study of the same name. In it, researchers looked at 400 fixed bugs in JavaScript projects hosted on GitHub. For each bug, the researchers tried to see if adding type annotations (using TypeScript and Flow) would detect the bug. The results? A substantial 15% of bugs could be detected using type annotations. With this reduction in bugs, it's hard to deny the value of static typing.
> Yes, I'm sure the LOAD instructions take pictures of cats as operands via their typeless instruction sets... What in earth are you on??
Have you written any assembly? Assembly generally has no or very rudimentary type checking, you're generally dealing with words/bytes and addresses, you can arithmetically add parts of strings, divide pointers, etc. Errors due to these operation will surface at runtime, not be typechecked.
> well once someone does that let me know.
You can use void pointers as return types and arguments for all functions in c code. The effect is significantly less type checking while having equivalent performance.
Actually, that makes it really easy to deny the value of static typing. If the total number of bugs in dynamic and static code is the same. But some bugs in dynamic code would be caught by static typing checking.
We most conclude that adding static typing results in a large number of non-typing related bugs being added to the code base. It's simple maths.
"Have you written any assembly?" Yes, I'm an emulator author, thank you very much.
The registers and opcodes are typed with things such as u8, u16, u32, u64, i32, i64 and only work with data of the right type.
"arithmetically add parts of strings, divide pointers" You mean standard C stuff, you know the statically typed language.
"You can use void pointers as return types and arguments for all functions in c code" Dynamic types is not the same thing as type eraser. That void pointer doesn't carry the information that it points to a picture of a cat for example.
The fact you don't understand the difference between a void pointer and dynamic typing doesn't exactly surprize me. It's more like a giant vtable.
Completely baseless assumption.
> Yes, I'm an emulator author, thank you very much.
Ahahaha let's have a link then.
> The registers and opcodes are typed with things such as u8, u16, u32, u64, i32, i64 and only work with data of the right type.
Those are sizes, not types, and the same opcode generally applies to signed and unsigned integers. You consider that to be a static type system and you think the purpose is performance and not correctness? Lol
> You mean standard C stuff, you know the statically typed language.
You're trying to debate against the need for static typing by pointing out unsafe parts of c? Ahaha
> That void pointer doesn't carry the information that it points to a picture of a cat for example.
First of all you can totally carry around runtime type info and value with a single void pointer. Not to mention many dynamically typed languages have type erasure and many statically typed languages have runtime type info. Also, you've claimed multiple times that there's no need for type checking whatsoever. You need runtime type information at runtime now?
Not really. Lets say you have an `add(a, b)` function. Static typing can guarantee that the function returns a number, but not that it returns the correct number. So you need unit tests anyway.
A unit test `assert_equal(4, add(2, 2))` actually tells you something about the correctness about the function, and it implicitly also verifies that it returns a number. So static typing does not save you anything.
Static typing does have advantages, for example for documentation and for automated refactoring. But it doesn't replace any unit tests.
> But if the tests verify that the output is correct, then it implicitly follows that the types are correct, so you get type checking for free.
Your unit test checks the correct behaviour for one out of infinite possible inputs. You don't 'get type checking for free' , what you get is no type checking at all.
If I understand you correctly, you a suggesting unit tests in a dynamic language should test all possible invalid inputs. Since statically typed languages eliminate a subset of invalid inputs (because they won't compile), statically typed languages saves you from writing some of these tests.
My argument is that you wouldn't write such tests in the first place since they provide little to no value.
If function `foo()` calls function `bar()` you want to ensure that foo passes the correct arguments. You don't do that by checking that bar() rejects all possible invalid types. You do it by checking that foo() actually works and provide the correct result. The types involved are an implementation detail.
But you're testing a tiny subset of possible scenarios of behaviour of the function. If you can anticipate all possible input types and value ranges, maybe unit tests are enough, but that's not realistic in a dynamic language/program where complex values are non-deterministically generated at runtime - based on user input, db, file, etc.
If you are talking about untrusted external input then yes, you need to carefully validate and reject invalid input. If you receive JSON or a CSV file or whatever, there are an infinite number of valid inputs and an infinite number of invalid inputs. But this is the same for statically typed languages, and the type-checker will not help you, since a valid JSON string and an invalid JSON string look exactly same to the type checker.
That's not at all the same problem, because here you're intentionally writing code that throws exceptions for some reason? Also, many languages, such as Java will not allow you to throw a checked exception that's not declared on the interface method - such issues will be statically checked and caught - thanks for proving my point ;)
> of valid inputs and an infinite number of invalid inputs. But this is the same for statically typed languages, and the type-checker will not help you, since a valid JSON string and an invalid JSON string look exactly same to the type checker.
Not at all, in a statically typed language the JSON would typically be deserialized into a structured type with all the benefits of type checking, validation and usage of the resulting values. Of course it's possible to just use a JSON string directly, but that's not idiomatic and generally not the way it's done in quality codebases.
If some other function uses `add()` but pass an invalid argument, then this bug would be discovered by testing the function which pass the invalid argument. Presumably that function would work incorrectly.
Ok, now that entire scenario goes away with a language that has a compilation phase and has type enforcement. It's boiler plate inside the function for type validation your don't have to do inside the python function (or blame the caller) and a testing scenario(s) you don't have to write. That's the work I'm talking about.
Therefore the scenario is ridiculously unlikely and lives in the realm of fantasy.
Unit tests which verifies a function has the correct behavior gives you "type checking" for free.
You just have to shift the mindset to testing behavior instead of testing types.
Lets say `add(a, b)` contains a bug. Perhaps it doesn't handle negative values or overflow correctly or something else. But it still returns a number, just the wrong number. This kind of bug is much more likely. If this wrong number is stored somewhere you also might get obscure errors down the line and a static type system will not discover it. And this bug will be much more insidious and harder to track down.
So you need unit tests anyway to protect against logical errors. If a unit tests verifies the behavior of a function it will also implicitly verify the types without any extra work.
Broadly speaking, this is the contract-based approach: at every data flow boundary, explicitly describe your contract, so that any violation can be detected immediately instead of propagating and poisoning derived state. Eager type checking is just a subset of that, true. That means that static typing is not enough, not that it's unnecessary.
OTOH unit tests cannot fully express contracts, because even with parametrization you can only test so many of the possible inputs.
Of course `add()`is a silly example since this is already a built-in in any language, but consider a function `calculateSalesTax(product)` - how do you define the contract which guarantees the result is the correct number?
I don't know if I can make the claim that if a program behaves to spec, the types must be right (that seems like an academic statement out of my own wheelhouse). But I've always scratched my head at the canard of "static types saves time from obviating the need to write a whole class of type validation tests", that just doesn't map to my own experience.
I challenged my intuitions on this recently and checked out some of my existing unit tests suites to see what my type error coverage looked like. It was exceedingly rare that I was able to fine a behavior assertion which after introducing an intentional change to cause a type error, didn't blow up the test.
The cases where I did find type errors that snuck through with false negatives, the test itself was actually written poorly and the same behavior assertion would fail in other ways unrelated to types.
I'm sure there are situations (in my own stack and in other folks') where my above findings are contradicted but, given that static typing for something like JS is not in any way a zero-cost abstraction, I'm not sure it's really a net win.
For example, if a function accepts a string as an input, test it by passing in all other possible primitives and classes and ensure that the behavior is strongly defined and results in the expected exception class.
Then realize that if you haven't written all of those tests by hand in a dynamic codebase, you don't actually have code coverage. Anyone who works on your codebase could pass you an argument of the wrong type, and it might be in a code path that is rarely invoked and you will find out it's broken in prod.
My point is that you wouldn't write such unit tests in the first place. Unit tests should test behavior, not types.
Otherwise the behavior is undefined, and that is really bad.
So no it's not much worse than usual.
Also, if there are input values of the right type that cannot be handled properly, that should also be tested! That is the kind of thing tests are for!
Example: you write a function that divides two numbers. It's not possible to divide by zero. You should absolutely have a test that passes in a divisor of zero and asserts that some kind of custom DivideByZeroException is thrown. If you know that certain inputs are out of range, you should test them explicitly -- otherwise you're being sloppy. If there is a large number of inputs out of range, you should test all of them, and you can use a `for` loop if you have to.
The difference is, if you have a valid input out of range, that is all you need to test -- the language itself takes care of all the "wrong type" tests for you. Otherwise you have to write them yourselves, and if you don't, you don't actually have good coverage. I understand that it is customary in dynamic languages to not write those tests, but that is merely another argument for the superiority of static typing.
Passing the wrong type is considerably rarer than the wrong value so need not be considered.
Well one programmer using dynamic typing can deliver the output of three programmers who use static typing.
That why people use dynamic typing, from a business perspective it's a far more compelling argument than the one you are making.
A codebase with 1000 people working on it benefits dramatically from static typing. Everyone who has tried to do the same thing with dynamic typing has learned that. And quite a few companies that started off with Python had to rewrite everything in Java.
If you honestly think that nobody ever passes the wrong type, you've never worked on a complex dynamic codebase with a lot of people. It happens all the time.
I've worked on dynamically typed codebases for a long time. Passing the wrong type is very rare. I mean how would the code pass the unit tests if it was passing the wrong type?
Passing the wrong value doesn't just happen randomly. It happens because of errors in logic many of which can be caught with a strong static type system.
(There is a separate argument over what proportion of functional vs unit testing is optimal.)
The increase in the code size usually results in significantly more bugs in the code.
Number of bugs is directly proportional to the number of lines.
The better way is to use unit testing instead.
Anyway, that "bloat" is documentation your computer can understand so it can tell you when it's wrong. The alternative is documentation your computer can't check. Or just skipping that and letting the next person eat the higher onboarding cost and defect rate, which I suspect is more often the case.
[EDIT] And it's not unit testing instead. It's in addition to. One's not a replacement for the other.
Bugs are directly proportional to total lines of code if you weren't aware
To me, these lie on the efficient frontier of the "compiler validation" vs "less boilerplate" tradeoff. Java is decidedly nowhere near that frontier.
Edit: A note on val/var w/ Kotlin, Intellij/Android Studio makes it easier to determine what the type of the variable is without having to dig for anything, looking upstream.
Your code should be obvious and self-documenting. Relying on language features to do a bunch of stuff behind the scenes more often leads to programs that are hard to debug and perform badly. How many times have you had to dig into C++ operator overloading compiler errors or someone being too clever with C macros (wait, that macro can return?!?) or someone overloading a Python built-in function in surprising ways. Coming across the short double example from the article in production code would make me groan. The "all sorts of uses for this kind of trick" the author lists sound like a nightmare for some poor employee to have to reverse-engineer five years from now.
I feel modern languages (Rust, Go, Swift) have moved away from clever, implicit behavior in favor of more considered and explicit code-efficiency mechanisms. For example making error handling (? error operator in Swift & Rust) an explicit first-class citizen, rather than some clever macro hacking; or the trend away from operator overloading and complicated class structures and towards explicitly declared interfaces; or the many restrictions on Rust macros[1] to help ensure they're maintainable.
[1] https://veykril.github.io/tlborm/decl-macros/minutiae/hygien... and related sections
And if you use a “statically typed” language, but have an algebraic type that holds int or string or ClassA or Array, how useful is that static typing? That’s why I prefer calling the code either weekly typed or strongly typed. And some languages allow you to have strongly typed code with ease, and some languages allow you to have weakly typed code with ease, and some, like TypeScript, allow you both in which case the whole dilemma seems unwarranted or at least outdated.
Why do you even assume classes are involved?
> And, bam, now your “statically typed” language actually has dynamic type checking.
Not in any way, shape, or form?
You have a ParentClass, it lets you perform ParentClass operations. If you want to do a ChildClassA operation if that's that, you check whether that's what you have, then do the operation. Both are enforced by the compiler (though how safe the check is varies).
If you have a dynamically typed language, you might think you have a ParentClass but you actually have a completely unrelated ToyotaSupra.
> And if you use a “statically typed” language, but have an algebraic type that holds int or string or ClassA or Array, how useful is that static typing?
Extremely. Because you know what the potential is, and you have to check for it exhaustively.
> That’s why I prefer calling the code either weekly typed or strongly typed.
So complete nonsense to make you feel good?
> Not in any way, shape, or form?
https://en.cppreference.com/w/cpp/language/type#Dynamic_type
https://en.cppreference.com/w/cpp/language/dynamic_cast
https://en.cppreference.com/w/cpp/language/typeid
It’s there in the name.
In the same way that the Democratic Republic of North Korea is democratic. Right there in the name.
Also that Dynamic Programming involves not knowing anything. Right there in the name. A Dynamic Microphone? Not actually a microphone at all.
Because classes are almost ubiquitous in modern commercial programming?
> If you want to do a ChildClassA operation if that's that, you check whether that's what you have, then do the operation. Both are enforced by the compiler (though how safe the check is varies).
Go to GitHub and search for "GetType language:csharp" or "getClass language:java" and see millions of hits for runtime checking in “statically typed languages". And who knows how many ad hoc runtime checks there are in C++ and other languages without a default reflection framework.
> If you have a dynamically typed language, you might think you have a ParentClass but you actually have a completely unrelated ToyotaSupra.
So is a dynamic language that supports type annotations in comments (eg JavaScript + JSDoc) actually a statically typed language?
> Extremely. Because you know what the potential is, and you have to check for it exhaustively.
It’s still going to be a mess that should be refactored.
> So complete nonsense to make you feel good?
You can do runtime type checks in “statically typed” languages. You can use type inference in “dynamic” languages. That is just facts.
Most "statically typed" OOP languages have this escape hatch for dynamic typing because it is very useful, and is usually only performed after a check for the correct type has been performed. or if it is known by the programmer that the correct type is present, but there is no way to express this knowledge in the type system.