Why do dynamic languages make it more difficult to maintain large codebases?
programmers.stackexchange.com
programmers.stackexchange.com
Please note (if you can't be bothered to read the actual post before responding), I am _not_ saying that static typing has no value. I'm saying it has both a value _and_ a cost, and too many arguments seem to consider only one of these.
These are not the same thing. See Haskell for a counterexample.
EDIT: I see that Haskell is mentioned in the comments. Still, I don't see any understanding there that static vs. dynamic is not the issue.
Declare types of all named functions (i.e. not lambdas). This serves two purposes.
First, it makes your code more self-documenting. When I look at a function, probably the first thing I want to know is: "What are its parameters?" Type declarations on functions give you a trustworthy and highly readable answer to that question.
Second, it makes type errors easier to understand. If you don't explicitly declare types, the compiler will perform type inference. When your code has a type error, the compiler can infer a nonsensical type for a function, yielding a hard-to-understand error message. If your function has a type declaration, the error message is more likely to point you to where your actual mistake is.
Rarely declare types on anything that's not a top-level function definition. For example, if I want to call (foo (bar baz)), I don't need to add separate type declarations to baz, (bar baz), and (foo (bar baz)), even though each quite possibly has a different type. On the other hand, if you want to remind yourself of the type of some inner expression, or you think it'll make the type errors more readable, you can add the declaration. It won't hurt you.
While type inference is cool, it's only really useful when I'm playing inside GHCi. Also - it can often improve the quality of a program a lot by thinking about your types first, declaring them, then filling out the function definitions (that is often how I do it).
note: Yes I know the compiler can't actually infer all types all the time, but the exceptions are obvious and require type information in dynamically typed languages too (ie. convert a string to an int: you have to tell it you want an int, you can't just say "convert this string" and have it guess what you want it converted to).
From TFA:
"They write test cases for every identifier ever used in the program. In a world where misspellings are silently ignored, this is necessary. This is a cost."
So in a dynamic language the cost is still there. In fact, it's probably higher than in a static language because you have to build at least some of the infrastructure yourself. Well, I imagine some people would argue that in a dynamic language the cost is optional. And again, I point to type inference where sophisticated static languages achieve the same thing.
How does "dynamic typing"="misspellings are silently ignored"? Maybe in some languages but that's not a corollary of dynamic typing.
Disclaimer: Haven't read TFA so shoot me down if the context makes this clear
And yes, you'll need to do a little bit of infrastructure on the test cases, but you'll gain the extra power that comes with dynamic types, that makes everything (even writting tests) much easier. And, of course, your codebase can be much less "huge" with that extra power.
Anyway, all that thing is a big red herring. Nobody decides to create a huge codebase.
What "extra power" are we talking about here? If we're talking about the power to pretend a bool is a string or vice versa, then no thanks, I'll pass. I suspect that a number of the things you are referring to that give you "extra power" are actually things that don't have anything to do with the type system like first class functions, tuples, etc. Static languages can have those things too, you know.
That's weak typing. JavaScript does this, Python and Ruby don't.
Typically when people bring up the power of dynamic languages it's in the ability to alter definitions. The ability to extend existing classes/objects is pretty powerful, and isn't always trivially duplicated in static languages.
Many concepts that can be painful in statically typed languages (creating/programming against new interfaces to existing types, reflection, delegation, etc) are easy as breathing with dynamic types. It's not a matter of can vs. can't so much as easy vs hard.
s/pretend a bool is a string or vice versa/not know whether an object even has a certain method because it isn't added until runtime/
I'll grant you that this type of thing is more difficult in static languages. But in my experience, the need for this kind of dynamic behavior is quite rare in terms of lines of code. So I'll trade it away any day of the week for the ability to reason about my code with more guarantees. At their core, all bugs are situations where the programmer's expectations didn't match reality. In my experience programmers who are good at debugging are good at being able to question as much of what could be going wrong as possible. Strong static types are about reducing the amount of things we have to question.
Agda is offended.
You can hack code together and forgo documenting it, but it's a form of technical debt. When creating a minimum viable product for a startup, taking on some forms of technical debt probably makes sense, but by the time you're talking about huge code bases, you're probably past the startup phase (I hope) and you're already up to your neck in technical debt.
It should be easy to integrate into any editor.
Anyway, my point is not that it's difficult or bad: it's that it's has _some_ cost.
Not the case in Haskell, at least. Most Haskell code out uses type signatures.
> Conversely, a requirement that for explicit type declarations makes the use of generics more cumbersome.
True enough, but you can't have your cake and eat it. Either you want explicit documentation-by-types, which has a cost, or you don't (which also has a cost in terms of maintenance).
But then the compiler has the capacity to write/emit the type declarations.
For me, I've looked inside one of the biggest Python codebases in existence (the EVE Online client). It's a freaking mess. A simulation of static typing but mostly they just passing freaking tuples around. Everything is a tuple in a tuple in a dictionary in a tuple in a homegrown table.
There are a variety of languages that don't impose explicit typing on you to enable static typing. OCaml and Haskell are popular options. Scala makes the compromise choice of having explicit function signatures with inferred typing elsewhere.
If the only thing you're using the type system for is to make sure arguments line up, you're barely using it at all. When your types can encode (and thus, ensure) aspects of your program logic, that is when static typing becomes useful.
For C++, note the now accepted use of `auto` and `decltype` which together remedy a number of issues from C++03 and prior.
You mean "in which I pose the rather obvious strawman". What scaffolding? What is this mysterious "extra stuff" people give vague names like "scaffolding" to that I've never encountered despite programming in a statically typed language all the time?
And while it's valuable to talk about popular languages in terms of the kind of static typing people are likely to encounter... they also have really outdated static type systems. While you can argue that it's likely that you'll bump into Java/C/C++ when working with static types, it's pretty invalid to argue generally about static type systems using them as examples. Things have just come a really long way.
No, because the affordances and limitations of static typing are not the same as the affordances and limitations of static typing as implemented in Java/C++. You can't infer very much about static typing from statements made about those languages, and you can't infer much about those languages from statements made about static typing, so it follows that you can't freely substitute one for the other in these discussions.
(not necessarily directed at you personally)
If the function has to handle different types in different ways, you'd have to write it again anyway. Otherwise, you could have used templates.
I'm not totally convinced, though. Take the simple case of "tell me what range of values to expect here". Java / C++ deal with this with their types, and you can be reasonably confident that you're getting the object you expect. As I'm sure somebody will mention, those types are a bit weak, and can't guarantee the object you get is exactly what you expect. What if you don't test with the right sub-class? What if it's an int, but it's out of the range you expect?
"No, that's not what types guarantee!" you may say. And it's true.
I think one of the problems is that language features aren't perfect. Forcing everybody to admit they're dealing with integers is more desirable than no contract at all, but would it be better if you dealt with Scores, which where integers in the range of 0-100?
In a big enough codebase, the language is never powerful enough to handle all your checking for you, so you need some level of discipline. I worry that by having something "good enough" when your code base is moderate, it allows teams to slip by into "large" without ever considering what their tools, conventions, and limitations should be.
Type systems can do all this and more.
https://github.com/milessabin/shapeless/wiki/Feature-overvie...
They can assure you handle all possible control flow cases with ADTs and pattern matching, they can facilitate automated testing, etc.
Just because it's not how java/c# do it, doesn't mean it's a limitation of 'typing'.
That's dishonest. That's the kind of stuff Ada or Haskell do. The linked repository is a Scala library.
1. I have a function 'saveScore' that has a parameter 'score' that is an int.
2. Users of my function frequently misuse my function, passing in negative scores, so my team decides a range-checked Score parameter makes more sense.
3a. In a static language, users of my new code get compilation errors saying 'saveScore' takes a 'Score', not an int.
3b. In a dynamic language, at best, devs that use my code get an easy-to-read exception thrown at runtime. This assumes I'm thorough enough to check my preconditions and throw an understandable error message. This also assumes all code paths are covered before the code is released to customers.
In short, static languages provide facilities for library and framework authors to enforce a little discipline in the consumers of their code.
Airplane? Hah, think bigger: http://en.wikipedia.org/wiki/Mars_Climate_Orbiter#Cause_of_f...
That's what value objects are all about, and yes - you can do that in a mainstream language such as Java/C#. http://en.wikipedia.org/wiki/Value_object
I strongly recommend that pattern.
Some dynamic languages even allow for type specifications and static analysis, optionally or through an external tool.
If you want to check preconditions rigorously and pervasively, use a statically typed language (and be a little realistic in what you can check). If you want short code without ceremony, use a dynamic language and don't uglify it with more boilerplate than the worst statically typed language. I've seen that kind of code in ruby, and it's really not a joy.
Real life in a dynamic language: people introduce some checks in some high-risk functions, and call it a day. And that's... Just Fine.
So in practice both (a) and (b) are weakened. Usually it's a whole lot of (a) and a little bit of (b). You can reduce the problem by establishing domains, but that limits reusability and requires its own kind of stringent documentation and discipline to handle.
That is something that you won't find that often in the enterprise space.
Which is actually where most large codebases tend to exist.
What utter rubbish.
What you won't find that often in the enterprise space is dynamic languages, so your statement is irrelevant as well.
Your "decent programmer" is about as common as a Sufficiently Smart Compiler.
Consider this common mistake:
interface Greet {
void sayHello(String firstName, String lastName)
}
class PrintGreeter implements Greet {
public void sayHello(String lastName, String firstName) {
...
}
}
Where's your compiler NOW?Notice http://hackage.haskell.org/package/happstack-server-7.3.2/do... isn't a string like "GET" or "POST".
-ddump-* is great for this kind of thing.
That said, it works well enough for deriving monad from a transformer stack. Do people use monad transformer stacks? It's been a while since I've written any real Haskell code.
There are more adventurous and specific ways being explored to handle things like "effects", but that exists mostly in SHE or Idris.
mtl is still the practical way to go for now.
I keep wanting an excuse to futz around with this: http://hackage.haskell.org/package/layers
(I swapped key and value in a string, string API and ended up writing 5+ MB "keys".)
If you language makes that problematic, it is the fault of the language implementation, not a problem with static typing.
(I've seen plenty of Haskell programs that are not compile-time checked for, say, XSS problems, even though it's trivial to do that with the type system. The problem is that nothing is ever automatic; if you want to write the best possible code, you have to make it a goal and carefully execute that goal. There is no silver bullet.)
double :: Int -> Int
double = (3*)
The only amount of typing that is going to help with that is an amount nobody is willing to use."The best feature of the script when it comes to higher-level product is that large portions (if not the entire thing) can be scrapped and started anew in a different way without a significant loss of investment."
Large portions can be scrapped and started anew without significant loss of investment?!? Is this code being auto-generated somehow? If it's legit actual lines being written by your developers, then that is the definition of a significant investment. And if you're redoing it, then that's by definition loss of the original investment.
"There are probably only a handful of apps that wind up being 'large' (in the 10k+ LoC range)."
What are you smoking? I've got more than 10k LoC in an app that has only been under development by < 2 developers for a few months. My company has another app built by one guy that's more than 30k LoC. And you say you have 30+ million LoC in production?!? If 10k is "large" for you, then you must have a shitload of apps out there.
Also, there's a huge difference between maintaining 3000 apps with 10k LoC and maintaining even one app with >1m LoC.
In our setup, apps leverage key lower level libraries that are maintained by a small team and represent a much smaller, more manageable native codebase. A very complex app can be written in 10k LoC. Yes, there are a lot of "apps".
> Also, there's a huge difference between maintaining 3000 apps with 10k LoC and maintaining even one app with >1m LoC.
I said exactly this at the end of my post, but I don't feel it is always black and white. I said "apps" in quotes above because many times they act more like modules than standalone apps. Many of them interact in non-trivial ways, making the comparison more muddled. It's definitely a lot of code either way and in our situation JS has definitely made evolving the product much easier from an end-user perspective.
edit: ^once
The most damning thing is perhaps that C and C++ require additional testing that JavaScript does not require, and that is testing for memory safety. In C++ you can kind of stay safe by using references but there is still a lot of risk and you always have to keep object lifetimes in mind. C and C++ are probably among the least productive languages in common use today for other reasons as well—such as for their header files.
My personal opinion/observation/"gut feeling" is that dynamic typing has helped, but I haven't done any kind of concrete study to figure out its impact vs, say, removing the compile/link phase. I observe some of the code being written and definitely see programmers taking advantage of the dynamic nature of the language.
Whenever I recall programming in Java though, I remember highly verbose, boilerplate-ridden codebases, and bits of code that are little more than abstractions for bits of code underneath them (with unit tests alongside them that test bit of code A calls functions B and C, which of course has no value whatsoever).
Nowadays I do Javascript, much better. Today I frowned upon a colleague who is used to Objective-C who proposed writing documentation. Pfff.
So developers with limited experience encounter some big "enterprise" Java code base and recoil in horror. Instead of making the deduction that enterprise Java is terrible, they jump all the way to compile time types being horrible.
While Scala/Haskell/F#/etc are great at concision, they aren't that common in traditional SW industrial applications.
IME, I've tried to make super concise modern C#, and it simply isn't concise compared to Haskell or even Python. YMMV, of course.
2. C# supports dynamic typing as well as static.
3. C# and F# are the easiest to mix and match within one solution.
Are you talking about automatic types? Because that's the only thing that resembles dynamic typing I ever found in C#. And it misses about all the reflection available in most dynamic types languages, thus, even if it's somehow possible to not define types, it's useless.
[1] http://msdn.microsoft.com/en-us/library/dd264736.aspx
[2] http://msdn.microsoft.com/en-us/library/system.type.aspx
ExampleClass ec = new ExampleClass();
//ec.exampleMethod1(10, 4); // compiler error
dynamic dynamic_ec = new ExampleClass();
dynamic_ec.exampleMethod1(10, 4); // runtime errorAs for scala, it's not quite as functional or well typed inferred as the others... interested in what the LOC savings would be.
Realistically though what's being compared is Ruby/Python vs. Java. It's much more than just the language it's the culture, java is just plain bureaucratic, a lot of routine paperwork for not a lot of functionality.
I would say that the most stable codebase would be a static one with lots of tests, but I would also guess that this would be the slowest to develop, and would be the most difficult to implement in a real world team of developers.
I know most of us would balk at a huge codebase with zero tests, but it happens all the time. I would definitely prefer a dynamic code base in which I could get my team to write decent tests, but give up type safety, than to have type safety at the cost of decent tests.
In the context of software engineering approaches, "What research exists?" is basically a conversation killer of a question. The scientific method, when applied to medium-number, complex-interaction systems like most business settings, offers very little in the way of predictive power. At best, it may illuminate some dynamics and raise some questions to ask, but rarely predict the future in the way that hard sciences research does. This is why, if you a predisposed to label something like TDD as "cargo cult", you will likely never find any published research to be "convincing".
But that said, a simple search for "published tdd research" turns up quite a few hits, including this one from Microsoft right at the top: http://research.microsoft.com/en-us/groups/ese/nagappan_tdd....
However the benefits of testing can seem rather obvious if you are in the software industry. We hire people to do QA, and if the developers write tests, we can hire fewer people to do QA and still meet reliability targets. Software development is a loop: design -> code -> test, over and over again, and automated testing means that the code is still fresh in your mind when you fix it.
But it sounds like you're wondering, "why test at all?" Well, would you sell a car or a table saw that had never been tested to a customer?
It's absolutely not worth it. By the way, I have a codebase with 240k line of early 2000's guaranteed-untested PHP, coded by the best and brightest and complete with best-of-breed programming practices such as global variables I could maybe interest you in. I also have a full complement of bridges available.
> I never read about any actual convincing research that shows TDD as actually being beneficial.
TDD? That's probably good when you know what you are doing, but I must admit I'm a "test after the fact" man most of the time. Probably because I don't know what I'm doing most of the time :)
Why guess, when you can look at real examples?
Look at standards like MISRA C, or projects like the computers on the STS or Mars rovers. Using a static type system and testing things is but a small part of the process typically used for high-reliability systems.
And that's the key. Producing high quality systems is something that is emergent out of a process designed to produce high quality systems (and in fact, you need to have a process for improving the process). What is your process? Has it gone under rigorous examination with an eye towards output quality? Then any quality that results is a happy accident of competent engineers and a particular time and place.
Unit tests, types, are simply a small part of how quality is created.
I only see one link to that answer on SO that points to a single study. It was provided by another commentator asking for the OP to provide evidence for the "strong correlation," claim. Not very good.
Though I'd hate to work with a team that used a statically typed language and tools that didn't write tests for their software. It's not magic soya-sauce that frees you from ever introducing bugs into your software. Most static analyzers I've seen for C-like languages involve computing the fixed-point from a graph (ie: looking for convergence). Generics makes things a little trickier. Tests are as much about specification as they are about correctness.
In my experience there are some things you will only ever know at run-time and the trade-off in flexibility for static analysis is not very beneficial in most cases.
Some interesting areas in program analysis are, imho, the intersection of logic programming, decomposition methods and constraint programming as applied to whole-program analysis. Projects like kibit in Clojure-land are neat and it would be cool to see them applied more generally to other problems such as, "correctness," and the like.
The beginning of the project is wonderfully productive. By the time you regret having chosen a dynamic language for the project, it's too late to switch.
"Rather it is also everything else [correctness facilities besides static typing] that is frequently missing from a dynamic language that increases costs in a large codebase. Dynamic languages which also include facilities for good testing, for modularization, reuse, encapsulation, and so on, can indeed decrease costs when programming in the large, but many frequently-used dynamic languages do not have these facilities built in. Someone has to build them, and that adds cost."
...I think that's a little generous to statically typed languages. C and C++, both heavy-hitters in the static language world, have anemic modularization, reuse, encapsulation, mocking frameworks, etc.
A language with strong typing facilitates communication about the intent of code as well as its function. Strong typing is easier to build into a static compilation phase, but this is not an absolute requirement.
Granted, three of the four quadrants are well covered. Dynamic, strongly-typed languages will take a speed hit unless designed by a wizard.
Sounds great on paper, I know, and is much harder in practice. But if done correctly then dynamic languages are awesome boosts to productivity.
So I don't know if I have an answer to this question -but certainly it appears, to me, that the more dynamic the language is - the easier it is to maintain .. certainly the mere fact of lighter tooling means this is true?
Does that mean you have a large set of small-to-medium-size code bases? Because, though that's also a hard problem, it's different than maintaining one huge code base.
Needing to do this indicates worse problems with the developer than the fact of a codebase being too unwieldy.
Compare that to languages like C++ where despite its age and being used to widely there are still very few working refactoring tools.
(All that being said, I usually used method argument reordering when removing or adding arguments to a method by changing what the method is doing and then noticing that a different order makes more sense when reading.)
Is this really something people are doing?
So for our highly used method we realize that there is actually a reasonable default for the first parameter (maybe because we notice tons of usages of that value). In order to change this from a required to an optional field, we must move it to the back of the list.
Without an automated and safe way to do this, I'm unlikely to make that refactoring, leaving us with a code base that is worse than it could be. Again, is it necessary? No, but if it is trivial my code will get better, if it isn't my code will stay worse.
I'm starting to build a pretty large piece of software and am struggling with this question right now. I'd much rather work in Lua or Python, but I've seen and worked on so many large Python projects that are absolutely unmaintanable and poor-performing that I hesitate. However, those projects all share a gratuitous use of threads, which I think is a poor idea in any language, but particularly in Python.
But .. I really do think I have a future as a 100%-Lua user, across the full stack. From mobile/desktop client to backend, and persistence. I'm just not finding much energy to deal with all the cruft of the other languages, while Lua just gives me all I need, and leaves me alone. I don't know if I'm becoming myopic in my old age - possibly - but the more I think about where I can place that sticky little VM, the more I think I just don't care about much else but doing that, a lot...
Object.class_eval do
def to_s
"a"
end
end
# in some other context perhaps a controller action... def lookup_user
user = User.find_by_name(params[:name].to_s)
if user.blank?
flash[:error] = "No user found by that name"
redirect_to "/users"
return
end
# good stuff goes here
end
#
# try debugging it... it's tricky... use grep?
#the thing is this is actually not too bad... in C++ try debugging a memory leak...
(formatting...)
On a codebase writen by a team scattered around the globe, with the customer shouting on technical support and on site support team to know when the fix is ready.
I like C++, but I don't miss manual memory management.
Statically typed languages should just provides code that need to be reused often. Also needed when you need performance or low power consumption.
Scripting should be about using those libraries together.
People should try to understand how gmail works, because I think javascript is just used for presenting data to a browser, while gmail servers do the heavilifting.
"Let me begin by saying that it is hard to maintain a large codebase, period.... That said, there are reasons why the effort expended in maintaining a large codebase in a dynamic language is somewhat larger than the effort expended for statically typed languages. I'll explore a few of those in this post."
Static types are not enough to ensure that your algorithm receives correct input, to ensure you that you will need to write tests anyway. So with static languages you get one check "for free" out of the many that you will have to write yourself.
If you rigorously write tests for a JavaScript code as you would have for a Java code, it becomes just as reliable.
Remember, great software can be made with any programming language! This is why all your favorite apps are written in Visual Basic. Hope this helps!
> So if dynamic languages are making it harder to maintain your code, that means you chose wrong, tool-wise and/or job-wise.
The issue with large code bases is that rewriting everything is too expensive, so by the time you have data and experience showing your code base is difficult to maintain, it's already too late.
I ask this because it seems like a lot of "big code" issues stem from poor development practices polluting the codebase over a long period of time, destroying any sort of design that once existed.
Does it help if the language takes a hardline stance against mutable state and side effects? I'm certain you can screw those up, but it's confined to a single place.
Maybe the idea that you can avoid side effects is a fiction that the horror of manual memory management cures you of.
And you really, really should try harder to make small isolated components (services, daemons, sites, commands, whatever) and avoid large codebases.
If you're using Java or C++ or mainstream OOP, you're not getting the benefits out of static typing anyway and a Python or Clojure is a big improvement. That said, there are domains where compile-time static typing is extremely useful (and justifies the associated costs due to, e.g., the loss of a Lisp's intuitive macro system and first-class REPL) but those are going to call for a Hindley-Milner type system, not Java's weird mess of one.
What is it that about Java or C++ or OOP that you think makes their static typing useless?
> but those are going to call for a Hindley-Milner type system, not Java's weird mess of one.
Hindley-Milner is actually a poor type system. It trades away powerful features like overloading and subtyping in exchange for global type inference. Scala, like C++ and C#, does not use Hindley-Milner but a kind of local type inference.
On the other hand, subtyping does add a large amount of complexity and is misused about 90% of the time.
He's not saying their static typing is useless. He's saying it's not an accurate representation of the state-of-the-art in static type systems.
> Hindley-Milner is actually a poor type system. It trades away powerful features like overloading and subtyping in exchange for global type inference.
Something that you may think is a good idea in one context may not be such a good idea in another context. I'll take Haskell's type classes over subtyping any day.
Perhaps he meant to say "full benefits" and it was only a typo.
> I'll take Haskell's type classes over subtyping any day.
Type classes and subtyping are not incompatible.