I think the whole "lack of static typing" thing is a red herring, for two reasons:
1. Large-scale code is more about contracts that specifically "static typing". There are many contracts which you can't express using a type system.
2. Clojure doesn't "lack" anything. You can implement contracts for functions using pre/post conditions, use clojure.spec to implement arbitrary contracts over your data (and I do mean arbitrary), or even go for core.typed if that is your thing. I started to rely on spec more and more and I believe that coupled with pre/post conditions, it is way more useful for larger systems than static typing.
One other factor which is often forgotten in these discussions, because so few languages get there: please remember when "choosing" a language, that Clojure can be paired with ClojureScript on the frontend side. That means you can use the same data structures, same transport format (no JSON or XML quirks, ship EDN over websockets using transit and sente) and much of the same data model code on both sides. This is a huge advantage if you can make use of it. And guess what, your spec definitions work just as well on the client side, too.
I think the static/dynamic divide speaks to deep divisions in programmer personalities, but if you're open to suggestion, I would put it like this: As Haskell/OCaML are to static langs like Java, Clojure is to dynamic langs like Python/Javascript.
Clojure has a deep vein of pragmatic simplicity throughout its core libraries and community that I think does more for taming large-scale development than static typing does, but that just my $.02.
I'm curious if we're going to see a big new synergy between Kubernetes (represent everything as data) and Clojure (great with data).
Let's take a simple example. I have an object with the method named "get". But I call "fetch" in my code. When will I see this error? During compile time or run time?
All spec errors turn up at run time. See the guide for more details: https://clojure.org/guides/spec.
You can combine spec with test.check (Clojure's quick check library) https://clojure.org/guides/spec#_generators
Suppose you have an XML format, <library> that has books and authors[1]. The elements look something like this: <book author-id="0000"/>, <author id="0000" name="Plath"/>. Obviously, you want to make sure that book's author-id attribute will always refer to an author that actually exists. This is something a (Java-style) static type system can't do: it doesn't know at compile-time what the contents of a variable will be. But spec can do this because it is a runtime check[2]: you'd just write a function that ensures all author-ids refer to extant authors and register it with spec, telling it that this must be true for valid <library>s.
As for your example, I don't think that is a use-case that spec was intended to handle. You could write a spec to ensure that an object has certain properties/methods, but I'm not spec would be too useful for a function invocation on that object.
[1]: For a worked out example of this in spec, see a blog post I wrote: https://lgessler.com/posts/2018-07-12-choosing-the-right-too...
[2]: It sounds expensive, but there's a compiler flag that lets you turn off all spec checks, so you can have them only run in dev builds if you like.
So is it similar to or different from assertions?
I realize this is just one example, but you might be interested to know that I've done some work making this kind of invariant enforceable via static types (though in practice, you need a combination of features that aren't found in many mainstream language besides Haskell afaik).
See the README here for motivation and some examples (in Haskell, but hopefully the idea is still clear): https://github.com/matt-noonan/justified-containers
Or the tutorial module here: https://hackage.haskell.org/package/justified-containers-0.3...
Or this paper, if you really want to go off the deep end: http://kataskeue.com/gdp.pdf
This post had some interesting tidbits about technical challenges, but I didn't fully grasp everything on first read: http://blog.ezyang.com/2016/04/hindley-milner-with-top-level...
I really like tome's observation in the other comment: universally-quantified types have a direct encoding in System F as type abstractions, but there is no such direct encoding for existentially-quantified types.
It's about inspecting and verifying minimums at runtime.
Spec is designed so that you setup automated generative tests on your specced functions. These perform brute force search of the input space.
So, they won't give you FOR ALL guarantees, but will still catch quite a lot. For functions with small input domains, it would actually prove FOR ALL.
The trade off is that, you can test for much more. You can test for properties of the values, not just type. Like say making sure that the output is always smaller than the input. Or that the input never is a blank or empty string.
A downside, it doesn't work well for unpure functions. Since the generative brute force doesn't have a way to brute force the side effect, or assert properties about it. For those, you'd need to write your own tests.
All in all, don't expect it to be at all like a static type checker. It's a very different beast, which you'll want to use very differently, and which offers a very different value proposition. For example, you could want to use Spec even if you had a static type system. Just like people still write tests. Spec is a new kind of tool that can be leveraged to mitigate software defects.
Clojure's Schema is similar to spec, if a little different. There's also Typed Clojure, which allows you to introduce gradual typing to a Clojure code base and get compile-time static typing (though I believe it's still alpha/beta quality).
I would say that that the reason type errors do not hurt Clojure as much as say Javascript is because Clojure data is immutable, and all the common functions operate on collections or sequences. Basically, if you write idiomatic code, you constantly use datatypes that implement the same interface which higher order functions (map, filter, reduce, etc) expect. In contrast, in OCaml for example you have to think whether you are dealing with an Array or a List. Clojure also provides mechanisms to enforce that functions (Macros too with clojure.spec) are applied the correct data parameters (clojure.spec + pre/post conditions). Spec shines when you are integrating with "foreign" data like the results of a HTTP request.
In my experience, it is faster to make changes to a project, then test quickly with the clojure repl than it is to spend a lot of time refactoring types. This is especially true when you are experimenting with how to process data a different way... Having to have all types work out at compile time is a real time sink, when mentally you already know how you want to process data.
I'll just throw my anecdatum out here: I found that dynamic typing was a major pain even on my own personal projects. I also never found REPL-driven development to mesh well with my workflow.
In Scala, for example, entire classes of error that I just shouldn't be able to make don't exist. Maybe I'm just the kind of person who works better with static typing.
Clojure overall is an excellent language with really good features. It just won't get out of my way sometimes.
One issue with trying to adopt the habit midway through a project is having a design that makes it easy to load small portions of the project with test data into the REPL. It's not quite the same as having code that's easy to test.
Do you use an interactive debugger? It's the same principle, but earlier in the lifecycle of development.
This response is pretty common: "You probably weren't doing it right. Trust me, this is the way!" While I agree that suspending beliefs/old practices and trying new things is important when using new technology, I'm also pretty confident that I gave Clojure a more than fair shake, and just find that I don't get much out of the technology surrounding the REPL in that environment.
I was more productive in Scala in days than I was after years of using Clojure, and the difference between what I'm saying and your response is that I'm not claiming my experience is universal.
What people mean by REPL driven development in context of Clojure is that the REPL is running within the context of the application you're developing. The editor is connected to the REPL, and you can modify any part of the application at runtime. You can even connect the editor to a REPL on a remote machine such as a production server and inspect its state from your editor.
Having the entire application loaded into your repl is even more powerful. I did it once in PHP using psych and was able to make use of any controller and try parts of it out. The main thing that is necessary is the ability of the REPL to either drive change to the code, or react to changes on file (the former is temporal, the latter is more permanent) and perform the necessary incremental computation to become eventually consistent.
> In Scala, for example, entire classes of error that I just shouldn't be able to make don't exist.
guess what programming practice allows to almost entirely avoid those classes of error?
That said, we have a micro-service architecture. Thus the overall system grows horizontally, in that more and more components are built and developed. So individual components don't really grow that large. This would be different say to building a giant monolith like a AAA game, or a massive application like Photoshop. So I can't speak to how such a code base would scale.
Also, this might sound strange, but I don't think we've ever had to refactor the code base in those three years. The paradigms of Clojure (data driven/functional/immutable/meta), combined with our architecture, it just doesn't really need you to refactor things to add features.
Most adding of features tend to be... Write functions for it. Plug the functions inside the outer orchestration chains of function composition which dictates code flow. Add some more keys to our data-structures. Write tests. Release. Where as in Java, we used to have to shuffle all the Class/Interface arrangements around all the time to make room for new features.
Clojure is great if you like dynamic typing...by far the best dynamically typed functional language out there. Beware of claims that clojurescript and clojure are the same language. While both are dynamically typed, clojure's type discipline is strongly enforced (type errors will throw exceptions) while clojurescript will fail silently. I personally know of a team that experienced a million dollar bug because of this subtle distinction...you absolutely have to be prepared for it if you plan on using both. If you like gradual typing, there are some great options.
Haskell is good but the language's strictness (especially around effects) is very demanding. This can be both a good thing and a bad thing. In my experience/opinion, it's a good thing in some use cases and a bad thing in other use cases...but there's no way to opt out of it when it's a bad thing. Further, the ecosystem is far less useful than the jvm ecosystem.
Scala definitely has its faults, but its type system is extremely expressive without being extremely restrictive. I've never felt restricted by it (like I did with Haskell), but it regularly blows me away how easy it is to refactor and fix errors. The Java interop is far more intuitive than Clojure. Drawbacks: The only really useable IDE is Intellij. Some libraries tend to form an ecosystem of dependencies which don't play well with other ecosystems (ie some libraries are "Scalaz-only", some are "Cats-only", etc), which can be really annoying at times. Java interop isn't extremely straightforward in the case of collections (if that is a big concern, go with Kotlin which is phenomenal).
The compiler correctly infers and warns of the problem, but that is of no use if your inputs change types on you. In the case of the service I'm referring to, an external service started serving up json strings for large numbers, whereas it had previously been json numbers. Since there was no exception produced, it silently corrupted data for a couple weeks before it was caught.
And documented to what extent types are automatically coerced.
Also got the rational behind this: performance.
And that every new release of the compiler performs smarter and smarter type inference and can warn about more of these cases. Though it still misses some as of now.
Also, yes, any dynamic language suffers from difficulties with refactoring, runtime errors, needing lots of test coverage and not scaling well to large teams and codebases. IMHO
So, in the following, I'm not ranting at you, but I am saying I see lots of arguments for types that jump to QED without nearly enough evidence, with the "everybody knows" kinda explanation. And you happen to have put a handy list right in front of me when I had time to respond.
> ... difficulties with refactoring, runtime errors, needing lots of test coverage and not scaling well to large teams and codebases.
I was really impressed with Scala when learning or toy projects, but every time I saw it in production it was pretty horrid. Small, technical refactoring was taken care of by the IDE, but large scale refactoring takes the same amount of thought as before, but now there's 10x as much code to understand first.
There were less runtime errors, but not much less, and I saw a lot of code that didn't error, but didn't get done what it needed to.
When it comes to scaling to large teams, I saw lots of "The compiler says it's ok, and the tests pass, so commit!" where the code didn't make sense! The old thing about write code for programmers to understand first, and the machine to understand second...
Outside of certain specific situations, I just haven't seen the benefits outweigh the increased code weight and slower time to market.
Whats even better is higher kinded types let you assert semantics to engineers that are hard to do in Clojure.. this type must be appendable or foldable, this type must handle async, this type maybe missing and this list cannot be empty.
All these things are a joy not a burden once they start working for you
Static types are overrated. Okay, so you caught the obvious errors a few minutes earlier, but while you were wrestling with types, I ran 30 quick tests validating the behavior, and made them unit tests.
I've only ever had trivial type error bugs, all the nasty ones are behavioral or related to race conditions, or at the boundaries where types don't exist.
How about plain old OCaml? Jane Street seems to find some success with it.
See this comment: https://news.ycombinator.com/item?id=18345672
Primarily been developing with Clojure for 5 years now with some pretty large codebases. It does depend on how you write your code but favoring pure functions, pushing immutability to the edges of your programs allows you to refactor without fear.
Clojure also has many things to aid in this such as pre/post conditions, clojure.spec (which allows you to build complex type definitions), and of course test.check (property based testing).
I would suggest that if you have runtime bugs popping up in Clojure programs then that would suggest the inputs to functions (since they should be primarily pure) are not being validated which can easily be accomplished. I would imagine this needs to be done in Haskell as well since just verifying types does not indicate valid data.
For these reasons and others, I don't personally find run-time checks to be an adequate replacement for compile-time checks. But there is no really convincing research on the subject, and I certainly don't begrudge your preference here.
Well, the statement you were disagreeing with from the post you responded to was:
> You can't refactor Clojure without fear like in Haskell.
You go on to suggest you can achieve a similar experience in Clojure by "depending on how you write your code". This simply hasn't been my experience. Just "writing your code the right way" solves almost every problem that arises in programming, but just isn't always feasible in practice (on a team of developers with mixed skill levels, operating under deadlines, etc).
Re. validation, as Matt Noonan mentioned, in Haskell the goal is often to build/leverage correct-by-construction data types which obviate the need for any validation.
A simple example of this would be the `NonEmpty` (list) type.
If you have a function that pulls a list out of some key in a clojure map and it's intended that it always be a non-empty list you still need to check if actually is or not before using it because you have no control over what the caller passes to you. If the caller never sends an empty list you're fine, but if they do and you don't check for it, you've got a bug. Even if you do check for it, there's often nothing sensible you can do at that point since the local function shouldn't know anything about it's calling context, so you have to raise an error or return a nil or something.
On the flip side, in Haskell instead of using a map you would be likely to create a specific data type, and in that data type you would declare the non-empty field to be of the `NonEmpty` type. The first immediate benefit you get is that you no longer have to do any of these checks for the list being empty (or nil, or something else instead of a list) and instead just write your algorithm over the list. Among other things this results in cleaner, simpler, and less code in your immediate function.
But there's another benefit which is now anyone that calls your function has to have constructed a `NonEmpty` list before they call your function. The impact of this essentially naturally propagates the need to construct that `NonEmpty` list to the right place in the code. Maybe it really is just the caller to your your function that needs to take a regular (possibly empty) list that it has and package it up into a `NonEmpty` to call your function... and in that case you still get the benefits of code that shorter, simpler and more clearly communicates it's intention, but where this really shines is when that requirement to pass a non-empty list makes you realize something about the nature of your problem, and you let that `NonEmpty` propagate all the way out towards the boundaries of your application.
Then you end up in a situation where a) all of the code that touches that field anywhere is simpler, clearer, etc. but more importantly b) if someone does send you malformed data with an empty list (say via JSON over a web API or similar) then the code that deserializes the JSON into the `NonEmpty` will fail and you will get an error that says something like "Couldn't decode a YourCustomType from {whatever it was trying to decode}" at the very moment that the bad data tried to enter your system -- instead of just reading the JSON into a map because it was well formed and then letting the record with the empty list in it bounce around until it hits a function that assumes it's non empty at which point you may have little to no information about the provenance of the data or other details that would make easier to solve the problem.
Your examples of the type safety can all be mimicked with spec in Clojure. Granted spec is opt-in (but I'm guessing so is some of the more detailed type safety attributes you are talking about like NonEmpty).
Anyhow, to each their own and one persons experience isn't likely to be the same as the others so I'd encourage everyone to try out many languages. Some languages click with people more than others do so it's always worthwhile to experiment.
Every tool has a sweet spot and large code bases and maintenance is well outside the sweet spot of any dynamic language.
Many static typing arguments remind me of the Air Force's old "We'll bomb them so hard we won't have to send in ground troops." I just haven't seen it in practice, and the few studies that have looked at it empirically haven't seen a clear advantage either. If you know of such a study, please point it out!
In the end, all sorts of combinations have succeeded or failed, to the point where now when people start talking about "the right tool for the job", I add in "the right tool for the right people in the right environment for the right job..."
This is basically right; there are not a ton of studies, and the ones that exist mostly have pretty bad methodologies. Dan Luu summarized a bunch of them, circa about 2014. [1]
Since then, there has been one study in this area that I think has a solid, well-defined, and plausible methodology [2]. Plus a "Threats to Validity" section, sorely missing from many other papers in this area. They work from a corpus of real-world public bugs in Javascript programs and quantify how many are detected by simple type annotations via TypeScript and Flow. They cap the amount of time for trying to resolve a bug with type annotations at 10 minutes. The result is that over a corpus of 400 bugs, they were able to resolve about 60 using either of TypeScript or Flow, suggesting that 15% of Javascript bugs can be eliminated by using either type system.
That's not a huge difference, but it isn't trivial either. The authors quote an engineering manager at Microsoft: "That’s shocking. If you could make a change to the way we do development that would reduce the number of bugs being checked in by 10% or more overnight, that’s a no-brainer. Unless it doubles development time or something, we’d do it."
I suspect with languages that are more amenable to static types the results would be even better, but there is no solid research that I know of to back that up.
[1] https://danluu.com/empirical-pl/
[2] http://ttendency.cs.ucl.ac.uk/projects/type_study/documents/...
Wikipedia defines it as "the process of restructuring existing computer code without changing its external behavior", which sounds about right, but it's also so completely generic that it could mean almost anything.
Claiming that Clojure can't do 'meaningful refactoring' sounds to me about like claiming that Kanji is bad for transcribing Welsh, or that sign language doesn't work well for audiobooks. They're technically languages but the fundamentals are so different that all the comparisons are talking right past each other.
The style of refactoring that one does in Haskell (disclaimer: it's been many years since I've written any) is not really possible in Clojure, but it's not really necessary, either.
It's true that Clojure has its own perks that aid in refactoring, like how you may have less code altogether. But in my experience it's still a drop in the bucket compared to the zoomed out view of dynamic vs static typing.
For example, could you give some examples that make Clojure particularly good in this regard? I'd have trouble coming up with many that quantify favorably against compile-time analysis. And it would be unfair to take this reality and say "Clojure sucks at refactoring." I think that people do say that is what clouds these discussions where one then needs to point out that Clojure has advantages over other dynamically typed languages which is definitely true.
Fowler wanted to write the book using Smalltalk, but because the techniques he wanted to write about were fairly language agnostic, he did it in Java, as that was more popular. Unfortunately I think a lot of programmers missed the point, and now think refactoring is only something that can be done well in languages like Java (static typing) and with IDE support.
He's announced a second edition[0] in which he'll use JavaScript, to again prove the point that the techniques matter less than the language (other than sometimes techniques that work for class-based designs aren't as relevant as for function-based designs) and because JS is so popular. I'm not optimistic it will help anyone see the underlying point, but at the very least it might kill the idea that refactoring has to be hard in non-static-typed languages.
[0] https://martinfowler.com/articles/201803-refactoring-2nd-ed....
You're spot on. It just isn't really necessary to refactor Clojure code. In my many years using it professionally, I didn't have to refactor it once. The code just stays clean, it doesn't need cleaning up. So I find the whole refactoring argument moot. Why does your language lend itself to code that needs to be refactored? Sounds like a disadvantage to me.
Dynamic typing is punitive if we make mistake (e.g. using the wrong architecture, now we have to refactor), and we'll make mistakes.
Writing perfect code with perfect architecture and perfect test has been the argument "for" dynamic typing, but I don't think that's a realistic assumption.
Static typing requires all types to make sense together no matter how small the addition is.
In dynamic typing, we can just add a class/object/field here and there because it'll only be used narrowly anyway. So, it's often fine.
This would be great if we can think of a small example to illustrate this trade off.
One could write a comment here in Haskell describing why static typing is better, yet nobody ever does. I think even the most ardent static typing proponents are implicitly admitting that dynamic typing is fine, and other concerns can be more important, in some contexts.
Bleacher report obsesses about information push latency, and they use an all elixir stack that deploys to at least a million users.
Not to mention WhatsApp.
If most of what you're doing is a little minimal data transformation, but otherwise shuffling data to different network locations, and you want consistent latency/throughput, Elixir is a dream. If you're doing some tight number crunching, then not so much.
Why? Because most of my problems are things Elixir can solve well, and Java can solve moderately well, after a bunch of pain and tweaking. Default to the thing that can solve most of your problems well, and pick a different tool when needed, rather than default to the thing that can solve all of your problems moderately well, and you never pick anything else, and so never have any elegant, quickly built yet well designed, solutions.
My experience is that dynamic typing is problematic in imperative/OO languages. One problem is that the data is mutable, and you pass things around by reference. Even if you knew the shape of the data originally, there's no way to tell whether it's been changed elsewhere via side effects. Keeping track of that quickly gets out of hand.
What I find to be of highest importance is the ability to reason about parts of the application in isolation, and types don't provide much help in that regard. When you have shared mutable state, it becomes impossible to track it in your head as application size grows. Knowing the types of the data does not reduce the complexity of understanding how different parts of the application affect its overall state.
My experience is that immutability plays a far bigger role than types in addressing this problem. Immutability as the default makes it natural to structure applications using independent components. This indirectly helps with the problem of tracking types in large applications as well. You don't need to track types across your entire application, and you're able to do local reasoning within the scope of each component. Meanwhile, you make bigger components by composing smaller ones together, and you only need to know the types at the level of composition which is the public API for the components.
REPL driven development [1] also plays a big role in the workflow. Any code I write, I evaluate in the REPL straight from the editor. The REPL has the full application state, so I have access to things like database connections, queues, etc. I can even connect to the REPL in production. So, say I'm writing a function to get some data from the database, I'll write the code, and run it to see exactly the shape of the data that I have. Then I might write a function to transform it, and so on. At each step I know exactly what my data is and what my code is doing.
Where I typically care about having a formalism is at component boundaries. Spec [2] provides a much better way to do that than types. The main reason being that it focuses on ensuring semantic correctness. For example, consider a sort function. The types can tell me that I passed in a collection of a particular type and I got a collection of the same type back. However, what I really want to know is that the collection contains the same elements, and that they're in order. This is difficult to express using most type systems out there, while trivial to do using Spec.
[1] https://vvvvalvalval.github.io/posts/what-makes-a-good-repl.... [2] http://danboykis.com/?p=2293