Simplicity Matters – Rich Hickey (2012) [video]
youtube.com
youtube.com
If I had a time machine and could make one change in an open source project it would be to go back and remove the design choice of putting using hashes as internal config objects in Rails. It makes understanding all the possible edge cases almost impossible unless you're really familiar with the project.
The second is the claim of speed that this "simplicity" buys you. I agree that functional programming is extremely fast when it's parallelized and it's extremely easy to make a functional program parallelizable. When we're dealing with sequential tasks, mutating data is much faster than allocating new data structures. I think clojure helps with the speed of allocating new immutable structures by using copy on write strategies behind the scenes.
I think Rich is an extremely smart and very accomplished programmer. I think functional programming is really good at some things, I don't feel like many people talk about the things it's not good at. To me we if we're not embracing and explore all a new concept/language/paradigm strengths and weaknesses, we're not growing by being exposed to that thing.
As someone who used to swear by Ruby but now prefers FP langs, part of learning FP is as much about learning what FP is good at as it is about learning what FP isn't good at. Given FP's proclamation for declarative style and immutability, clearly FP isn't good at domains where imperative styles are important or really, really, really[0] CPU intensive work where mutability is needed.
Although, thanks to compiler research (which is mostly done by FP language users, ahem) languages like OCaml are bridging this gap when you take a look at projects like MirageOS. You get to write your code in a declarative, immutable style that then gets compiled to very fast native code. It's having your cake and eating it, too.
As far as your original complaint of everything in Ruby either being a hash of hashes or a class that just wraps a hash of hashes, I completely agree. That is what drove me away from the language. I'd rather just blow that hash + class relationship up completely. Not to mention, in ML-dialect languages you can use features like Enum/Sum/Product types to define the configuration/syntax of your programs which is a lot better than reading arbitrary keys in a hash, as you have stated.
0 - No, I mean really, really, really, really. "I think I need C to do this fast enough for my use case" is the "I need Cassandra for my 2GB database of 'Big Data'" of the FP world.
No, it's not copy on write in the sense that COW means in other languages. Data in clojure is persistent, so usually even an altered piece of data is not actually copied, only the tiny bit that changed is added, if necessary. This is very different than, for example, Swift that implements copy on write, but the whole data structure is copied even if the change is minor.
Additionally, you say:
>When we're dealing with sequential tasks, mutating data is much faster than allocating new data structures
That would be true in cases where the entire data piece is new; but so much of what you do in any language involves interative changes over existing data, and again because of clojure's persistence, new allocations are often not happening at all in many cases.
...However, I do agree with most of your other points :)
That part about speed was more about my general frustration with the "immutable is fast" and "mutable is slow" meme that I hear too frequently. While that can be the case, it isn't always 100% true. Your language and how you use it can play a huge part.
I'm hoping to learn more about clojure data structures in the coming months, so much good stuff.
I don't think anybody is arguing that trie lookup is faster than array lookup. But when it comes to passing data between parts of the system, zero cost copies are much faster than copying buffers or objects.
Not if your data structures are immutable, which is exactly what he is advocating.
I think my other point still stands, things like optional keys, or missing or required keys still must be shared throughout the app. Everything that touches that hash needs to know its structure.
And in case you want some security you can use libraries like prismatic/schema, that allow you to declaratively describe your expected data-structures akin to a type system.
The web is build on json and not corba because plain maps and vectors are far easier to work with than domain specific objects.
There’s tradeoffs. Shallow datastructures are under-used.
But OOP objects share fundamental flaws with deep hashes. (And add some more.)
The two major sources of complexity are the fact they are mutable and the fact that keys are mutable but hashes are computed when the object is first added.
So if you construct a valid hash, then call three functions passing them the hash, you have no idea if the hash is still valid for the second and third functions without reading the code for all the preceding functions.
What this tends to mean in practice is you are always having to check all over the place that your hash is valid.
They might look the same as clojure hashes but in practice they are used very differently.
I agree here
> So if you construct a valid hash, then call three functions passing them the hash, you have no idea if the hash is still valid for the second and third functions without reading the code for all the preceding functions.
I don't find that to be true, even in the Rails codebase. You pass a hash to a function (method) and expect that it won't be mutated. If you're writing a method and you mutate a hash, you're expected to dup the argument so it won't be mutated. This is the convention. There are times when hashes are mutated, but generally that's reserved for methods who's purpose are to mutate its arguments.
Where Rails gets into trouble in its functional passing of hashes isn't in the mutability of the hashes, it's in the composability of the functions. I've never written about this problem before, so i'm not sure I have a great example but it's definitely a problem.
One example in rails is `url_for` it takes a hash argument. The problem is that there's multiple `url_for` methods that all do slightly different things. Some are needed for generating links in email, while some work in your views, and others are designed for you to use programmatically outside of a view context. One of the hardest things about this method is that one `url_for` can call another `url_for`. Since we can never be guaranteed the order of the calls it is really it makes things like having default values, or optional keys. You have to replicate logic in different functions since some may never be called in the order you might expect. This significantly impacts our ability to refactor which in turn impacts our ability to make performance improvements.
I recently did a bunch of perf work in https://engineering.heroku.com/blogs/2015-08-06-patching-rai... and some of my biggest perf improvements were getting rid of duping and merging of hashes. If Ruby had an immutable and performance efficient hash then it would have helped a bunch, however I don't think it would make the general awfulness that is an entirely hash based API to a very complex action (such as url_for) that significantly better to work with.
Because it passes every request as a map you can hook functions (middleware) in between that change the behaviour of the request handling, for example they could add a field for passed params or fields for user authentication.
This is possible because maps are easily extendable and functions down the line don't need to know about additional keys. You can't do that with OO properly.
Edit: It's one of the first thing he says in the video actually - "It's about the interleaving, not the cardinality."
Rich absolutely argues that basic data structures like hashes are simple -- something that is generally agreed upon. They are simpler than 'objects' because you can use basic comparison operators on them, and you can operate on them with higher-order functions etc.
The question which parent is digging at is whether a larger, complex app, that leans heavily on hashes actually results in an app that is on the whole simpler, i.e. does 'a simple thing plus a simple thing equal a simple thing'. This is something I've heard discussed well on the Ruby Rogues podcast -- I think that either David Brady or Josh Susser may have a good blog post on the subject of 'simple + simple != simple' but I'm struggling to track it down.
Whilst I love Rich Hickey's talks I do find myself coming round to the same conclusion as parent -- if you forgive my possibly incorrect interpretation of their argument -- that the idea that simple data structures are simpler is fairly useless if programmers use that fact naively.
PS. Everyone should watch the linked talk, it's brilliant, and one of my favourite programming talks ever. I recommend it to every programmer I meet.
Why not just write accessors and mutators for your hashes? Also, Python hash a function dict.get(key, default). Does Ruby not have this?
This is a wonderful tenet in general: let's think about what X is bad at, rather than what it's good for.
While I'm not much of a Rubyist (I've hacked a little in my time, but I fall on the Python side of the divide), working with trees (hashes of hashes are a kind of tree) has never been easier than in Clojure in my experience. I suspect this is true of any lisp.
(def g {:x "a" :y {:k 71}})
user=> (assoc-in g [:y :k] 72)
{:y {:k 72}, :x "a"}
With this you can walk and mod deeply nested datastructures with ease (and others can reason about it well).
What natrius says is completely right on this point (which explains why those functions don't exist commonly outside of lisp) but it should by now make you wonder: if that's true, why haven't they?
The rest of your points have some merit. :)
Go learn Clojure.
I understand if you don't want to take the time, but I would love it if you could you give an example of an abstraction that you can build in Closure that you would consider "easy to understand", yet which couldn't be built just with, say, anonymous functions, structs, arrays, and simple loops?
Macros allow things like core.async (like Go's channels), which is just a normal library; you didn't have to upgrade your Clojure version or anything.
Immutable datastructures make your life simpler because you're not worried about values mutating suddenly. Keeps you from cloning or locking an object. And undo is simpler: you don't destroy old state by mutating it, so you can just hold onto old versions.
I'd say a specific example would be pmap. It's very difficult to parallelize code as simply as pmap does without functional programming constructs.
I'll dig in some more though. Thanks for the reference.
[1] https://github.com/clojure/clojure/blob/master/src/clj/cloju...
I could write a few simple routines in Go that certain get the job done without much code, but when I read the code I have to perform more mental translation from how the code is written to what my intent was.
In contrast, pmap or other constructs such as PLINQ get the same work done with less code that expresses my intent more clearly.
Besides, I know lispers hate the idea, and it may not even be necessary, but uniting behind a single good-enough lisp like Clojure will reap more rewards then advocating for multiple (while still good) Scheme implementations.
It's an uphill battle to advocate for something that isn't the status quo, so I'm proud of the little victories.
edit: state of adoption and perceived coolness, not birth year. Coincidentally, Python 1.0 released in 1994 and Clojure 1.0 was released in 2009, and both had a couple years of unstable releases predating 1.0
Clojure macros are the antithesis of simple, and the need to indent a scope for every new variable actually fights against TDD in my experience.
I recently wrote a good chunk of a RESTful SQL-backed application in Python in two days that took a team of 3 people in Clojure over 2 months to just get to the basic level of library support one would expect from things like sqlalchemy and flask.
Clojure isn't simple -- it's basically a step up from assembler in how little it provides.
Simplicity is having all the power tools and being able to put them together and be instantly productive, and to support programming in multiple paradigms.
While it's not the norm, I sometimes feel many FP purists spend so much time debating purity and giving basic concepts complex names - when they could be using something else and getting much more done.
Side effects aren't the devil and are sometimes neccessary to get real work done. Bad code can be written in anything, and it just takes experience.
I'd much rather see a language focus on readability, maintaince, and rapid prototyping than side effects.
Functional programming concepts have benefits - I love list comprehensions and functools.partial in python is pretty neat, but when you can also have a decent object system, and embrace imperative when steps are truly imperative, you can get a whole lot more done.
For example, under this definition, the simplest thing possible is immutable data, since nothing can affect it. The next simplest things are pure functions, because they're affected only by their arguments. Clojure is a language built around Rich's idea of simplicity, so Clojure prefers data over pure functions, and pure functions over side-effectful functions.
What you're describing is what Rich would likely term "easy". Something is easy if you can do it with little effort. Something is simple if few things affect it.
My approaches to Clojure have been seriously hampered by the fact that some of the abstractions above those that are "simple" are remarkably complex, and that the tools that surround the ecosystem are still pretty frail.
Macro bugs are certainly something that have scared me away for a while.
My experiences were around looking for a quality ORM, job scheduler, and web framework - things like korma and ring exist - but they lack a large amount of features compared to equivalents found in /most other/ languages.
I came to the conclusion that Clojure is an acceptable way to call Java SDKs if you want a bit higher Java velocity AND Lisp fits your brain already, but I'd rather pick up Scala or Groovy instead for that purpose.
Libs in pure clojure, for which I tried dozens, were usually incomplete and error-prone even if they were community favorites, which I attribute in part to the fact that it's a small circle of developers using it, and the language is still newish.
Nonsense. Java library interop is great. If you could write it in 2 days in Python then it wouldn't take more than a week in clojure. In the very worst case scenario you could simply write "Java in Clojure" and directly use Java libraries.
> Clojure isn't simple -- it's basically a step up from assembler in how little it provides.
It gives you full access to the JVM and its tens of thousands of man years worth of high quality libraries.
As for the language itself it has macros, first class functions, built in vectors and maps, structural editing, fantastic REPL support, and a top rate concurrency story to name a few features.
- Inside Transducers (11/2014)
- Transducers (09/2014)
- Implementation details of core.async Channels (06/2014)
- Design, Composition and Performance (11/2013)
- Clojure core.async Channels (09/2013)
- The Language of the System (11/2012)
- The Value of Values (07/2012)
- Reducers (06/2012)
- Simple Made Easy (9/2011)
- Hammock Driven Development (10/2010)
- Are we there yet? (09/2009)
https://github.com/matthiasn/talk-transcripts/tree/master/Hi...
Here's what drives me nuts about this one, though - it gets passed around a lot where I work, and people say how strongly they agree with him. There are real, concrete things he claims are not simple here! Things like for loops. And these people I'm talking about, they say they love this talk and then they say they love for loops. Really! I don't get it.
Now, personally, I don't have much to say about for loops one way or the other. But I do disagree with him on types, which can be far less complex than he intimates. System F, say, which largely covers F#, OCaml and Haskell, can be completely defined as a handful of self-evident rules. A single page for both checking and inference, if in somewhat dense notation. That's not complex at all, especially since it follows fairly naturally from the way the lambda calculus works even without types.
To me, that seems like a perfectly consistent view. Nothing forces you to take everything he says or leave it: you can find it accurate piecemeal. Take the mental framework but apply it with your own knowledge and experience and you could very well come up with your own conclusions.
Seems like the perfect way to use ideas like this.
Some links:
* Definition of Simple (2011) - http://www.rebol.com/article/0509.html
* Fight Software Complexity Pollution (2010) - http://www.rebol.com/article/0497.html
* Contemplating Simplicity (2005) - http://www.rebol.com/article/0127.html