FP vs. OO, from the trenches
blog.fogus.me
blog.fogus.me
HN has a tremendous reach these days, from pioneers in the field to total (and often clueless) newbies. Fogus felt compelled to write down his thoughts, someone else saw fit to post it and as of this moment 100+ people thought it was useful enough to vote it up. If that's not validation enough for you then remember that time when you spent half a day tracking down an endless loop and you had to ask a more seasoned programmer to point it out to you.
Fogus is a respected member of the community because he tends to write down what he thinks in a way that others much further down on the totem pole can grasp it too, his articles are not upvoted because he's 'well known' but mostly because people find them useful.
If you feel like this article was 'devoid of content' or whatever slur you feel like directing at it then please, write a better one.
The irony is your argument is just as flimsy and opinionated as everyone else's.
This article is devoid of content AND is written by a respected member of the community but neither of those things make it worth reading.
The great moral of the story here is, if you let the internet bother you, the internet will definitely bother you.
Some of the deeper links between linguistics and programming can be quite eye opening and it does not require a large number of words to put those down. If you find one of these, especially through your own work (as in 'from the trenches') then that should carry some weight with you.
If you're doing this independently and re-discovering something that apparently a lot of people such as you have already found out or learned about elsewhere then indeed it is devoid of content. But I'll bet that it wasn't the case for the majority of those that read it.
Lots of programming wisdom is so terse that we have acronyms for it (DRY for instance), that does not mean there is no content there.
Anyway, enough said, I think it is worth reading, and worth 'grokking' for want of a better word, in case you had not already discovered this. Naming stuff is one of the harder things in programming, that different streams of programming should lead to different groups of words being used to describe the code should probably come as no surprise and yet I find myself amazed that there are such deep connections.
Without a clear definition or concrete examples of what the OP meant by "simulate people" versus "data about people", the post is almost devoid of information (to me; apparently, a lot of people understood what OP meant).
They're not mutually exclusive: the inverse is not necessarily true. This is basically the entire point of the article. When you're just working with data, don't use constructs which are meant for simulation. It's a subtle way of saying "Don't model data about people (or any other piece of pure information) with mutable objects; prefer values instead."
And, if you respect that man's uncommitted, unexplained opinion, then his article is a good starting point for further research that draws conclusion.
The rest of his blog I found fascinating, however, and I'll be reading through it with great pleasure.
Software organized around nouns looks different than software organized around verbs. Not better. Not worse. Different. Use what is appropriate.
The root of that discussion was http://www.perlmonks.org/?node_id=318257 which I think is interesting in its own right.
most 'objects' as they exist in today's OOP have methods, which makes them both nouns and verbs. Idiomatic code in most of today's mainstream OO languages doesn't make heavy use of closures, mostly only Javascript.
Anyway, the closures vs methods debate has been going on for decades is not a closed topic[1], and the answer certainly depends on the language you're using.
[1] http://www.c2.com/cgi/wiki?ClosuresAndObjectsAreEquivalent
As a concrete example I point you at http://www.perlmonks.org/?node_id=34786 where I demonstrated how to take transform a markup language into HTML through the technique of grabbing a token, then passing to a closure that is associated with that token to consume a bit of the document. Porting that closure-based code to an OO style would be a nightmare.
Verbs/functions are easier to compose and recompose; they're easier to reason about (there is no "state" in verbs; verbs are declarative.)
Nouns (assume: with behavior; not just data -- otherwise you are in FP-land again) are harder to compose and recompose, harder to reason about (state transitions).
Consider calculating age. Assume general FP and general OOP ways of solving the problem. Then:
ageOf(p)... ...ageOf() can potentially be applied (or modified to apply) to different value types. ...self-evidently does not mutate p or cause a state transition (assuming FP).
p.ageOf() ...can only be applied to value types/classes that support that operation. ...may or may not mutate p/cause a state transition of p.
One is much easier to reason about and reapply when the world changes.
However let's limit ourselves to functional techniques. Whenever someone talks about something in the very abstract and says "easier to reason about" I mentally insert "for you". Certainly it isn't true for everyone or every situation.
A basic principle in programming is that starting at the hardware level we build abstraction layers that let us think about things in higher level ways, hopefully ultimately in a way that matches the terms that the end application uses. End applications have concepts like account balances, user status, and other such stateful pieces of information. If you're going to have the state, workaround that you pretend you don't, you will cause confusion. Pick your favorite 5 Monad tutorials for proof.
Consider as an example the use case, "Provide a form to let us mark spammers, and block them from our site." For all the theoretical advantages of avoiding state, this is going to be easier to model in a system where users have a piece of state known as is_spammer.
The idea about mixing OO and FP at different levels is a good one. Have had similar experience.
Anecdote-Driven-Development will justify almost any approach, especially to vague problems.
Thanks!
Be aware that this seems very much like criticism, even if it wasn't directed at the author:
> ...seems to be the only way posts devoid of value push forward around here.
I wonder how many people here do this (i.e. upvote topics for HN discussion, not for the submission itself). Personally I do that quite often.
For an article like this I'm split. I'm a passionate programmer and would love to read a good article on the tradeoffs between functional and object-oriented styles. I'd also love to read a good discussion about the two styles.
For some topics, say, a site being down, I don't really care about the discussion in most circumstances.
For many other topics, I'm more interested in a good discussion about the topic than I am in any particular article on the topic. It's harder for me to articulate which topics those are, but I know them when I see them.
You will get a "huh?" from the majority of developers management likes to hire as cogs.
Fogus: Say my name--
Declan: You're Heisenberg.
Fogus: -- to 1,000,000 Java programmers and--
Declan: WTF?
--
On a more serious note, the above is my way of giving props because Joy of Clojure is a really good book.
The idea of using an object oriented approach to simulation goes back to at least the sixties.
The way I read his post is that working on sets of information tends to be easier in functional languages event based simulations tends to be easier in object oriented languages.
Especially considering how many don't ever come around to publishing anything?
I actually don't think it's that bad, though...
I always thought that FP and OO were orthogonal concepts and never considered them mutually exclusive. To me, the antitheses of FP would be imperative programming. Maybe I'm missing something?
CLOS, OCaml, Scala, I can think of many examples of languages that include components of both FP and OO. Or maybe we're simply considering that OO means this special paradigm used in C++/C#/Java (that really only claims to be the true and only OOP by the commercial entities backing these languages)?
PS: I googled around a bit, this Clojure article [1] seems relevant.
OO comes from the imperative end of the spectrum (after all, objects are an abstraction over mutable state), whereas FP comes from the declarative end (a pure function is a relation between sets and mutability is a foreign concept).
However, most programming languages are impure and support both paradigms, so you can get away with thinking about OO and FP as orthogonal ways to reason about and structure your code; language syntax and semantics may of course favour one concept over the other...
Unfortunately, discussions on the merits of FP versus OOP tend to be asinine just because every participant has their own personal definition of both 'functional programming' and 'object-oriented programming', often based on implementations in a particular language. How does Haskell's idea of FP compare to Java's idea of OOP? What about Agda versus Python? Erlang versus Smalltalk? Does FP require purity, or types, or pattern matching?[2] Does OOP require inheritance, or private members, or classes? Without knowing these details, who can even tell what schools are being advocated or why?
[1]: In OCaml, objects can contain mutable values but they must be explicitly marked as such. An object with no mutable values is effectively a pure object, and there are many good reasons to use such an object—analogous objects are useful even in other languages not traditionally considered functional, e.g. new PureObject().someMethod().otherMethod() creates a series of objects which themselves needn't expose a stateful interface at all.
[2]: I'm not asserting that FP requires these things, but I have seen discussions where people clarify that languages can't be functional without function composition or algebraic data types or typeclasses or what-have-you, usually working off an informal definition of 'functional language' based on their language of first exposure.
I find it especially interesting because under normal circumstances an object with no mutable state is no different from a collection of methods (which can be encapsulated in different ways of course; say as an OCaml module).
So I am curious as to what advantage an object system that allows only pure objects provide.
data Set elem = Set { contains :: elem -> Bool }
insert :: (Eq elem) => elem -> Set elem -> Set elem
insert key set = Set { contains = c }
where c key' = if key == key' then True else contains set key
union :: (Eq elem) => Set elem -> Set elem -> Set elem
union set1 set2 = Set { contains = c }
where c key = contains set1 key || contains set2 key
emptySet :: Set elem
emptySet = Set { contains = (\ _ -> False) }
-- mySet corresponds to {1,2,3}
mySet :: Set Int
mySet = insert 3 (insert 2 (insert 1 emptySet))
-- myVal is True
myVal :: Bool
myVal = contains mySet 2
Depending on your definition of object-oriented programming, this could be considered an object or something else. William Cook argues that it's an object because access depends only on a public interface, i.e. I could add another implementation fromList :: (Eq elem) => [elem] -> Set elem
fromList list = Set { contains = c}
where c key = key `elem` list
which uses a different representation underneath the surface, but the previous implementations of union and insert will work with it because it exposes the same interface... which is not the case with ML/OCaml modules, which must select a single implementation. This is a simplified version of the example Cook himself uses, so I urge you to read the paper if you want to understand more. (On the other hand, people in the Smalltalk/Ruby/&c camp will say, "No, of course that's not object-oriented—you're not sending messages!" So... it's a complicated debate.)[1]: http://www.cs.utexas.edu/~wcook/Drafts/2009/essay.pdf also linked by https://news.ycombinator.com/item?id=6084465
If you have only pure objects, you can still perform open recursion and leverage structural polymorphism (including row variables). Additionally, your objects are just values and, because they aren't mutable, can be held for backtracking.
For instance:
```ocaml class virtual biggable x = object val x = x method x : int = x method virtual embiggen : biggable end
let o = object (self) inherit biggable 0 method embiggen = {< x = x + 1 >} end
let q f o = object (self) val x = o#x method x = x method embiggen = f {< x = x * 2 >} end
let a = q (q (fun x -> object inherit biggable x#x method embiggen = {< >} end)) o#embiggen
;; # a#x;; - : int = 1 # a#embiggen#x;; - : int = 2 # a#embiggen#embiggen#x;; - : int = 4 # a#embiggen#embiggen#embiggen#x;; - : int = 4 ```
Or in Scala, you have to mark object's variables either as `val` (immutable) or `var` (mutable).
Like /u/Fishkins said, the Scala course in Coursera gives some examples of the use of immutable objects ( https://www.coursera.org/course/progfun ). I am sure they are used in a lot of places in Scala code, not only in Odersky's course.
OOP is about subtype polymorphism [1] by means of runtime single-dispatch [2] and modularity, as in abstract modules or components that help with decoupling, a notion supported explicitly by SML, or exemplified beautifully in the Cake-pattern used in Scala.
Your definition is of course an equally valid one that might be more appropriate in certain contexts - it just doesn't fit my mental model of computation equally well.
Likewise I see FP as a reactionary mechanism to write fewer and simpler protocols, because when the data is immutable the problem can usually be greatly simplified.
I believe CLOS would beg to differ.
Multiple dispatch is one of the things I really miss about Common Lisp. I do like OOP, but often the most intuitive way to think about something is multiple dispatch, and trying to shove it into a single dispatch model can produce some really ugly code. I'm pretty sure that is the main reason there are so many anti-OOP zealots in the world.
Also, I didn't realize Dylan had multiple dispatch. I should take a look at it sometime.
The cake pattern has as much to do with encapsulation as it does crosscut composition.
I never implied mutual exclusion.
It's imperative vs declarative - FP is on a different level of abstraction, as a sibling to, let's say, logic programming.
It really depends on what you mean by FP. If you refer to lambda-calculus or to say, the head/tail decomposition so common in FP, then those are inherently sequential and not really declarative. If you refer to functors / monads, then you're on to something, but then again, OOP doesn't really imply imperative.
Good point, I think the common tendency to refer to procedural programming as either procedural or imperative interchangeably got me doing the same thing with functional and declarative.
I hope I hit the upvote button instead of the downvote button. I'm typing this on a shaky train and my browser will only zoom in so far.
VS.
...you already have data about a set of things and those things aren't going to be changing their own data.
Think, Video Game VS. Data Visualization
I'm curious, has anyone ever implemented a run loop in a functional style? I know there are recursive approaches to passing the time deltas and what not, but it seems rather difficult to make video games with a purely functional approach... Can anyone offer any insight?
http://prog21.dadgum.com/23.html
EDIT: but not that hard, according to http://prog21.dadgum.com/37.html ;)
The author's explanation was incredibly clear and I feel like I got a handle on it!
Just 'lift' the relevant game state and pass mutatations along in a simular way to how you might pass exceptions along with a promise based system... Then it eventually feeds to a makeSounds() or an updateAnimations() or whatever! That way the code can keep all related functionality in one place, instead of spreading sound or animation related methods all over the code base encapsulated in each object with its own logic! Brilliant!
Caveat lector: I'm not familiar with the current state of the art in gamedev, so this may not be new or interesting.
The gist of this is that you prefer composition over inheritance, which is a bit more functional than an OO inheritance hierarchy. It's not strictly functional and it's definitely not purely functional. But it's a step in that direction, as it lends itself to viewing your system as a series of transformations over data rather than a system of actors.
The second piece is also possibly only tangential, as it does have state. But it shows an interesting way to incorporate concepts germane to FP like continuations, and it does use a homegrown Lisp as its scripting language.
I would like to hear a more detailed treatment of the distinction between "dealing with data about people" and "simulating people". If these overlap (and surely they do), then you get into a circularity where your favorite paradigm leads you to see a problem its way, just as much or more than the problem nudges you toward a paradigm for solving it.
On the other hand, my experience matches the OP's that OO's sweet spot is close to its origins in simulation--where you have a working system that exists outside your program, your job is to model it, and--critically--you can answer questions about how it works empirically. When you aren't simulating anything, OO is cumbersome because it requires you to name and reify concepts that become "things" in your mind; these ersatz "things" continually draw attention to themselves and get between you and the problem.
This is a really underutilized design strategy. Much too often architecture is discussed in OO vernacular, and as a result is made much too complicated. I think we'll come to see OO as a low-level technique more and more.
[1] https://www.destroyallsoftware.com/screencasts/catalog/funct...
On the other hand, simulation is quite state-full, often has tons of little details that you don't want to expose to the world (encapsulation), the thing that you are simulating usually has a collection of varying behaviors that you can represent via polymorphism or composition, they have attributes, and the 'engine' is often easier to write if you use dynamic(virtual methods)/static(generics)/duck polymorphism to get your things to do their, well, thing.
Which is not to say that either is undoable or really hard in the other paradigm, but if I am simulating something I am often thinking of 'objects' and doing things to them; the OO paradigm helps you program about that in a very expressive way. Ditto for FP and functional programming - I'm thinking about filtering and transforming data.
That is just my view of the author's article, with a caveat. I've seen terrible, beastly simulations done in OO that basically thought that single inheritance is how you spell OO. A single type hierarchy bites you so hard: 'Oh, I'm make everything a widget with draw/eat/process/whatever methods'. Then requirements change, and widgets in this subtree need to do something you thought only widget in that other subtree would do, next requirement has you needing to reflect on who can do what, and so on. You end up with a huge Blob object at the top of your hierarchy, or you endlessly end up with huge blocks of RTTI/reflection to figure out what kind of thing you are dealing with, or otherwise try to work around the problem that the world is not easily decomposed into a single hierarchy.
In practice I find all of this terribly reductive. Need some small thing whose behavior varies in well defined ways? Use polymorphism. Have a bunch of stateful stuff that you need to keep track of? Encapsulation is nice. Want to perform some operation on a large collection of data? Look to transform(). Need to make a lot of inferences and conclusions? Perhaps logic programming will work there. Need to provide a framework where people can build larger systems? Hopefully you are using component or services based ideas. And so on. A programmer should have a bunch of tools handy, and use the best one for the job. On the other hand, just about every person I interview opines that we should inherit for re-use, and I silently despair.
I think this one is open for debate. There are other models for managing state which make things simpler to reason about than classical OO with encapsulation. One such example is Clojure's epochal time model.
I would also suggest you (perhaps through equal haste in writing) made a category error. Encapsulation != OO. For example, I can achieve encapsulation in C just by putting variables in my .c file, and not distributing the .c, but only the headers and a lib. I am not trying to nitpick, but wondering if 'classical OO' is part of your assumption.
In any case, I have never programmed in Clojure, and know nothing about epochal time models. It looks interesting enough, but is it a tool I can readily reach for if I am programming in C++, Python, or what have you? Will others understand it? Googling provides only a dozen or so relevant links. I think all in all I stand behind "Encapsulation is nice". It is nice, it is not the only way or necessarily the best.
Yes, indeed it was haste. Where I really want to draw a distinction is between values and mutable objects (which may or may not use encapsulation). Encapsulation is a leaky abstraction when applied to mutable objects because the hidden state of some object may impact other parts of the system in various ways. Values (and functions of values) are a much sounder abstraction because they are referentially and literally transparent.
know nothing about epochal time models
The epochal time model is a mechanism used to coordinate change in a language which otherwise uses only immutable values (e.g. Clojure). It provides the means to create a reference to a succession of values over time. This means that any one particular list is immutable but the reference itself is mutated to point to different lists over time. The advantage of this is that these references can be shared -- without locks, copying or cloning -- because the succession of values is coordinated atomically.
OCaml becomes interesting given the above due to syntax similarities. Scala/Clojure due to Java library support(?). Haskell because it seems to get a good amount of attention.
As far as data processing goes, I've seen both scala get some attention in the data game, because you can easily write hadoop jobs, and you can fairly easily interface with libraries like ATLAS for matrix work. However, working in a functional style in the JVM with lots of data can bite you sometimes, particularly with GC hiccups where the GC just wasn't designed to work efficiently with functional resources usage.
Personally I think Haskell is very well suited to data work. You get pretty good speed out of the box, concurrency and parallelism are baked in, and the library support is pretty extensive. It's also quite cross platform. In addition to that, you've got brilliant folks like Simon Peyton-Jones working on stuff like automated stream fusion, which can mean huge speedups. To me the main disadvantage to using haskell for data processing is the relatively long compile times. Since "data science" is typically fairly interactive, until you are pretty good at haskell you might find yourself waiting a long time to realize you've done some silly things. It's really a situation where things liker IPython's html notebooks are invaluable.
Sidenote: As far as string processing goes, I've found working with haskell's parsec library to be simply great for creating parser combinator libraries in a very natural way.
Well... and this post, now.
I always come to the comments to throw a little balance or extrapolation into my reading, and its disappointing to find people mostly just complaining.
Got a few interesting nuggets, though. So not all bad.
And whats the difference between e.g.
agents[0].location.x += 5;
versus
(update-in agents [0 location x] + 5)
I.e. a language that keeps a clean orthogonality between the three dimensions of a program --state, functionality, and event handling-- not favoring one over the others. The hard part is to make hierarchical modularization mechanisms work simultaneously along all three dimensions.
These are (mostly) good tools, but I still avoid hierarchy more than one level deep—in my experience it just introduces more complexity than it saves.
First, working in a single language allows you to accumulate what, for lack of a better word, I call "IP". Components/libraries/frameworks; a body of work. Having a single language gives you the leverage of previously written code solving prior problems in a debugged fashion.
Second, having a single language allows easier social operation; people can review each others work, a common body of knowledge can form around the language under common use which is difficult to maintain for multiple languages simultaneously.
Third, having a language which is a bit of a melting pot allows idioms to be used in which people are comfortable with their specific idiom - OO/FP, etc.
Many languages implement multiple flexible looping/filtering structures to discourage shoving everything into for/while loops (which are susceptible to confusing and messy continue, break, goto, yield statements strewn about). Furthermore the rationale behind Clojure[1] is stateful objects are a new kind of spaghetti code, and managing state across scopes requires breaking mental boundaries.
Of course, it's not impossible to have referentially transparent objects. I wrote such a language (Reia)
> most OO languages allow two objects that represent the
> same states to have distinct identities, even if
> compare-by-value claims they're equal.
You can do this in languages that offer only compare-by-value by attaching a unique ID to every object. Of course, you may have to write your own equality relation if you want compare-by-value semantics in addition to compare-by-identity semantics.Lots of languages have concepts from multiple paradigms. What is the correct mix of concepts is up for debate.
I honestly believe that Rust has the potential to become the perfect mix for me, but OCaml and D are pretty good second places with the benefit of exponentially greater stability/maturity.
OCaml - basically only gives you pattern matching on primitive types. Object types end up having dramatically different code style
Scala - any syntax that deals with types quickly becomes so complicated that you can't explain it to non-experts
I don't know much about F#.
I've heard that Clojure's core.match grants pattern matching over abstract interfaces in a nice style. It's not an official part of the language, yet, though.
Guidelines seem sensible though.
I think we're wrong on one thing. Personally, I think imperative code is more intuitive. It matches how people think of operational behaviors.
For a 20-line program that will never grow, I'd rather see an imperative straight-line solution. FP wins when things get a lot more complex: various special cases and loops within loops. Stateful things compose poorly. Debugging a 300-line inner for-loop is no fun. Functional programming, done right, means people factor long before it gets to that point.
Factories and Visitors prove what hellishness comes into being when one doesn't have those functional primitives, but this does not make the superiority of FP intuitive or obvious.
Now, OO is so variable in definition that it's very hard to know exactly what people mean. There's the good OOP (Alan Kay's vision; encapsulate complexity) and then there's bad OOP (overuse of inheritance, auto-generated classes, Factories and Visitors).
One can conceive of core.async as an OOP win in Clojure; lots of complexity (macroexpand a go block at some point) is being abstracted behind the simpler interface of a channel.
The masses agree with you, apparently. That's why they sit in their cubes all day writing the same god damn re-implementation of map over and over again
List<String> result = new ArrayList<>();
for (int i = 0; i < list.size(); i++) {
result.add(input.get(i).toString());
}
instead of input.map(_.toString)
because functional programming is confusing.I still don't understand why recursion is confusing to learn, as for me it was obvious from day one. Then again I am good at maths.
Additionally, I had lots of fun doing Caml and Prolog when I was at the university.
So even though I tend to do the typical boring enterprise JVM/.NET/C++ stuff, I do welcome the FP contamination of those ecosystems. :)
However on my last project, I had to rewrite some LINQ stuff back to plain imperative code, because few people on the team could manage it (LINQ). :(
The solution using map is more intuitive, once you know what map is.
I agree with you, as well, that the functional style should be used far more often than is the case. I just think that one shouldn't assume that our way of doing things is inherently more intuitive, even if it is most often better.