Stop writing classes...
pyvideo.org
pyvideo.org
3DPoint = collections.namedtuple('3DPoint', ['x', 'y', 'z'])
p1 = 3DPoint(3, 2, 4)
p2 = 3DPoint(10, 1, 10)
dist = math.sqrt((p2.x - p1.x)**2 +
(p2.y - p1.y)**2 +
(p2.z - p1.z)**2)
I don't mean to say that classes aren't useful in Python, just that I've gotten a lot done in the language just depending on hashes, lists, tuples, named tuples, functions, comprehensions and generators. I've never felt the need to define a a class - the closest I came, I realized I wanted a named tuple. It certainly would make sense in the above example to define "dist" as a member function of a 3DPoint class, but it's just not how I write code in Python.Contrast this with C++, where I define classes all the time. I think the reason for this discrepancy is that if you want to do anything interesting in C++, you almost have to define a struct or a class; that's where most of the abstractions happen. In fact, if I want functional style code in C++, I have to define my functions as classes. I think that Python's native tuple types, first-class functions and comprehensions make the difference.
Freud said sometimes a cigar is only a cigar. I say sometimes a function is only a function.
Classes get in the way between the programmer and the problem. They force you to reify your thinking into "things" (classes) that demand names and citizenship rights, so to speak, in the system. These "things" don't really exist, but we act like they do, and so increasingly see the problem in terms of the classes we've defined. This is a distorting filter that we can ill afford when the problem is complex (if it's not complex, who cares - anything will work ok). The solution is not to define such pseudo-things in the first place, but rather to focus at all times on the problem at hand. Functions, for me, are the right level of abstraction to do this. They don't get in the way.
Once a bunch of classes have taken residence in your system and your brain, it's hard to remember that they aren't real, and thus the pain threshold is quite high before you realize that they're causing problems. At that point they're hard to change: you have to replace the old class hierarchy with a new class hierarchy. The difference between that and reworking a few functions is the difference between working in concrete and clay.
There's also the cognitive tax of having to think about what classes to create and what to name them. This too is overhead we can ill afford given our limited mental resources. If you've ever found yourself putting some code in a new class (perhaps because the old one got too big, or because you need to use it from two places) and gotten stopped in your tracks by the question, "What should I name it? What is this thing?" you are experiencing the overhead I'm talking about. It's the programming equivalent of Mr. Heavyfoot: http://www.youtube.com/watch?v=R9d2Y1We-d0. There is no such thing; you're just working in a way that forces you to pretend that there is and fulfill a bunch of duties towards it.
These two issues together are a double-whammy: you pay up front for what you pay for later.
It is shocking to see how much simpler systems become when we stop doing the things that make them complicated.
Classes are nothing but:
1. a group of data attributes, and
2. a defined set of operations that can happen on instances
of those data attributes.
If you are opposed to classes, which of these characteristics are you opposed to? Since I think it would be pretty much impossible to be a programmer and be opposed to structs or tuples, it must be the latter that you don't like: giving structs an associated set of methods that operate on them.But having a set of methods that define the state transitions that are possible on an object is incredibly useful, for so many reasons. It helps organize your code. It provides nicer, more intuitive syntax. But most importantly, it reduces program complexity by narrowing down the set of code that can directly modify the attributes of an object.
Without this encapsulation, any piece of code anywhere in the program could be modifying data attributes in any way, and it is extremely difficult to reason about the state transitions that any particular struct could be taking.
There is no rule that says classes have to be in any kind of hierarchy at all, and in fact "prefer composition over inheritance" has been common wisdom for years now.
I recently rewrote a very useful but hard to follow JavaScript library in a more object-oriented style. The old code has a bunch of maps everywhere that can be read or written by any part of the code, and it's very difficult to follow. In a more object-oriented style the structure of the code is far more obvious, and the state transitions that each object can take are far more clear.
I don't accept your breakdown of what classes are "nothing but". You've omitted at least one critical thing: the name. But even if the list were complete, it's a fallacy to say "if you dislike the combination, you must be opposed to one of its components".
(The name is critical because the human mind leaps irresistibly from names to things, something you appear not to be considering. But I already wrote about that.)
Organizing code, making it nicer, reducing complexity? all I can say is that my own programs got far better when I stopped organizing them into classes. They became shorter, easier to write, easier to test, easier to change, and more fun to produce. But we're in YMMV territory here. I do feel like stating, though, that I worked for years in the style you describe. I even taught it. And I said and believed many of the same things. Should that count for something? Maybe not. Maybe I just got bored and went off.
Encapsulation? The most grossly overrated allegedly simplifying mechanism ever. But "any code could be modifying blah"? Any code can do a lot of horrible things. Trying to rely on technical constructs to prevent it just adds weight and bloat and impediments. Let's have flexible languages and be good programmers.
Difficult to reason about? The hardest code I've ever tried to understand has been complex object models where Thingy depends on Fooey which needs a Bingy and a Batty, and to construct a Batty you need a... Compared to this sort of conceptual glob, bad procedural code has always, in my experience, been easier to understand.
Inheritance vs. composition? Not relevant here. Maybe I should have said "graph" instead of "hierarchy". Whether A "is" or "has" a B, that's still an edge in a graph, and it's those object graphs I'm talking about. They are much harder to rework than functions. The trouble is that when we believe that programming is making object graphs, we assume that difficulty to be part of the problem and don't notice it.
Rewriting? You rewrote something to be clearer and more intuitive to you and maybe to others who share your beliefs. That it happened to come out in objects is a consequence, not a cause, of what you find clear. It's a lot harder to judge clarity across assumption sets. For example, I work a lot in JS and defining JS objects is the one thing I never do. Maybe if each of us put our code in front of the other we'd recoil to exactly the same distance :)
Then, fairly recently, I dabbled in some OO programming when I was making a mod for Unreal Tournament (I've also coded in C++ in the past, but not extensively); and I realized that I had already been using the same line of thinking with my functional programming in the way I was using structs, tuples, etc. When you really get down to it, it's really all just syntax and semantics.
For whatever reason, my brain has always naturally leaned towards the functional style; and it's just my opinion (of which gruseom seems to agree) and my opinion is probably only the result of how my brain works, but something about OO programming always seemed unnecessary for me. It seems to me like it's just another way for people with different thought processes than my own to represent their program logic. To me it seems a little too verbose and entirely too repetitive, and I always felt like the overhead required to "classify" everything wasn't worth it; but it's all in how you look at it, how your brain works or what you're used to. So what seems unnecessary to me (and makes it harder for me personally to trace the program's logic) might be perfectly comprehensible code for someone else; and some of the relatively minimalistic code I attempt to write might seem perfect to me, while an OO-minded programmer might feel like it's scattered about and hard to understand.
Good code is good code, regardless of the paradigm.
> but something about OO programming always seemed unnecessary for me.
I actually learned to program making mods for Unreal Tournament (1999). I'd love to hear more about your feelings about how OO seems unnecessary.
For me, UT was a prime example of OOP (both the general bundling data and functions together as well as inheritance) being used to simplify things. Key is that any instance of 'actor' is a single physical object rendered in the world. A subclass of 'actor' is "inventory" - items that can be picked up and held by players. Below that is the "weapon" class; which a player can switch between, fire, and are rendered in 1st and 3rd person views.
Obviously you don't need to be OO to pull off a game; many FPS games are not. But it makes things so easy to handle. If you build a new weapon, all of the functions to handle picking it up, rendering it, etc. are written in super-classes. But if you want, you can over-ride those. Furthermore, the weapon interface is well-defined to external entities; e.g. the 'playerpawn' knows his weapon has a 'fire' method.
I'm not sure how you could build an equally powerful and customizable engine without using OOP techniques at some level. Sure a weapon could be a struct, but how would you know what associated fire() method you would call? Unless you want a bunch of non-extensible case statements, you'd need a function pointer in the struct itself.. and once you couple the data and methods, you are getting close to OOP.
One thing I loved about UT compared to Quake-engine games was its extensibility. OOP allowed reflection; just 'spawn weaponpackage.weapon' and you can use a custom weapon. This also allowed gameplay "mutators" to blend together. Quake games in the day required a custom -game argument at startup; and consequently custom content could not easily be merged together.
Of course, this is how I learned to program so my brain may be overly-wired that way.
The influence that is driving the new stuff - component systems and composable entities - can be seen as derivative of the actor model. Actor models go beyond what is present in most game engines, which tend to stay well on the imperative side of the line. It takes a heavy dose of purity to grok actors in full, but it dovetails well with functional style as an alternative view of computation. (I also recommend the Tim Sweeney POPL06 "Next Mainstream Language" talk to see exactly what problems are motivating the changes)
First, aim to never store values on objects - prefer instead listeners that return values independent of the objects themselves. Suddenly you start to enjoy "code=data" just as in Lisp, because the semantics of getting data and processing data are identical and interchangeable - you can take a getter on one object and copy it to a different one, and they'll return the same data, even if their other methods are completely unrelated!
Decoupling occurs naturally in such a system - instances can be as similar or different as desired. The usage of classes shackles these abilities into monolithic, heavily-coupled data structures again, which is the core point of difference from "OO" as we know it in industry.
If you also stay purely asynchronous at all times, there is no longer a difference between "now" and "later" to the callee, so order-dependent operations can naturally transition from one-to-one mappings into arbitrary ordering and queuing. Pure asynchronicity was lost when actor-style systems first tried to enter industry(e.g. Smalltalk-80 opted for a synchonous system - a concession to single-threaded performance), but it's come back into vogue in this heavily-parallelized era.
Fortuitously, it's straightforward to achieve these properties in languages that support closures and reflection. And you don't have to write entirely in this style; raw data structures can be written imperatively, and then tied together using actors. The listeners can access a common database to look up values. You can mix and match.
State-of-the-art game entity systems aim to cut down this level of power into a palette of performance/power options for composition, but if you can live with high overhead, taking all the power is pretty simple to do. I wrote a prototype for it yesterday - 169 lines of Haxe.
The actor system (Haxe highlighting is broken): https://github.com/triplefox/triad/blob/master/dev/com/ludam... And example: https://github.com/triplefox/triad/blob/master/examples/Sour...
This restriction was often frustrating, but limited types meant the language was very approachable for beginners. Once a programmer understood the entity and its attributes, they knew what they had to work with. They only needed to find interesting ways to transform data, not interesting ways to organize data.
While sometimes frustrating, the simple type system rarely limited the creativity of developers. This is evident in the wealth of mods that were created for the game. People built Diablo style RPGs [2] with it, racing games [3], and much more.
In Quake's sequels, the modding system grew more flexible (and also more complex). The number of mods seemed to decrease. The amount of mod content for Quake 2 was less than 1, and Quake 3's was far less than 2. The mods that were released were often far more polished, but I think this is representative of the skill level and determination that was required to actually build mods for those games.
Garry's Mod [4] is only modern game I've seen with a similar level of successful modding activity. It also uses a scripting language (Lua). Garry's Mod defines a strict set of structs that you use to interact with the host game engine (effects, entities, tools, and weapons). It feels very Quake-like.
I personally feel like these kinds of restrictions create flatter code, and structure that is easier to hold in my head. I find myself concentrating more on features, rather than on code organization.
[1] http://www.inside3d.com/qcspecs/qc-menu.htm
[2] http://www.inside3d.com/prydongate/
[3] http://www.quakewiki.net/archives/qca/reviews/patch70.htm
But, outside of simulations, a lot of programming problems come down to things that don't have a convenient real world mapping. A lot of it is just data transformations, and data transformations are awkward in an OO setting, because the focus is on the object rather than the process. This is where more functional styles really shine, (for instance, in compilers and such), where the existance of an object would be incredibly transient and short lived, because the goal is to transform one set of data into another representation. OOP basically sucks at this.
Objects are certainly useful, but I think they've been fetishized to a weird degree.
It also tends to work for things that can be represented as opaque handles (such as files, sockets, ...). Polymorphism helps there to be be able to treat different but similar objects in the same way. No need for deep, nested hierarchies there though.
But I agree that outside the domain of UIs it quickly diminishes in value.
Once I move past the UI though I really start to feel the burden of OO start to drag me down as far as productivity goes. It requires a lot of organizational overhead to force overtly service based functionality into OO structures, something like getUserList is functional at it's simplest form and should remain so. When working with Java I tend to discard OO by creating static functional classes that acts as a service and util layer. I find that doing so simplifies the architecture of the system and eliminates a lot of spaghetti code that has to be written to deal with making functional logic OO.
For that reason I tend to prefer languages that can do both, or that don't force a pattern on you at all, and that leave it up to the library developers to work out patterns. JavaScript despite it warts is a good example of the library developers building out that layer of the language. I also find myself writing more and more back end code in Clojure for that reason, while it is a functional language it does not preclude the ability to write OO code if it is the right pattern for a particular problem set. I like not having to fight the language to break free of the constraint.
From a business apps point of view, this is exactly it. You might find that methods are subordinate to objects, but in the business domain itself, the objects themselves are manipulated by processes. It's a whole other layer above the OO layer, and failure to realise this means you either end up with lots of FooManager classes, or you push overarching responsiblities back down into low-level objects that don't really want them, and find you have too much coupling and not enough Demeter
(Of course, you might be writing FooManager classes with this upper layer explicitly in mind - in which case fair enough and you'll probably like DCI. But I think there's still a lot of mileage in plain ordinary functions that don't have to belong to anything)
Maybe a reasonable conclusion is that there is no silver bullet for managing complexity and making things easier to reason about.
They already have some results: http://www.vpri.org/pdf/tr2011004_steps11.pdf
To sum it up, they do "personal computing" in about 4 orders of magnitude less code than current mainstream systems (from more than 200 millions lines to about 20,000). That includes the self-implementing compilers. The only thing they left out was the device drivers (which by the way takes less than 0.5% the size of current systems).
Now the Complexity Werewolf just went from "Unspeakably Awful Indescribable Horror" to "Cute puppy".
> 1. a group of data attributes, and
> 2. a defined set of operations that can happen on instances of those data attributes.
It must be nice to live in a world where a given piece of data belongs in only one group.
vals = []
avgs = []
stds = []
file = csv.reader(open('data.csv', 'r'))
for row in file:
vals.append(row[0])
avgs.append(row[1])
stds.append(row[2])
Providing this abstraction without using objects would be difficult, I think. The key differentiator is, I think: is what you want to represent with a class a concrete thing with a clear interface? If those two things don't hold, then it may be better not to represent the concept with a class.Again, I have to break this reasoning all the time in C++. I try to use free functions as much as possible, but when I want behavior based on runtime values, I often have to represent abstract concepts using classes because those are the tools that C++ gives me.
def readcsv(file):
for line in file:
if line:
yield tuple(field.strip() for field in line.split(','))
No explicit class definitions necessary! And you can write the same kind of code. Or, if you want to be a bit too clever for your own good, you can shorten it to this: rows = readcsv(open('data.csv', 'r'))
vals, avgs, stds = zip(*[map(int, row) for row in rows])
At no point here have I used the "class" keyword!In fact, I think that's how generators are implemented in python, the language just hides this fact away from the programmer.
def read_csv(file):
for line in file:
yield parse_csv_line(line)
-or- def read_csv(file):
return (parse_csv_line(line) for line in file)
-or- def read_csv(file):
return map(parse_csv_line, file)
-or- from functools import partial
read_csv = partial(map,parse_csv_line)
None of which involve defining a class.In Python 3 it is (the `map` builtin is lazy, it's essentially `itertools.imap` and the old eager `map` has disappeared)
(let ((reader (csv-reader "data.csv")))
(loop :for row = (reader) :while row
:collect row[0] :into vals
:collect row[1] :into avgs
:collect row[2] :into stds
:finally (list vals avgs stds)))
That's in the bastardized Lisp that I favour, but any language with closures should allow it. The point is, no classes.I don't want to seem fanatical about this - your example is a fine use of objects. But unless I'm missing something, one can do just as fine without them, and with a little less code (since you don't have to define a class).
I (maybe!) have no problem with classes for things that, as you say, are concrete and have clear interfaces. But then, when things are concrete and have clear interfaces, there are a lot of nice ways to do things. Making things well-defined and clear in the first place is the hard part.
I almost invariably find I need more than one closure down the road, outside of simple predicates handed off to algorithms. Even for such simple things as specifying a sink for output (out :: string -> ()), later on I find myself wanting a flush() routine. That's just my experience: problems grow thornier over time, and starting out with something more object-like adapts to that growth more gracefully than a closure does.
I think a lot of the problem comes from the syntactic weight of classes. I've implemented closures in a commercial compiler; the actual implementation, when it uses captured state, is very similar to a class, and in fact the implementation method I chose behind the scenes was a class. So I tend to view closures (rather than simple function literals that don't capture state) as just a concise way of declaring and instantiating an object.
I've had some success with blending the two. For example, a read-only collection facade is often implemented in .NET with ReadOnlyCollection<T>. But the normal way of implementing that thing is to pass off an IList<T>, which you may need to implement wholly if the logical read-only collection is not already a list. I think that's a complete waste of time; instead, I wrote a function that takes two closure arguments, getCount :: () -> int and getItem :: int -> T, and henceforth I can create a new collection with new behaviour without worrying about the "name" of the descendant, but with the benefits of a class for flexibility down the road.
Those are just names for values though, we don't have to create a name for a new general principle just to use one.
Not to mention values often do not need names at all, in an expression oriented style where function calls are composed.
Variables are easier to name than the more general notion of class or type (being specific to what is immediately at hand and thus not a distraction), which I think was the original point.
And even assuming that my structures will be named is assuming too much. Consider common uses of the fold function, e.g. which act on the pair of (accumulated-value, item). It's common that the pair isn't named. It's parts are named usually, but that's beside the point, I wasn't forced to name the structure itself, much less create a "AccumulatedSumWithNextItem" class.
For example, List<Customer> describes both the type and its semantic content. List<Tuple<string,DateTime,int>> just describes a type; it's a lot less self-documenting, and code using it will be cryptic. In a dynamic language it will probably be customer_list or somesuch. Point being, "customer" is a name and it crops up somewhere, whether it's in the variable or the type.
I don't know the ins and outs of the loop macro, only that it has a lot of them, so perhaps this is simply Common Lisp and I need to keep reading my CL books, but the "row[0]" sugar makes me think that there is something else going on.
You may be able to get something close to what I wrote with a function that returns a closure, but I don't think you can get exactly the for x in y syntax. As explained in another post on HN's front page, "A Guide to Python's Magic Methods": http://www.rafekettler.com/magicmethods.html#sequence, you need to define the __iter__(self) method on an object to support that style of iteration. (And that method needs to return an iterator object.)
And that brings us back to the same point I already made, but I think it's so important that I'll say it again: the abstractions you choose to represent concepts in your program will be determined by what is easy in the language you're using.
def fibonacci_sequence():
a, b, c = 0, 0, 1
while True:
a = b; b = c; c = a + b
yield b
>>> fibonacci_sequence()
<generator object fibonacci_sequence at 0x108055550>
>>> list(itertools.islice(fibonacci_sequence(), 0, 10))
[1, 1, 2, 3, 5, 8, 13, 21, 34, 55]
>>>
for n in fibonacci_sequence():
print n
if n > 13:
break
1
1
2
3
5
8
13
21
>>>most of the time, OO is a waste of time, and a few arguments is all you need
That reminds me of another key point. I try not to let function signatures have more than a few arguments, and I try to keep those arguments primitive. (A good litmus test is how easy it is to call from the REPL.) When my code starts to break these guidelines, that's a sign of design weakness, and I do what it takes to break up the complexity. It still amazes me how far you can get in terms of simple, decoupled design just by doing this.
The widespread OO practice of factoring some of those arguments into a new class, so that now you need pass only one thing (the new composite object) instead of several old ones, does nothing to solve the problem, but instead makes it worse: you've both added complexity and lost transparency. It's like a kid saying yes when mother asks "did you clean your room", having shoved all the mess under the bed.
The same can be said about functional programming. But once you degenerate to passing a ton of parameters around in each function call, you're probably better off defining an explicit object anyway.
Measuring complexity by number of lines and saying defining a class is adding complexity is off the mark.
Coupling data and operations is also a problem. It's non-extensible, both in that you can't add more data and use the same operations, and you can't add more operations that work on the added data. Inheritance works around some of this, but it doesn't solve the general problem. What if it makes sense to add the same piece of data to objects of classes from different parts of the hierarchy?
The thing about functions is that you can always write a new one that works with anything you want. You don't need to add it to a class or an interface (you can, if the language allows that), and if possible, with immutable data that essentially makes the function a unit of perfect encapsulation.
Not content with repeating myself in this thread, I will repeat myself from threads of the past as well: it's astonishing that the evidence is so consistent on this, yet almost nobody takes it seriously. But I do! And I have a really interesting question, too: what it would it mean, not just for a programmer to organize their work around this principle, but for an entire company to do so?
We all know that lines of code is a poor measure of code complexity, but it does correlate with it. It is far more reasonable that error rate is proportional to code complexity, which is weakly measured by LOC.
Unfortunately I don't have any studies to back up my intuition here. But I would bet that my definition of complexity, aka "interaction space", would correlate much stronger to error rate than LOC does. Also, that decently good usage of classes would in fact reduce the complexity, and thus error rate.
Intuition is often surprisingly wrong - lesswrong.com shows a lot of examples. That doesn't mean it's wrong in this case, but now you have code which is hidden so the interaction space is limited, so there could be:
1) Now you have to add complexity to get information into and out of this interaction space. Similar to how it's easier to walk onto a field than it is to walk into a castle - if the walls protect the contents, they also contort your path around them.
2) Leaky abstractions. It might be the case that your protected code is needed elsewhere. Now you have to duplicate it, or break around the protection/expose a new way through, all of which would not be necessary for a free function. At the very least you have overhead considering this. All of a sudden your neatly wrapped class is two classes, one in use in two places where you could again change code and have an effect far away if you're not careful.
3) In a language such as Python, you don't have your code inside a class protected by anything except convention. So it seems that you either must: a) accept that what protects groups of code is programmer attentive care, not the system of classes itself, and thus that any other programmer paying the right kind of attention could have the same benefits and less code, or b) declare languages like Python to be excluded from the benefits of classes by not implementing them properly. Do you agree?
We all know that lines of code is a poor measure of code complexity
We certainly don't! Program length is the only good measure we have. The best measure of program length is up for grabs, but LOC is probably as good as any (there was a big thread about this a few months ago).
But I would bet that my definition of complexity, aka "interaction space", would correlate much stronger to error rate than LOC does
But the best work on this points to so exactly the opposite of what you say, I would really recommend you take a look at it: http://www.neverworkintheory.org/?p=58 - and tell us what you think.
I'm going away now.
If we can come up with more metrics that are based on our perception of a codebase's complexity, I think we would see interesting results.
Any language that isn't a fairly close C descendant. Haskell, Lisp, Python, Perl are languages I've done some work in with minimal OOness.
When I used Java a few years ago, it required classes. C++ winds up being very class-ish. C of course is limited.
I have come to be very pleased with not using classes unless the solution really demands it. It's sort of a "grow the solution" idea, instead of "waterfall the solution".
At least in Python - which is the only one of those four I can speak for - all you are doing is basically OO (in fact, everything in Python is an object). The beautiful thing about Python is how it's abstracting a lot of this away by providing functions which implicitly call "magic methods" which are either generated by the interpreter or can be defined by the programmer (ie, __iter__(), __next__(), __init__()), and by virtue of Duck Typing, which means an object defines itself by what it can do, not its identity or origin.
You'd never even notice a lot of this unless you dip into the internals of Python a bit. And ultimately, I don't think this is actually possible without using OO or some close equivalent. So in the end, a lot of the "solutions" people have proposed in this thread ultimately depend on the very thing they want to "solve": Object Orientation.
Or of "sockaddr_in".
sockaddr_in is the way it is because of sockaddr, sockaddr_in6 etc. Judicious polymorphism could have led to fewer headaches converting IPv4 apps to IPv6.
Couldn't agree more. I've long suspected that OO is a diversion or misdirection from actually solving the problem, rather than an inherent part of it. Discovering functional programming recently with Scheme and then Haskell has only reinforced that.
OO may be the appropriate abstraction for some things like GUI programming, which seems to be what it was originally created for, but not everything.
That doesn't mean everything needs to be a class -- that's crazy too. And except for special cases, my inheritance hierarchies never have more than two levels. (And I probably wouldn't allow Point3D, for instance, to be derived from.)
But classes are an incredibly powerful tool when they are needed.
Is it the interface it provides, the types, the coupling of methods to data, or something else?
1) Naming it vastly clarifies further data structures and interfaces. (This is shared with having a Point3D data-structure, obviously.)
2) Providing functions and operators specific to the type.
At it's simplest (and possibly most compelling) it's the difference between needing to say
p.x = p0.x + t * (p1.x - p0.x);
p.y = p0.y + t * (p1.y - p0.y);
p.z = p0.z + t * (p1.z - p0.z);
and p = p0 + t * (p1 - p0);
The latter is easier to type, easier to understand, and less error prone. It's a huge win, especially in more complicated expressions.Now of course, your language may allow you to define data structure-specific functions and operators. But at that point, what you've got is essentially a class, even if you call it something else.
(PS Note that data hiding is completely unimportant here, and inheritance would be nothing but trouble.)
That aside, you're quite right that when one has a common construct that permeates the entire program, one badly wants to give it first-class status in the language. What I've noticed, though, is that these constructs don't tend to get involved in complex object graphs (or if they do, they're leaf nodes). Rather, they tend to be atomic and universal (i.e. interoperable with anything) and simple. So I wonder if what one really wants here is the ability to add new primitives to the language. That's potentially quite different from classical classes. It's more related to the "build up a language that fits the problem, then code the solution in that language" idea.
So clearly I agree that adding primitives is an extremely important ability. But it's already nearly if not completely there in all the OO languages I use, using classes.
I don't think primitives are enough, though. The interface / implementations of interfaces combination is also essential to my programming. Being able to say "Curve" and have it mean Straight or NURBS curve or Circle or Offset curve or Composite curve or... is another huge simplifying device. In fact, I've just looked at my code, and I have 15 distinct curve types, at least 5 of which I'd have missed if you asked me to list them, despite having written all of them over the years. That is exactly the power of simplifying abstraction.
Let me point out that given a sufficiently capable OO language, this doesn't need inheritance to implement. This is just the implementation of a role or interface. But it does need classes.
And surely you're not arguing the defining feature of a class is that it's closed?
And in cases where you really do need class-like behaviour, it's easy to extend it with additional methods or operators, while still getting all the other convenience stuff for free:
class Point3D(namedtuple('Point3D', "x y z")):
def __add__(self, other):
if other.__class__ != self.__class__:
raise TypeError
return self.__class__(*[a+b for a,b in zip(self, other)])
print Point3D(1, 2, 3) + Point3D(5, 4, 2)Much of the benefit of C++11 and Boost is in not having to write functor classes.
I played with boost::lambda, and I found the resulting code less clear than a simple for loop. I also had difficulty gaining an intuition for what I could and could not do with it. Rather than constantly asking myself "Can I use boost::lambda here...? Will it look better or worse than the alternative?" I just decided to never use it.
While I prefer a functional style of coding, I recognized that it was silly to adhere to it if doing so made my code worse - C++'s support for what I wanted to do just wasn't there. If what I wanted to accomplish was several lines of code, I wrote a separate function or function object. If it was brief and I could not express it with combinations of existing functions, then I just wrote a for loop.
Twice now, I've been faced with a metric crap-ton of Java source files for reading / writing to a device.
Anything you could possibly imagine, gets defined as a class. There's dozens of classes, and more SLOC than I can count.. Meanwhile, an equivalent C program comes in at 2k SLOC, and 4 files.
You().Can('not').Do(any(thing)).Concise().In(java())!
When I was primarily doing C++, I gradually switched to generic programming once visual studio supported STL... And then when I started using Perl I found I didn't need to write classes at all.
Reminds me of this Paul Graham quote: "This practice is not only common, but institutionalized. For example, in the OO world you hear a good deal about "patterns". I wonder if these patterns are not sometimes evidence of case (c), the human compiler, at work. When I see patterns in my programs, I consider it a sign of trouble. The shape of a program should reflect only the problem it needs to solve. Any other regularity in the code is a sign, to me at least, that I'm using abstractions that aren't powerful enough-- often that I'm generating by hand the expansions of some macro that I need to write."
This is just so wrong. The shape of your program should reflect your understanding of the problem, and the clearest way to communicate that to other humans. Patterns will be a part of this, as it reduces the amount of mental effort required to understand the solution.
It is true that writing too much boilerplate to create a pattern is a sign that your language is lacking necessary abstractions. But patterns themselves will always be a part of good software.
I swear this entire thread is just trying to rationalize bad programming practices.
In the neighboring programming-language kingdoms, taking out the trash is a straightforward affair, very similar to the way we described it in English up above. As is the case in Java, data objects are nouns, and functions are verbs. But unlike in Javaland, citizens of other kingdoms may mix and match nouns and verbs however they please, in whatever way makes sense for conducting their business.
For instance, in the neighboring realms of C-land, JavaScript-land, Perl-land and Ruby-land, someone might model taking out the garbage as a series of actions — that is to say, verbs, or functions. Then if they apply the actions to the appropriate objects, in the appropriate order (get the trash, carry it outside, dump it in the can, etc.), the garbage-disposal task will complete successfully, with no superfluous escorts or chaperones required for any of the steps.
There's rarely any need in these kingdoms to create wrapper nouns to swaddle the verbs. They don't have GarbageDisposalStrategy nouns, nor GarbageDisposalDestinationLocator nouns for finding your way to the garage, nor PostGarbageActionCallback nouns for putting you back on your couch. They just write the verbs to operate on the nouns lying around, and then have a master verb, take_out_garbage(), that springs the subtasks to action in just the right order.
Dictionaries and lists of things don't relieve you from the job of naming and encapsulating the data you're dealing with. But full on classes are usually only needed when behaviour needs dispatching from polymorphic locations, and that isn't too often.
It's the OO and design pattern culture that results in much of the Java code you describe as over-engineered. But you can write bad code in any language.
Taking software engineering advice from algorithm geeks is a mistake we've been making for decades over and over again.
In my professional life, I've not had a need to engineer a single algorithm. I have had to make plenty fault-tolerant, maintainable and understandable control software, though.
The same holds for most, if not all, of my colleagues. Common professional software engineering has nothing to do with sitting in a room on your own figuring out how to best store a dictionary in memory such that it is fast under insertions, removals and lookups.
Or is it that classes in Python specifically aren't that much of a win over the language's implementation of other paradigms? I sure haven't hit into any problems using classes or objects in Ruby, even when sometimes the use feels a little contrived to begin with, but.. I also have no choice :-) (maybe Rubyists have a Stockholm Syndrome with OO ;-))
In Python, it's reasonable to have a package with just functions in it, whereas in Ruby, writing a top-level method means polluting every object in the system. You can write modules that have singleton methods on them, but you still don't have anything as flexible as Python's "from pkg import foo, bar"—the caller needs to either write "ModuleName.my_function" at every call site, or use "include" and end up with the module's methods as part of the consuming class's interface.
After working with Perl (which has a similar namespacing), I must it's something I really miss when I come back to Ruby.
Python - take your __init__+1 method class and just make it a function
Ruby - take your initalize+1 method class and just make it a static method (and maybe make that class a module instead)
It's more that Python programmers prefer not to proliferate excess complexity or layering. They hate ravioli code.
This isn't to say his underlying message is wrong (I 100% agree with Jack's message, too much classes create unnecessary complexity in your library/app), but I don't want people destroying their classes without fully understanding the consequences.
try:
import simplejson as json
except ImportError:
import json
I'm pretty sure I've seen those exact four lines of code in the internals of Tornado and Flask, among others.He should explicitly catch the ImportError
In plain C, your main "method" can be a simple external function, and the "helpers" can be static functions. Python offers similar functionality.
Making a class is a step backward because the internal helpers have to exposed to users, even if they're intended not to be used.
(The reality is simple. Classes are Python's state-capture abstraction. Using a bunch of lambdas to do the same thing means you are writing Haskell in Python, which is a dumb thing to do, because while the computer knows what you want, other programmers don't. If you want to use Haskell, use Haskell.)
My experience confirms this.
EDIT: Can't reply to Tommabeeng's comment (probably too deep). Yup, and I agree with the presenter's main point - that classes are not necessary in many situations, even for decoupling.
That's wrong. I like the key to his talk but the plague of large complicated systems is exactly coupling and incoherence. Classes are NOT necessarily the way to achieve these.
I think his intention was to explain that these are the same words (mostly from academia and taken up with reckless abandon via language architectures like Java) that may be thrown around by some to create the code mess (anti-pythonic) in the first place.
It's easy for people who are very smart and experienced to dismiss the rules - that applies for all things, not just programming. But if you never learn the basics then you'll stumble over silly things that have be solved by much smarter people.
var o = {
name: 'David',
greet: function () {
console.log('Hi, I am ' + this.name);
}
};What can you possibly believe a class is other than exactly that?
Very often when writing Python (or whatever) you'll create a class which is only ever instantiated once and the instantiation is done simply for the benefit of calling one method.
In such cases I find classes to be overkill. What I really want is the object. I don't need to factor the object's behavior out into a class.
Many of my sympathies are expressed much more thoroughly and eloquently in the many papers you can find online by Walter Smith outlining the motivations behind NewtonScript.
function greet(name) {
console.log('Hi, I am ' + name);
}
var o = function() { greet('David'); };I suppose that "true" objects are Erlang processes. (I'm only partly joking)
OOP gone wrong when it became structs with methods.
Edit: here you go:
http://news.ycombinator.com/item?id=2855500
http://news.ycombinator.com/item?id=2928672
http://news.ycombinator.com/item?id=2856567
http://news.ycombinator.com/item?id=2855508
The first two are teaser comments. The second two are the important pieces. Also the paper Kay praises at the start of the talk is short and very worth reading.
Unfortunately, this is the only way to perform this activity in Java. Also, I adore Go's style of pseudo-OO programming.
.. or, from reading the HN threads, is that decoupling?
If so, I'm using decoupling right now!
Alternatively, if you are working with things that you expect to follow a certain interface, Python does not really require a formal definition. Duck typing should suffice.
I agree on some of the presenters philosophical ideas, mainly that adding complexities into the codebase before they're needed is almost always a terrible idea, but saying that classes lead to complex code is simply not true, at least in my experience.
I think the distinction between 'end-users' and people who make libraries is an important one and one that Jack seemed to overlook.
That's a platitude which everybody can and does say about everything.
But I think one can make a case that there is something wrong with OO in general. The original arguments for OO were: it's better for managing complexity, and it creates programs that are easier to change. Both of these turned out not to be true in general. So we need to ask, when is OO a win? But for simpler problems it doesn't matter, because anything would work. So we need to ask: what are the hard problems that OO makes significantly easier? I don't think anyone has answered that.
I suspect it's that OO is good when the problem involves some well-defined system that exists objectively outside the program - for example, a physical system. One can look at that outside thing and ask, "what are its components?" and represent each with a class. The objective reality takes care of the hardest part of OO, which is knowing what the classes should be. (Whereas most of the time that just takes our first hard problem - what should the system do? - and makes it even harder.) As you make mistakes and have to change your classes, you can interrogate the outside system to find out what the mistakes are, and the new classes are likely to be refinements of the old ones rather than wholly incompatible.
This answer boils down to saying that OO's sweet spot is right where it originated: simulation. But that's something of a niche, not the general-purpose complexity-tackling paradigm it was sold as. (There's an interview on Youtube of Steve Jobs in 1995 or so saying that OO means programmers can build software out of pre-existing components and that this makes for at least one order of magnitude more productivity - that "at least" being a marvelous Jobsian touch.)
The reason OO endures as a general-purpose paradigm is inertia. Several generations of programmers have been taught it as the way to program -- which it is not, except that thinking makes it so. How did it get to become so standard? Story of the software industry to date: software is hard, so we come up with a theory of how we would like it to work, do that, and filter out the conflicting evidence.
Wrong. Even the worst Enterprisey mess of Java classes and interfaces that you can find today, is probably better than most of the spaghetti, global state ridden, wild west that existed in the golden days of "procedural" programming.
If you consider that software is composed of Code and Data, then OOP was the first programming model that offered a solid, practical and efficient approach to the organization of data, code and the relationship between the two. That resulted in programs that, given their size and amount of features, were generally easier to understand and change.
That doesn't mean OOP was perfect, or that it couldn't be misused; it was never a silver bullet. With the last generation of software developers trained from the ground up with at least some idea that code and data and need to be organized and structured properly, it's time to leave many of the practices and patterns of "pure" OOP and evolve into something better. In particular, Functional has finally become practical in the mainstream, with most languages offering efficient methods for developing with functional patterns.
With the global state code you have to understand the entire codebase to be sure who exactly is modifying a particular bit of state. This is far more fragile because the "interaction space" of a bit of code is much greater. The dangerous part is that while you must understand the whole codebase, the code itself doesn't enforce this. You're free to cowboy-edit a particular function and feel smug that this was much easier than the enterprisey code. But you can't be sure you didn't introduce a subtle bug in doing so.
The enterprisey code is better because it forces you to understand exactly what you need to before you can modify the program. Plus the layered abstractions provide strong type safety to help against introducing bugs. Enterprisey code has its flaws, but I think its a flaw of organization of the code rather than the abstractions themselves. It should be clear how to get to the code that is actually performing the action. The language itself or the IDE should provide a mechanism for this. Barring that you need strong conventions in your project to make implementations clear from the naming scheme.
I agree that a project needs strong conventions, consistently practiced. See Eric Evans on ubiquitous language for the logical extension of that thinking. But this is as true of OO as anywhere else.
>interactions between objects can be just as hard to understand as any other interactions.
While this is true, there are strictly fewer possible interactions compared to the same functionality written only using primitives. To put it simply, one must understand all code that has a bit of data in its scope to fully understand how that bit of data changes. The smaller the scopes your program operates in, the smaller the "interaction space", and the easier it is to reason about. Class hierarchies do add complexity, but its usually well contained. It adds to the startup cost of comprehending a codebase which is why people tend to hate it.
In my experience, classes don't give the kind of scope protection you're talking about. They pretend to, but then you end up having to understand them anyway. True scope protection exists as (1) local variables inside functions, and (2) API calls between truly independent systems. (If the systems aren't truly independent, but just pretending to be, then you need to understand them anyway.)
In python you can access any classes internals if you're persistent enough. However, the difficulty in doing so and the convention that says that's bad practice gives the perception of a separation, thus one does not need to be concerned about its implementation. It lowers the cognitive burden in using a piece of functionality.
Haven't you noticed this yourself? A simpler interface is much easier to use than a more complicated one. If you're familiar with python, perhaps you've used the library called requests, its a pythonic API for http reqeusts. Compare this to the core library API using urllib2. The mental load to perform the same action is about an order of magnitude smaller with requests than urllib2, and its because there are far more moving parts with the latter.
My contention is that if bits of data and code are exposed to you, it is essentially a part of the API, even if you never actually use it. You still must comprehend it on some level to ensure you're using data it can access correctly, aka interaction space (I feel like I'm selling a book here).
Part of the allure of Java: it gives the illusion that more progress was made before you entered that region.
We've all heard that you can shoot yourself in the foot with any language, but I think the implications haven't fully sunk in: There are no bad paradigms, only bad codebases.
How so? I find OOP code to be much easier to understand and change.
Considerable size of the system is because of considerable complexity of the task at hand. I don't see how you can write 1000-page book of requirements into 2 screens of code. And if you need the system that is able to be customized to satisfy different books with minimal change - you need more complex code.
If people aren't moving on, it's not simply because of 'intertia'. It's because the alternatives aren't offering enough benefit yet.
What will change the paradigm is the desire of a future generation to make its mark by throwing out how the previous generation did things. One can see this trying to happen with FP, though I doubt FP will reach mass appeal.
When you do anything as intensely as writing software demands, your identity gets involved. When identity is involved, emotions are strong. The image of being a rational engineer is mostly a veneer on top of this process.
Edit: 'intertia' - your making fun of an obvious typo rather illustrates the point.
So this is basically a contentless argument that says people are stuck on OO because they are irrational.
As you me 'making fun' of a typo - I don't know what you're projecting onto me. I simply mistyped the word inertia. I put it in quotes because I don't think it's a cause of anything.
As for OO, I think our disagreement has probably reached a fixed point.
Edit: nah, I can't resist one more. I believe I explained how the inertia was overcome before: people who were not identified with the dominant theory at the time (structured programming) came up with a new one (object-orientation) and convinced themselves and others that it would solve their problems. Why did they do that? Because the old theory sucked and they wanted to do better. Their error was in identifying with the new theory instead of failing to see that it also sucks. The only reason they could see that the old theory sucked was that it was someone else's theory.
As for "what would replace OO if not for inertia", I know what my replacement is: nothing. I try to just let the problem tell me what to do and change what doesn't work. Turns out you can do that quite easily in an old-fashioned style. But if you mean what paradigm will replace OO (keeping in mind that we ought to give up our addiction to paradigms), who knows? FP is a candidate. One thing we can say for sure is that something new and shiny-looking will come along, until we eventually figure out that the software problem doesn't exist on that level and that all of these paradigms are more or less cults.
Perhaps I should add that I don't claim to know all these things. I'm just extrapolating from experience and observation of others.
The trouble I have with what you're saying is it suggests that a better paradigm (e.g. FP), higher-level and more powerful, will improve upon and succeed OO. But the greatest breakthroughs in my own programming practice have come from not thinking paradigmatically- to be precise, from seeing what I was assuming and then not assuming it.
Edit: My own experience has been this weird thing of moving back down the power continuum into an old-fashioned imperative style of programming, but still very much in Lisp-land. For me this has been a huge breakthrough. Yet my code isn't FP and it certainly isn't OO, so I guess it must be procedural. How much of this is dependent on the language flexibility that comes with Lisp? vs. just that Lisp happens to be what I like? Hard to say, but I suspect it's not irrelevant. If you can craft the language to fit your problem, you can throw out an awful lot. Like, it's shocking how much you can throw out.
There were also some revelations - like printer().getSingleton() (or instead of printer, put keyboard, mouse, memory, etc.) - at some point some of these "singletons" might be become "multi-tons" - for example gamepad().getSingleton() already is a trouble...
"The singleton pattern combines all the perf benefits of a global variable with all of the code maintenance benefits of a global variable."
1) Do not use a class if it doesnt have state
Agreed
2) Keep it simple and do not think about the future, fix it later if needed
He probably never wrote API code in libraries needed for other people. Once you choose something you are bound to it. In this case classes and interfaces are very helpfull. For example the API class example in the beginning (no state) is a good example. In the future the api might be extended with more functionality. The last thing you want is have 3rd parties rewrite their code cause you changed the API. So a class is good to provide a Facade (yes Facade Pattern).
3) package names are not for taxonomy but for preventing nameclashes
True. However in Java the editor is so powerfull that you never ever have to lookup package names. This goes automatically (IDE parses the code on the fly). Java package names also use URI patterns like: com.google.android.hardware.vibrator. Which can only be used by somebody who has the google.com domain name. This prevents name clashes ALL the time. In python you sometimes have to rename a whole project. And the taxonomy gives you an indication what the code actually does.
4) Dont overuse exceptions and keep shorts names.
Agreed. However i am not sure what the best granularity is. In a public Library API, i would be very carefull about this.
Otherwise I agree with him.
Classes are fine. Over-engineering classes, like over-engineering anything is bad, m'kay?
The benefits of Classes are still related to data abstraction. But it is pointless if you are still accessing the data directly anyway (getters and setters be damned). So, I would only legitimately prescribe a class only when you are building a data structure.
Nowadays, my classes are either - composition of primitive data (as I don't have tuples(python)/structs or dynamic objects (javascript)) - or function groups (like a business logic layer of some sort, which could have just as easily have been just bags of functions in modules).
I'm going to be teaching my students Classes in a couple of weeks... food for thought before then.
Facepalm.
eg complaint about nested packages. (not that I agree that this is a necessary class): instead of MuffinMail.MuffinHash.MuffinHash, type MuffinHash and have your editor add: from MuffinMail.MuffinHash import MuffinHash.
No
Ooooh, you wanted to make a case for not using classes in certain situations? Then enough with the linkbait title.
class foo:
pass
f = foo()
f.bar = "yes"
which is easier to type and more readable than: f = {}
f['bar'] = "yes"
(credit where due: learned this trick from TB) >>> x = type('Foo', (object,), {})
>>> x.bar = 1
>>> x.bar
1