In a FPL your graph will be an ADT. The advantage is clients can pick it apart, but there is no easy way to extend it. If you wanted to "use the same object for computing distances to multiple vertices", you're screwed: the ADT is laid bare, and there is no data hiding.
In an OO language, your graph will be an opaque object, and you can write things like CachingDijkstra or ParallelDijkstra and make these work transparently for your clients. The price is that you are bound to your APIs: a client cannot pick apart your data structure, because it's opaque.
It's not like when I use a sorting function that the implementors couldn't swap it out for a caching one or a parallel one.
One of the amazing things I remember is changing an int to a float in a very well used "struct" that was touching thousands of functions but because the functions in Haskell usually so extremely general I didn't have to change a single other line (Number type is heavily used and doesn't incur an extra performance cost, which it very much does in Java).
In fact we can be even more precise or loose in what assumptions need to be made via dependent types in functional programming languages like Agda.
I encourage you to take a look at any popular Haskell-framework to see how it's done in practice.
Functional programming languages tend to have good support for algebraic types (which I think you mean) but they can often do data-hiding too! In Haskell, for example, you do this by not exporting the data constructors for your type, so they can only be used within the module that defines the type.
DijkstraLib =
there exists GraphState with {
construct: List(Edge) -> GraphState
getDist: GraphState -> Vertex -> Vertex -> Num
}
Now we can write two different implementations of this type that use a different data structure for graph. The data structure will be opaque to the client (even though the client obtains and passes around graph objects!). Like so: -- Library
type CachingDijkstraLib : DijkstraLib
with GraphState = (HashMap((Vertex, Vertex), Num), List(Edge))
with {
construct(a_list) = (HashMap.empty, a_list)
...
}
-- Client
DL = CachingDijkstraLib : DijkstraLib
some_obj = DL.construct(edge_list)
DL.getDist(some_obj, u, v)
DL2 = SomeOtherDijkstraLib : DijkstraLib
some_other_obj = DL2.construct(edge_list)
DL2.getDist(some_other_obj, u, v)
The types of some_obj and some_other_obj could be completely different, and neither would be accessible to the client. For example, the client couldn't assume some_obj is a pair and try to get its first element. It would also be an error to call e.g. DL2.getDist(some_obj, ...).It's an age-old discussion. As old school software engineering became sufficiently advanced, some people bundled their favorite "state of the art" concepts, and called it OOP. (Also: Multiple people, multiple different concepts of what "OOP" actually precisely means.) It became a horrible hype, and from then on out everyone was taught to see all these concepts as "OOP concepts" that elevate OOP over other programming paradigms.
But they're not and they don't.
The only real "advantage" is, OOP languages are opinionated and will try to force their interpretation of these concepts on the coder. Which is nice, if you happen to like it that way. Just not everybody does.
For instance, in Python, one would define a function with `def fun(state)` or a class method with `def fun(self)`. In the function body, one would use respectively `state.var` or `self.var` to refer to a variable `var` in the state object. Running it is then done using `fun(state)` or `state.fun()`.
There are some minor syntactic differences between the two, but I find both these examples to be just as explicit...
Haskell developers don't just create a typeclass to bunch related functionality together. Doing so is a common beginner's anti-pattern. Typeclasses exist to wrap up general operations that satisfy some laws.
When objects in OOP are used in the same way, I don't mind it as much. It's when they're not that I don't find the syntactic inconvenience worth it.
In OO notation, we have A.B(C)
In almost all languages, you can store C in a variable and supply it "later":
z = C
A.B(z)
You can also store A in a variable and supply it "later": x = A
x.B(C)
However, it is much more rare to be able to store B in a variable and supply it "later". Often, you must store A.B: xy = A.B
xy(C)
This means that at the time you select which method to call, you must also have at your disposal the instance on which you want to call it!With functions, things get easier. If you have B(A, C), you can generally just store B in a variable on its own:
y = B
y(A, C)
Some OO languages allow you to store plain methods that aren't yet bound to a particular instance (JavaScript, modern Java) but far from all do, without resorting to inconveniences like y = lambda o: o.B
Please note that what this workaround accomplishes is effectively to turn the method into a function!I suspect the main argument here is that it just doesn't make sense notationally to treat the first argument to a function specially. It only introduces wierd edge cases that have to be worked around. Plain old function notation has stood the test of time for good reason: it's flexible and handles most of the common cases we want.
That is also its advantage, since the method to call can depend on the instance. Not just for dynamic dispatch either. A statically dispatched call will look up B on classof(A).
class Animal:
alive = true
class Cat : Animal
def hello(): print "meow"
class Dog : Animal
barked = 0
def hello(d::Dog):
print "bark!"
barked++
let d = Dog {}
d.hello()
let a:Animal = d
a.hello()
In other words, A.B(C) is just syntactic sugar for B(A, C), and class definitions are just syntactic on top of that. The relation between OO and regular functions and stateful objects is completely transparent (the only real gotcha I can see is that a.hello() still works here because it was assigned a subclass where this method is defined). This makes it easy to store B for later usage, like you wanted to do in your example, and reason about its behavior.[0] https://en.wikipedia.org/wiki/Uniform_Function_Call_Syntax
Which ones?
>However, it is much more rare to be able to store B in a variable and supply it "later". Often, you must store A.B:
How is it rare? Smalltalk doesn't require you to specify the receiver when storing a symbol. Neither does Objective-C (selector). Neither does Ruby. Nor Java.
Which commonly used languages require one to have the receiver tied to the method when memoized? C++ is the only one I'm aware of. So it might be rare to do it in C++, but not in OOP generally.
Not for virtual member functions (not that I've been able to find). C++ is geared towards generating optimal executables. Since the type of the object is known at compile-time, the particular function used for that type is known, and the compiler will generate code that applies for it.
This is why C++ uses mangled names: to discriminate between identically named functions of different types and signatures.
The time you "select which method to call" is when you write the code, right? You select it once. Like, "printf(arg);". I selected printf (in my mind), I called it. I don't need to do "x = print; x(arg)".
> With functions, things get easier. If you have B(A, C), you can generally just store B in a variable on its own:
> y = B
> y(A, C)
>Some OO languages allow you to store plain methods that aren't yet bound to a particular instance (JavaScript, modern Java) but far from all do...
Sorry, I lost the plot here. If you have B(A, C) then you don't need "y = B". B is an invariant of sorts. Just like you don't usually store the address of a class method; you just call that method on an object. Nobody writes:
> x = &objectInstance::method;
> x(args);
People just write
> "objectInstance.method(args);"
Therefore in your case, if using functions, you just do B(A, C). No assignment needed, works in every language.
What am I missing?
myobj.cool_func(1,2,3); // C++
cool_func(myobj, 1, 2, 3); // C
Is this better?If you use a vanilla C++ style class then the class definition is in the header. All changes to the class require a recompilation and break the ABI of the class. This means you basically cannot release patch versions of your library because consumers cannot link against it if they have the old headers.
Contrast with C where you would forward declare a struct in the header and pass that around while the definition is hidden in a .c file. This means changes to the size of the structure don't matter because all calling code ever sees is a pointer. (True data hiding!) It's called an opaque pointer and many many libraries use this pattern. e.g. libcairo.
There is then a pattern to do this in C++ called pimpl:
https://en.cppreference.com/w/cpp/language/pimpl
This solves it for C++ but the alleged convenience is gone because now you're coding C++ with the supposed convenience but still your class implementation has to trampoline all the calls to the pimpl which is an extra function call (if you're writing C or C++ this might matter).
But RAII, though. Not having [standardized] RAII sucks.
This is pretty much only a problem in C/C++ though. Everyone else is using static linking (which you can of course also do in C++), or a VM of some form. IMO relying on a library having stable ABI is a bit of anti-pattern. It's like relying on the implementation details of a function/class rather than it's API: sometimes it's necessary, but it should be avoided if possible.
I think you're heavily over stating reliance on libraries having stable ABIs as an anti-pattern. If a mistake is made in the ABI promise, can it get hairy? Absolutely. If you work on an application and are somewhere in the middle of it and need to recompile it to test a fix, do you want to wait for the compilation of the entire application to test it? Obviously not.
definitely not all, only things that change the layout. adding a non-virtual method does not break ABI at all for instance.
> Contrast with C where you would forward declare a struct in the header and pass that around while the definition is hidden in a .c file. This means changes to the size of the structure don't matter because all calling code ever sees is a pointer. (True data hiding!) It's called an opaque pointer and many many libraries use this pattern. e.g. libcairo.
Does this have a point in 2020 ? Good modern practice (c.f. Rust, etc) is to ship LTO'ed mostly-static binaries. Such hiding is only relevant for legacy things such as what Linux distros obstinates themselves in doing for their userspace - hopefully in a few years everyone will switch to snap / flatpack / appimage to finally reach a sane application distribution model on Linux.
I'd definitely say PIMPL is 100% an anti-pattern in every possible case. It forces more memory allocations, and makes it much harder to hack your way out of things when it's 3AM and everything is crashing and you have 5 minutes to fix things before heads start to roll.
Yes, you are correct. The point was that C code can keep the ABI intact through (some) data layout changes but vanilla C++ classes cannot I misstated it.
>Good modern practice (c.f. Rust, etc) is to ship LTO'ed mostly-static binaries.
That is the state of play now because Rust does not have a stable ABI. I thought one might be coming but maybe no one's working on it. It's unfortunate because it would be really nice to have binary crates with signing.
>hopefully in a few years everyone will switch to snap / flatpack / appimage to finally reach a sane application distribution model on Linux.
Snap has a while to go before it can be adopted fully. Images are enormous, they take a long time to load, and they don't respect (or don't understand) hidpi settings. I tried to use it for spotify and telegram-desktop and other applications but they just don't work that well.
>it much harder to hack your way out of things when it's 3AM and everything is crashing and you have 5 minutes to fix things before heads start to roll.
Absolutely but it comes from enormous systems like Windows where changing a library and breaking ABI means recompiling Everything. And these compilation jobs are often 'overnights'.
What Go and Rust do (by default) is only useful for internal software that you have full control of (both in source and in updates), but it is definitely not "good practice" for general software distribution.
Going fully static (Go) or mostly static (Rust) means your users/clients cannot update dependencies easily. Users will suffer when the vendor is not replying as quickly as they would hope for (which is not a pleasant experience at all, specially if you run server software) or when they simply discontinue the software (cf thousands of games). Then there is the usual duplicated code pages with less effective cache utilization etc. (which is a minor issue nowadays but if everything is a 100 Mo Electron copy it starts being painful).
Linux packaging is interesting, since the delivery model of gratis distributions often have no such guarantees. Thus one may indeed experience difficulties at any moment after a recent upgrade due to that model. In that case, the complexity lies within upstream and distro, a shared responsibility that the distro should reconcile but may fail to do in minute detail.
Modern methods involve a pipeline, and no rogue upgrades that haven't passed multiple stages of tests and security checks.
If they don’t work, they are broken. Of course it can happen, but the opposite approach means never being able to do it.
> If there are security issues, 99% of that should rather be solved by safe language usage and security perimeter or encrypted channels
There is no mainstream language or operating system out there that solves "99% of security issues" unless you are talking about formally proven systems etc.
> The deliverables are provided and guaranteed by the vendor.
As I explained in the GP, this only happens for systems with support contracts.
For most software out there, this isn’t the case. Mainstream software vendors don’t guarantee you anything at all, for good reasons.
> Modern methods involve a pipeline, and no rogue upgrades that haven't passed multiple stages of tests and security checks.
That is not "modern". That is how it has always been done since the 80’s. Again, for software properly supported.
I am not sure why you talk about "rogue" updates, since nobody has mentioned such.
My experience as a Linux user is that I suffer even more from some minor update in /usr/lib/libwhatever.so suddenly breaking some feature in software that I use to do my job / have fun, which wouldn't happen at all if that software was self-contained.
The benefit of dynamically linking is that you (as a user) or upstream (most likely) can fix those issues you point out, including future ones.
Being self-contained simply means updates aren’t forced into it, which is a good thing and I agree with it.
I didn’t downvote you, by the way, since the point you make is valid.
It sure sounds to me like modern practice as you've defined it is that when a component has a security vulnerability you're not going to patch it. It's going to be statically linked and running in a container image maybe created by somebody else and you'll have no idea it's there.
Yes, it theoretically lets you ship a patch version bump containing security fixes for the few languages with a stable ABI. But even ignoring the deluge of unpatched IOT garbage running seriously out of date kernels - nevermind userspaces - how do you ship that version bump of your library to your Windows users? Your OS X users? iOS? FreeBSD? RedHat? Android? Mint? Debian? PS4? Redox? Xbox One? Solaris? Switch? I've dynamically linked into PuTTY before - a good 'ole C codebase - and I still had to rebuild and redeploy myself when a security patch landed.
Meanwhile, an updated package on crates.io can change deps.rs badges, or trigger dependabot - optionally auto-merging the pull request if CI passes, or altering you to the build failure and need for manual action if that "patch" version broke things. Rustsec and cargo audit/geiger/crev are all hideously platform independent ways to catch problems. My patches to an unsound Rust crate can automatically land on your WASM-enabled website if you have things so configured.
Ctors/dtors can break ABI as well: https://gcc.godbolt.org/z/XAx00v
Very few professional developers code in vim or notepad. Vast majority of us are using IDEs. IDE knows types of things, type `myobj.` and you'll get a suggestion list with methods of that class callable from the current context.
> There is then a pattern to do this in C++ called pimpl
There's another useful pattern in C++:
struct iObj
{
virtual ~iObj() { }
virtual void cool_func( int a, int b, int c ) = 0;
static std::unique_ptr<iObj> factory();
};All operations on `cairo_t` have prefix `cairo_`
Notepad++ and vim are the 3rd and 5th most popular development environment according to stack overflows 2019 report [0] and I and many others I know use Vim at least.
> IDE knows types of things, type `myobj.` and you'll get a suggestion list with methods of that class callable from the current context.
I've never used notepad but Vim absolutely supports omni completion [1].
in C++ the ratio of ppl I know using an IDE with semantic code completion ability (that is, not rtags but actually using e.g. libclang or something that does understand C++ to some extent for IDE completion) vs "glorified text editors" must be 95/100 using an IDE though.
But your assertion is weird. Either you work on Windows where everyone is on Visual Studio, or you have 95/100 people around you using CLion, XCode, and Qt Creator, two of which don't even have an entry on GP's survey link. One of which is probably only used to sign apps for ios. Could you give more info?
TBH, in my limited experience, I have never seen code completion working well on large C++ code bases. I remember years ago VS literally becoming unresponsive when intellisense was enabled.
my_lib_cool_func(myobj, 1, 2, 3);
so you don't pollute the global namespace excessively. Having it accessed through the object avoids that.It doesn't really matter that much how things are actually implemented: most languages are flexible enough that you can do this with a class or a function or even implement the algorithm as a generic interface. But if you want the algorithm implementation to be reusable, you have to remove as many assumptions about the state as possible because otherwise some users won't be able to use your algorithm.
If you pick up a class, and the class owns the state, now you need to make the class generic to allow the user to customize the state, you need to allow the user to break the state invariants, to be able to unsafely read and write from it, because otherwise you are forcing the user to read into a separate buffer, and then make a copy to your class, etc.
Somebody that knows how to avoid all these issues when writing their algorithm as a class, probably also knows that by just using a function most of these issues just cannot happen anymore (or are much harder to introduce).
You can also write the algorithm as a function, that takes some generic state, and provide a "class wrapper" for convenience, so that those who don't want to customize anything don't have to. But then your class doesn't implement the algorithm anymore, it just wraps it.
I was only talking about algorithm state here. Check, for example, how the C++ string searching algorithms work to get a feeling of how these APIs work in the wild. For an example of this gone wrong, look no further than C++ stable_search, which does not let the user reuse the merge sort buffer across searches.
I don’t know what stable_search is, but it would be a design decision if/how it supports different memory management options for buffers, and isn’t really affected by the decision to use a class or top-level function for implementation. Either way, buffer allocation could be hidden in the implementation and out of the user’s control or exposed to some degree and under the user’s control.
A lot of your functions aren't going to need access to the whole state, only to some of it. If we're not using a class, we can just make these functions take only the needed arguments, rather than the whole state. This makes them easier to reason about, as we can know from the function signature that it only looks at the params we pass in, not the whole state. It also makes them easier to reuse elsewhere in the code, where we might not have the full state but we do have the arguments to the function.
For any function that doesn't need to access every member variable of a class, making it a class member function essentially creates an unnecessary coupling between that function and the class members variables it doesn't use. This makes it unecessarily harder to reuse elsewhere, as anyone who wants to use it needs to create an instance of the whole class, including any variables that aren't needed by that function.
Design trade-offs, as always. I get why you would do it your way, and it makes sense in some applications. But for some applications the fact that you're hiding which specific bits of state each function depends on, is actually the point of using a class.
Generally using classes when modelling persistent state is not an anti-pattern, because that's what classes are for.
void do(withThis, that) {
that(withThis)
}
Why be implicit with anything, when you can provide nice headscratchers for the next clown! ;-)Seriosuly though, most OO-languages brings not much to the table other than being "OO-centric" (methods+data vs procedures|functions). OO without encapsulation of data and implementation though, may even be worse abuse than no OO. So one can by convention even do OO without direct support for it in the language itself, by using any form of indirection that encapsulates data and implementation. REST typically fails this, because it often exposes direct access to internal resources (CRUD).
But why?
There's a lot of object-hating going around. Its silly. Use objects to encapsulate functionality, they are good at that and everybody understands what it means.
The context pointer for an encapsulated set of functionality, need not be passed around. It adds little or nothing. In a language full of encapsulated functionality, it stands out like Chekov's Gun. I'd argue, there better be a damn good reason to deviate for normal practice. And not just 'I prefer the explicit'.
I've seen monstrous state variables passed around to "pure functions" at least as much as I've seen monstrous god objects; the problem is around complexity and data flow design and has very little to do with function vs object.
You can encapsulate functionality without objects. There is nothing that you can do with an object that you can’t do with a closure, and the closure version will be 1/10th the size and have 1/10th the semantic programming language elements. That is why you see ‘object-hating’ all over the industry. It hasn’t lived up to its promise.
Objects aren’t simpler, inherently. It’s just what you know, today, and you’re unwilling to acknowledge your biases.
The reason information hiding is/was advocated was for reduced coupling. But in reality I find that this coupling becomes implicit which is arguably worse. But this is just my opinion.
If it does have internal state, then the answer it returns depends on the value of that internal state which means it's hard to verify that the answer returned is correct since you have to know the internal state of the object.
You test a class as a user by calling its methods in the desired order (per requirements and documentation) and you check that the results you get are correct. If the results are correct, the class fits your use case. If they are not, then you're either calling it wrong or it has a bug. Either way, no need to know the internal state.
This is my point. The order of calling the methods matters, because it manipulates opaque internal state. Since this ordering matters, the caller is implicitly dependent on this internal state as it must know the order to call the methods in to get the desired result.
So it's very relevant to the outside world in my opinion and nothing is truly hidden. It becomes part of the public API whether you hide it or not.
Think of a List class. Let's say it has methods to add, remove and retrieve elements, and it has a method to retrieve the size.
Let's say the internal implementation is an array with a fill pointer.
Obviously, before you can usefully call list.get(7), you must have called list.add(T) at least 7 times (assuming no deletes).
Do you need to know the the state of the backing array and the value of the fill pointer? Do you need to know whether there even is a fill pointer, or a linked list underneath (obviously, the performance would be different so you would need to know at some point)?
However, when it is hidden in the OOP sense it makes it much harder to reason about the behavior of a function without reading the source code. Because there is this internal "hidden" state that you do not know exists.
But a referentially transparent function is much easier to understand. For a given input, you get an exact output.
Testing is limited when you do not have the ability to fully control state, and understanding can be limited when information is hidden. The impacts of this are determined by the information being hidden, its relationship to the function's behavior, and the documentation of that behavior relative to the calling context.
Limitations on the ability to fully control state are orthogonal to the method of information hiding (opaque function parameters, internal object state, etc).
This is not an OOP/functional discussion, this is a discussion about the tradeoffs inherent in information hiding. And let's be perfectly clear here: These are tradeoffs, not black and white clear wins in either direction.
It's been a useful discussion though so I appreciate everyone's input.
I feel like it's difficult to talk about this without examples. The library I like thinking about when I think of passing around in a state variable is the lua C api.
https://www.lua.org/pil/24.1.html here is an example.
https://pgl.yoyo.org/luai/i/lua_State here are docs for the state variable's type.
This is pretty object oriented except that it's not using C++'s syntax sugar. You can't really do much with L except pass it into other lua "methods". The discussion up until your comment seems to be about the difference between using the syntax sugar and passing around a state variable yourself.
Information hiding and these other implementation details are kind of seperate. In my example you can't just configure L exactly how you want by messing around with it or inspecting its state directly. You have to go through accessor "methods".
I think if the argument is that OO promotes information hiding I can agree with that. I'm not sure about the point about the internal state becoming a part of the API though. APIs are contracts that can be met with different implementations right? If your implementation details are public then your contract is huge and inflexible.
Since using this API makes an implicit constraint that it is using a stack underneath the hood, why not just expose the stack for inspection? What advantage does keeping it an opaque blob have? I agree you don't want code manipulating this stack (although if it is immutable as in FP this isn't an issue), but since the external code already knows it is a stack and relies on that fact then it is part of the public API already.
I believe David Parnas introduced it in 1971 to mean that a program's design was sliced along shared units of concerns (things that vary together) rather than "steps in a flowchart".
https://prl.ccs.neu.edu/img/p-tr-1971.pdf
I believe what you are trying to convey is called "data abstraction", as for example used by Reynolds, 1975;
mentioned here: https://www.cs.utexas.edu/~wcook/papers/OOPvsADT/CookOOPvsAD...
explained here: https://link.springer.com/chapter/10.1007%2F978-1-4612-6315-...
Tests.h
#define private public
#define protected publicWhen you test opaque objects, you're generally not trying to test individual state transitions, you're trying to test that the class's interface adheres to the external promises it makes.
That said, many OOP languages have solutions specifically for when you actually do need to test those internal state transitions. C# for example has the "internal" keyword and allows you to declare friend assemblies, so you mostly get your cake and eat it too, at the cost of not hiding the code from yourself as the module implementer.
Whether you make it private or not hidden state leaks into the output since the output is dependent on it.
I don't want to worry about external methods modifying state, I don't want to deal with constructors, I don't want to deal with instantiating state or modifying state before I call my function.
My function just takes a state and returns a new state. Simple.
But it is - the programming language is now repeatedly quizzing you on something you've already told it.
method(state){
return doSomething(state)
}
Or this: class A:
otherThingThatMutatesState1()
otherThingThatMutatesState2()
otherThingThatMutatesState3()
method(){
this.state += 1 // or some other mutation
}Also, pieces of data that hide their internal structure and just expose an API to modify that data (Objects) are a very useful construct, and used in all languages with first-class mutation. Some languages have special syntax for this case (e.g. C++, Python, Go, OCaml), some don't (e.g. C). For example, no language that I know of exposes a mutable List type with a public count of elements. Neither does any language I know of expose a mutable List type as a closure.
I'm noting this because i think that number kinda detracts from your main argument, with which i agree completely. I think even if the code would remain roughly the same size, the removal of OO would still be worth it, since, as you mention, is much less jargon and conceptual baggage. And if i can achieve the same result with a much simpler conceptual model and language, then all the better :)
Closure and function return function can be used as substitute, and except for inheritance (which should be avoided too), I haven't found any use case for using class class. Example:
function MyList {
let data = [];
let add = (item) => data.push(item);
return { add };
}So for every list you have an additional allocation for the add function. It obviously continuously to get much worse as more "methods" are inevitably added. That's extremely inefficient for something that's used often and it's not something JavaScript engines can optimize away. The benefit is also extremely negligible since the closed over values are still implicit in the call to add.
This seems pretty central to your point, but a factor of 10 is also a very strong claim. I'd like to see sources.
Passing around state via variables is no more tedious than the extra syntax for class definitions, constructors, member variable access, extra semantics related to objects, extra keywords related to visibility, etc.
So instead of doing objects, you're doing closures that act like objects?
Closures - functions with data
No, I'm using using as-plain-as-possible data structures for my state, that I pass to functions that operate on it and return new state. Simple, composable, easy to test, easy to pass the state data elsewhere, serialize it or whatever else I might want to do. I'm not arguing against using classes or objects, but often its unnecessary and hiding mutable state in an implicit "this" variable is, in my opinion, something that often gets in the way of simplicity, as it often tends to imply mutable objects when in my opinion mutable data should be a careful decision rather than the default. Not necessarily of course, but even when you're operating on a constant "this", I feel that hiding the data that its acting on internally is still often not the right choice, especially if the object itself isn't constant/immutable: then its a coeffect instead of an effect, and if you're using truly immutable objects, then its really just a minor syntax difference to write obj.foo(bar) instead of obj(foo, bar). Object systems still carry around a whole bunch of other functionality that you may not want or need, but by using objects, a reader can't know if you are or not without reading the code or documentation. If its data and functions, it isolates the places things can happen somewhat. Classes and objects have their place, but I feel they shouldn't always be the default choice, just because.
That's my take at least, and a big reason I switched from Python and Java to Clojure, but to each their own :)
Often the best things to do in an OOP language is to remember that you don't have to encapsulate everything, and sometimes it's just about bundling data/state.
I may have highlighted the wrong point that I was responding to. I was making a point about the behavioral equivalence of objects and closures in response to this:
There is nothing that you can do with an object that you can’t do with a closure, and the closure version will be 1/10th the size and have 1/10th the semantic programming language elements. That is why you see ‘object-hating’ all over the industry. It hasn’t lived up to its promise.
If you're using a closure where it would have made sense to use an object, it's likely that the closure has the same problems to solve with encapsulation, reasoning about hidden state, etc. that make many languages' object related syntax complex, but instead of dedicated object syntax sugar you're solving it with dedicated closure syntax sugar.
Of course in reality its not that bad because you can also use classes as much or little as you want and can use them simply for single dispatch, or namespacing, if you wish.
I do find that in practice syntax and features matter because they push you to think in a certain way and it can be hard to break away from that. For example, I write a lot of Clojure and really like functional languages, but when I use Java, Python or even C++, its all too easy to slip into the OOP mindset even when I set out to just "do functional" in those languages. Maybe that's my failing and its certainly not OOP's fault, but I think in practice (at least from observation) it seems most people are like this.
But yeah, I don't actually disagree with you on what you said.
All you need to do is define a function and call it.
Literally, again I urge you to take a look at the following example, it's waay more simple than OOP:
Except that it's actually more tedious implementing your own bootleg object. Then managing the scope, namespace and pointers. When they could all be contained inside a language construct guaranteed to follow the rules.
How do you prevent memory “leaks” due to the banana referencing a forest problem? With JavaScript libraries returning a function pointer, the returned function pointer would often have a closure that contained a lot of irrelevant state. There is no easy way to reason about the leak, there is virtually no way to work out the cause at runtime (debuggers can help, but a problem with increasing memory usage certainly is not easy to diagnose why). Even when really conscientious about the problem, it is seriously hard to avoid the problem or inadvertently create a closure e.g. many programmers are unaware that arguments may not get GCed if you add an event handler within a function (closing over arguments and maybe other variables).
Objects have their problems, but the memory leak problem with JavaScript closures is insidious and very very hard to fix.
But no one is going to write a big industrial project in assembler today, because it's the wrong abstraction for the job.
Likewise with objects. They make it easy to switch conceptual levels in a domain in a way that projects don't. They take a bit longer to code, but with careful coding you can reuse them ad lib.
The OP is making the point that OOP is the wrong abstraction if there's no reuse. And that's perfectly true.
But there are situations where you want to say "Make a thousand of these items which respond dynamically and somewhat independently to their environment but which still support some kind of top-down management..."[1]
You're going to have a bad time trying to do that with exclusively with closures. Of course you can probably make it work - but that doesn't mean you should.
[1] E.g. particle and/or crowd simulation in games/CGI.
So yes, object-hate characterises the popular but tiresome anti OOP sentiments going around.
OO couples behavior with state. Sometimes that's exactly what you want (e.g. containers). And sometimes you just need data, functions and namespaces and all 3 are available in C++ outside classes.
If my state has some invariant, I want to be sure that I can trust it. I don't want to worry that someone will mess with it, breaking the invariant, or accidentally pass in stale state.
If that's all handled automagically by the object, it's out of sight, out of mind. Add in a get_state() and set_state() method if you need it as an escape hatch or for testing.
https://news.ycombinator.com/item?id=23338700
You will see that it is far more concise and intuitive to do it with a function.
Sadly objects are also used in an attempt to hide data (because direct access to data is considered dirty in OOP).
Now, consider a non-trivial OOP class hierarchy. You have an object that needs to mutate another object's data (an object that is completely unrelated to your initial object). In order to solve this you have options:
- completely rearchitect your code so that you take this use case into account
- find the shortest path between these two classes and add setters to each class
- hack the code and call it "technical debt"
First one is not feasible, second one is unmanageable (complexity increases dramatically), third one is usually chosen.
Speaking from experience, OOP is most of the time the worst tool for the job.
Yes breaking encapsulation willy-nilly is a bad thing. So don't.
I don't doubt there are bad programmers out there. Blaming the hammer for a bad carpenter is foolishness.
Btw objects are used for like 5 different things. Choose what works for your problem space.
Call it whatever you want, the requirement is for one class to produce a side effect in a completely unrelated part of the code (happens all the time). This doesn't fit with the current architecture, and there was no way of predicting it beforehand. Now what? Rewrite everything? What happens if you're in a department of 30 developers? Tell them "hold on while I go back to the drawing board to create a clean design"?
OOP isn't a hammer, it's duct tape, and carpenters are taught from school that duct tape is the main way of doing carpentry. It's duct tape because it ties data to functions in a way that makes it very difficult to extract or reuse that data somewhere else.
When OOP is your religion you don't blame the bad carpenters, you blame the master priest.
But even if you have free functions, as they become more complex, you tend to group them together, so at that point you might just slap the "class" keyword somewhere at the beginning of the file and just call it a class.
I don't understand this. The difference between passing a structure implicitly or explicitly will always be a single argument.
>But even if you have free functions, as they become more complex, you tend to group them together, so at that point you might just slap the "class" keyword somewhere at the beginning of the file and just call it a class.
Sure, but at that point you're not really using any class features. If you're just using classes as modules, why not use modules directly?
proc doThing(x: MyObject; value: string)
can be called both with doThing(x, y)
and with x.doThing(y)
where x by itself is 'just' a data object, but is treated transparently by the language as having associated functions, setters, getters, etc as if it were a class (even though they can all also be directly called separately).It's an approach that I find really elegant and that I wish more languages used in some way.
class MyObject:
def doThing(self, value: str) -> None:
# whatever
pass
obj = MyObject()
obj.doThing("value!")
MyObject.doThing(obj, "a different value!")
The only difference is that doThing is namespaced to the MyObject class.That's exactly what a 'class' is.
I've done this before for Munkres, also called The Hungarian Algorithm, which assigns jobs to workers. Logically the interface is a simple function, but internally it's very useful to have a class with multiple private methods.
class Foo:
def __init__(self, a, b):
self.a = a
self.b = b
def f(c):
return foo(self.a, self.b, c)
And then you create an instance `obj = Foo(a, b)` and when you need to call `foo`, you call it via `obj.f(c)`. This is a long winded way of saying an object is a poor man's closure. # Signature of f is (a, b, c) -> r
g = curry(f)
# Signature of g is a -> (b -> (c -> r))
# i.e. g is a -> blah,
# where blah itself is a function taking b, etc.
Admittedly you can get partial function application out of currying (but only with parameters at the start of the parameter list): g = curry(f)
h = g(a)(b)
# Signature of h is c -> r
But, although they're related, they're different procedures overall: sometimes you'll really want the curried function without immediately applying some arguments to it.Anyway, replacing "currying" with "partial function application" in your earlier comment: I suppose you're right to some extent, but the "this" parameter usually isn't that hidden and in Python you even pass the "self" parameter explicitly. I suppose the real major thing is you've put a bunch of parameters into a single class/tuple/struct rather than spelling them out individually; if you had a weird language that didn't support structs etc. then some partial function application would certainly be a very useful alternative to writing the same argument list out many times. But once you've got the struct, making your utility functions be methods is just minor syntactic sugar on top of that.
> I suppose the real major thing is you've put a bunch of parameters into a single class/tuple/struct rather than spelling them out individually
Yes, this is the mechanic, ultimately. Partial application vs structs/objects/tuples are just two different solutions for the same problem. Arguably they may even reduce down to the same underlying solution (to the extent that closures are objects behind the scenes).
(b) seems like quite a bit of trouble, since the clone function must either copy the underlying graph by default or provide a mechanism for swapping out the underlying graph that fixes up the internal map of (node->distance) and the internal priority queue to refer to things in the new graph.
Either of these cases are different from the provided example: The example will only ever compute a single value per object, always returning that value from a single method with no arguments (other than self).
class BFS(object):
def __init__(self, origin): self.origin = origin; ...
def next(self): ...There are plenty of cases where it makes sense to wrap a function in an object, but the point of the article is that this isn't one of them!
queue <- { root }
while queue is not empty
node <- dequeue one node
# One of the following:
# for a simple function approach…
if predicate(node)
return node
# or for a monadic approach, can be overloaded generically…
yield! node
for child of node
enqueue child
The monad-style extensions fits with the other pieces of the language much more naturally IMO. You could easily use a predicate monad to regain the functionality of the naive implementation, or a more complex monad if you want a complex BFS traversal, but the function looks almost identical either way.Python and Ruby and Lua and JavaScript even do this! C++20 will probably make this somewhat more common too.
(The word “monad” may set off sirens here, but it really just refers to a feature similar to but somewhat more generic than “yield” in a coroutine/generator.)
Below are two short examples of two different ways to do this:
type Node = Null | {value: int, left: Node, right: Node}
//a Node is NULL or a value that is an int with left and right nodes
//procedural with mutation, saved state is placed in parameter: found
def find_n_and_print_all_DFS (node: Node, search_value: int, &found: Node) -> Void:
if node == Null:
return
else:
print(node.value)
*found = node.value == search_value ? node : Null
find_n_and_print_all_DFS(node.left)
find_n_and_print_all_DFS(node.right)
//functional, no mutation, saved state is returned
def find_n_and_print_all_DFS (node: Node, search_value: int, found: Node = Null) -> Node:
if node == Null:
return found
else:
print(node.value)
found = node.value == search_value ? node : Null
left = find_n_and_print_all_DFS(node.left, found)
right = find_n_and_print_all_DFS(node.right, found)
return left != Null ? left : right
I would argue that if I used classes the code would be way harder to read and much more verbose.You can essentially use tail recursion on that DFS and pass in some accumulator as a parameter into your DFS or whatever. That accumulator can be some pointer representing some saved state or anything you want. There is still no need to group everything together into an Object.
It's a bit trickier but in the second example you can see that you don't even need to mutate anything at all. What you described can be achieved without redundant OOP classes and without mutating state at all! Additionally setting found to have a default value of Null negates the need for the extra accumulator parameter during function calls as you can just call it like this:
saved_node = find_n_and_print_all_DFS(x, 5)
I would still say OOP is an anti pattern because it overall promotes adding these mutating parameters in your code even when you don't need save states or anything of that nature. In OOP, the existence of getters and setters and methods for accessing variables outside of the definition of the method itself heavily promotes this type of coding style regardless of whether or not it is needed.The procedural style forces this additional "saved" feature to be evident as an extra parameter. If you never use that parameter it becomes obvious that the parameter is redundant while in OOP mutating external state is an intrinsic part of the style and like the OPs example you have to go through several logical leaps to see how redundant it is.