Death by 1000 layers: the perils of over-abstraction in Java (2017)
quantcast.com
quantcast.com
> Good Java should resemble C
The recommendations are a mixture of things that are... questionable, and things that really apply to basically any big refactor regardless of paradigm, language, or problem.
But getting back to the issue at hand: what I see from the Java/C# world is a mentality of "abstract unless there's a reason not to", instead of "don't abstract unless there's a reason to".
Charitably, this approach might be compared to the micro services pattern: if you're shipping your org chart, and your org chart is thousands of people across four continents, then maybe eagerly building out those fire-walls of decoupling can be useful. Stereotypically, at least, this seems to line up with the kinds of companies we tend to think of as using these languages.
But where it becomes a big problem is when it's used in a one, two, five-person project, and all those walls just become massive speedbumps, not only slowing you down but making it hard to drive straight.
When I was starting out and reading lots of books and blogs this was very much my mindset.
Must use all the patterns!!!
But after working on a huge code base with many many many layers of abstraction, I now am of the mindset of just keeping it dead simple. If I need to get something from the database. Just get it. Introduce an abstraction when it’s needed. It’s easy to add an abstraction, it’s harder to take it away.
Sure, you might want to say that I should just use IDE which will show me all public methods in a GUI tree control. But I don't like it, I like code. May be some feature which completely hides anything in the class except public method declarations could substitute, but I did not see that feature yet.
Honestly I don't know what's wrong with an interface, even if there's a single implementation. Idea and Eclipse allows you to jump to implementation with a single click if that's what you're looking for. JIT will replace virtual method calls with direct ones.
I once had to work on a code base where every SQL query was inside a DaoImpl (ok) which was an implementation of a Dao interface (questionable) which is then called by a BoImpl (why on earth?) which is an implementation of a Bo interface. By your logic it would be sensible to add even more interfaces even though the intermediate classes do nothing except pass a function call to the next layer. Want to add one parameter to an SQL query? Well vbezehnar said it's fine to modify 6 or 8 files to accomplish such a trivial task.
The irony is that the whole code base was garbage and hardly readable and the biggest culprits weren't the overuse of patterns (they were just the cherry on top) but rather the fact that a random bundle of libraries has been misused for purposes that they were not designed for. My personal favorite is keeping hibernate objects beyond the lifetime of a hibernate session and the hibernate session was configured to live as long as the http request. Combine this with a web framework that completely abstracts away http requests and it means you can never use hibernate correctly because you never know when the session has expired.
If you can, set up a real DB, cache, whatever, and run your services in tests with real instances of their dependencies. Obviously this falls through for external APIs and things that are very expensive, but your savings in "But the tests passed" confusion time will be huge.
*I still think unit tests are great for tricky logic, but a lot of coding is just gluing together different systems where unit tests don't have a lot of bang for their buck.
I too generally only write unit tests for tricky pieces of code, especially when it's dealing with concurrency which tends to be difficult to test deterministically at the API level. There are other good reasons to write unit tests, such as execution speed, complexity of setting up a testing environment, faster development feedback loops, etc.
Maybe that'll help you convince your managers to invest in the kinds of tests that make sense for your product.
If you want to test how your code behaves when getting/setting data from a database, you fundamentally cannot mock the database. If you want to test how your code forms arguments to send to a dependency, you probably can, as long as you do both. If you're just testing that your args are correct without actually using them, you'll never really know.
A specific example stands out in my mind when I watched a TDD-minded developer suggest that their code was complete because they injected a mocked version of DynamoDB into a service and then verified their tests ran. They committed and pushed without actually testing against DynamoDB. When tasked with _actually running it_, they found that there was a fundamental flaw in the way that developer perceived Dynamo's behavior, which obviously meant their mock was totally useless.
If you want to test that you're pulling the right ID off an object to send to a database you can mock. If you want to test how your code actually behaves with a database, including failure scenarios and invalid arguments, you cannot.
As usual, solutions are much more nuanced than I previously stated. I've moved away from mocks, not banished them entirely.
But mocks could be useful for testing that your code correctly handles errors that are returned by an external component.
This is what makes mocks of limited utility. In the end, they don’t prove the code works and they are expensive to create and maintain.
I also refuse to measure test coverage in my codebases. "How frequently do bugs show up in production" and "how frequently are bugs fixed without adding tests" are metrics I find valuable but are underrepresented in the testing space. If bugs don't make it to prod very often, and when they do they are fixed, your testing strategy is probably sufficient. There's no reason to write thousands of null checks and formatting validators if you don't need to. Tests require as much or more maintenance as code. There's no reason to write more of them than you need.
You should shoot for both, actually. If your tests require a live DB to run, they can't be automated (or they can be automated, but failures won't tell you much useful).
There's a non-unit-test value of having good code coverage, too: being able to run any function independent of the rest of the app, with arbitrary inputs. When you have a problem that's difficult to reproduce in a controlled environment, it's really useful to be able to run a function that's normally ten function calls deep all by itself over and over again while trying to isolate a problem. If you have to run the entire app to run any of it, that becomes impossible.
Can you define a scenario in which hitting a real DB wouldn't present useful information while a mocked DB would?
If you're relying on a shared database you're already not really testing effectively. I don't really think this is a scenario to worry about (meaning, if you're in this scenario, you have bigger things to worry about).
> if the DB is down.
An interesting suggestion.
The use case of mocks/stubs/etc for me is “this thing is expensive to execute or external and can’t be validated directly”. In the former case, validating the inputs to the expensive interface is buttressed by thorough and verifiable testing of the thing it calls into. In the latter case, you provide an interface that behaves according to your understanding of the external service without requiring its availability.
In pure unit testing, the “external” boundary is hungry and consumes as much interaction between code components as possible. In behavior testing, you extend end user validation as far as is practical or possible.
The benefit of that is you can verifiably satisfy user needs up to the limits of your testing/testing environment, in user language. Wherever you need to silo who the “user” is (eg service/api boundaries), you can do so. But your tests are designed to validate user assumptions first, and fussy technical details after.
It’s hard (okay it’s not hard, it takes study and practice) to know what the “right” abstractions are, and it doesn’t produce IDE solutions (unless you’re working from a strong theoretical foundation).
It certainly doesn’t help when the type system doesn’t help. Lots of great abstractions are available when you can derive and infer meaningful types and derivatives at the end interface without a lot of fuss. It doesn’t just improve DX for library and product devs, it also improves design instincts.
A more functional style often fares better in my experience. Functions just accept “some data” and return “some other data”. Even without Haskelly algebraic data types, “some data” in a function that isn’t explicitly tied to a concrete type is a good prompt to think about the kinds of things that are common or disparate about the data and how to handle them.
It doesn’t necessarily lead to fantastic patterns. Most functional-style-but-idiomatic TS is not as well designed as equivalent code in an ML, error handling can be hell. But at least in my experience it has better outcomes than the OOP equivalent.
There's still a milion ways to do the same thing, and there are still smartass people going out of their way to produce code that makes people notice how unique they are.
All at the expense of wondering what the memory and performance will be.
I'm sold on "pure PHP" but you can get a hell of a lot done with a few simple libraries and avoiding all the bullshit pushing you into classes.
Typescript is quickly falling down this rabbit hole as well.
The balance I do like to strike is to use interfaces for any classes that interact with the outside world, and then have those faked out.
That way, I can use them to provide data to smaller units, or record that business logic recorded the data I expected.
Obviously this doesn't fix if the DB interactions themselves are correct, but it's a start.
I vaguely remember language support for something like this in the SatherK dialect of Sather. Sather itself was inspired by Eiffel. One-interface-per-class actually is no problem if it's well supported in the language, by simply adding a single special character to the concrete class name.
It worked something like putting some sigil (dollar sign, IIRC) in front of a concrete class name meant that the code was referring to the public interface exposed by the class. (I could have gotten it backward in that maybe you needed the sigil in function signatures/field declarations if you wanted to force the concrete implementation instead of the interface. I think it evolved from the normal Sather dialect, where interface declarations looked pretty much like class declarations except that the name began with a sigil. If it hadn't evolved from standard Sather, I suspect they would have required the sigil if you really meant to intentionally restrict your code to the concrete implementation.)
Java could have headed off a lot of premature abstraction by making it easier to later replace a concrete implementation with an interface. Instead of making "new" an operator, Java should have made "new" a static method returning an instance of the class, with a bit of compiler flow analysis to allocate an object the first time you referenced "this" within "new". The chief advantage is that it would have minimized the amount of code change necessary to replace a concrete class with an interface if needed. If new is a static method, then the interface's "new" can return some reasonable default implementation, or decide among implementations at runtime.
You'd still need to recompile everything, so that all callers would use the invokeinterface bytecode instead of invokevirtual. Alternatively, invokeinterface/invokevirtual could be collapsed into one opcode, and rely on polymorphic inline caches/hot spot inlining to remove most of the method dispatch overhead. A third option would be to have the bytecode verifier perform invokevirtual/invokeinterface substitution at class load time in cases where a concrete class has been changed to an interface or vice-versa.
Making "new" static factory methods instead of an operator would also make it possible for a class to transparently swap in specialized subclass implementations (for instance, specialization for BigInteregrs that fit in 64 bits, or Strings where all codepoints are less than 256). PyPy does a small amount of this specialization internally for Python, but it would be nice to have language-level support.
The Gang Of Four design patterns are great when you actually use them to solve the problem they were designed for. I worked on a product that grew in complexity and a factory was needed. Implementing was easy in Netbeans - right click and extract interface and provide the new implementations and the creation logic.
I'm not sure when using patterns all the time got so prevalent. As the saying goes, when all you have is a hammer, everything looks like a nail.
As another poster elsewhere said, the annoying thing is only the single implementation for the interfaces, and no real code that makes use of the abstraction. It only really alleviates other devs from asking "where are the interfaces for these classes?"
In my experience, Java developers don't use all the patterns, they use exactly two: Factory and Singleton.
I started exactly like this, and the pattern I wanted to use/learn was 'classes'. After struggling for a while, I realized that what I wanted to do was much easier to solve with a 'table' (list of lists) and creating functions that return the resulting table modified. It taught me that there is value in simplicity, in particular at the level you understand right now, or are at least not too far outside of your knowledge-zone. If you want to get things done, that is ;) Note that this was a few years before scientific Python and its ecosystem came into full-swing, I'd probably have discovered and uses Pandas right now.
Later I learned that this was called functional programming, I learned how classes work and what they are useful for (far larger programs than I was writing at the start of my career). C++ is a fixture in the scientific community (I'm now into Monte Carlo and solvers), and while I'm happy C++11 happened, the functionality of the C++ stdlib enables so many different styles that it is really hard to get a team to stick to a subset, and not get too distracted with all the nice new tools. Sure, much of it is nice, but how we can write C++20 is pretty much another language from C++03.
I guess I rambled a bit there. What I'm saying is that I'm generally critical of using new features, unless I've seen a convincing and for me personally concrete demonstration of how it saves mental CPU-cycles or boiler plate. This makes it easy for me (in the end, I'm not here for the code but for what you can do with it) and onboarding others (who, in science, might barely know that C++ has versions).
KISS KISS KISS. Nothing will convince me of the opposite :)
why is that a problem ? in the end ideally each of your program should be its own mini-DSL that solves the problem at hand in the best way ; successive C++ versions just allow your to create more and more expressive DSLs
If that sentence from the article is the only thing you took issue with, I'd wager you interpret it different than the author's intent. I took it as low on abstractions and mostly imperative.
> low on abstractions
I feel with this that the author is going to the opposite extreme. The takeaway isn't that abstractions should be eliminated categorically, just that they have a cost and shouldn't be fetishized.
> mostly imperative
I take issue with this too, though this part may be more a matter of personal preference. Yes, it's worse when your imperative logic is split across 15 files, but I see imperative logic itself as something to be done sparingly. Imperative logic means mutation of state, means easy-to-lose-track-of side-effects, means unpredictable behavior. This is sometimes unavoidable at the systems level (where C is typically used), but it is absolutely avoidable at the application layer (where Java is typically used).
Here's another one that raised a red flag:
> A well-considered class...should be easy to name, ideally with a noun that corresponds to something in the real world with a lifecycle of its own.
Real-world-modeling was found a long time ago to be a bad approach to OOP. The author is right to say that classes should represent clearly distinct concepts that are easy to name, but there's no reason those concepts also need to map to real-world entities. Often, they shouldn't.
There is something to be said about making sure it's clear where and when state mutates. But functional code has it's own set of pitfalls. I've seen otherwise good devs struggle like crazy reasoning about complex functional code. (Especially when it comes to performance)
> Real-world-modeling was found a long time ago to be a bad approach to OOP. The author is right to say that classes should represent clearly distinct concepts that are easy to name, but there's no reason those concepts also need to map to real-world entities. Often, they shouldn't.
I 100% disagree with this. Most applications are modeling domain logic and business processes as code. And ideally as many of your classes as possible refer to a concept or should translate to simple and understandable objects in your domain. In an ideal world every class would be named after a concept/entity your domain expert understands and uses.
*Unless you read real world objects as physical objects, but I think the author meant real world as in domain concepts.
This is not about abstract/non-abstract, Java, C or C# .
This is about 99% of people not being able to write decent OOP code. This is not about "programmers" not getting that first awe inspiring feeling when they first see decent OOP code, it is about not caring, not event after 15+ years of programming OOP languages.
This is why I always see endless @Service and @Entity code, completely separated, like it's 1996 and we have procedures and structures in Borland Pascal.
I embraced KISS maybe 2 or 3 years after starting at this company. I NEVER regretted it.
I still maintain a few projects started in my non KISS period. I regret every time I have to touch the code base.
Of course YMMV.
I've seen beautiful abstractions that make a lot of sense, a la DDD that saved companies a lot of money. I've seen a senior architect almost shed a tear when I told him I loved the way he used the Bridge Design pattern in a project and how easy it was to implement some feature because of it (I don't remember the example.)
Is the _exact_ behavior I have seen in 100% of the Java and C# codebases I've used.
> But where it becomes a big problem is when it's used in a one, two, five-person project, and all those walls just become massive speedbumps, not only slowing you down but making it hard to drive straight.
I think that most patterns exist to prevent bad things from happening, or to make things easier to understand. I think that in a team of 1-5, there just isn't enough output happening to make either of those things huge concerns. If you're very small, there's not a ton of reason to try to protect yourself from future-scaries that don't exist in your company yet.
And then, if the response is "But what about when we grow", that's literally the exact right time to abstract behavior. _As needed_.
It feels so strange to make C++ a kind of better language, when it was one of the main creators of that abstraction culture.
Naturally, in the last 20 years, the only C++ developers are those that didn't moved on, so now 90's C++ code gets coined with "writing Java in C++", the irony.
The desire to write new code is often so strong that I've seen Spring being itself re-wrapped in a local flavored "frameworks" that only brought value by ignoring parts of what Spring could do out of the box. The solution is most often to strip off the overwrapping and bring the app back to Spring primitives.
But I never gave in to Lombok until recently, I was seeing it as gimmick at best. Like.... Ok, it will generate getters and setters, so can any IDE.
Lombok actually makes writing and refactoring concise java code way easier. Eg: Macro like @builder allow to introduce immutability instantly. Notation of the form Dog.Builder().size(5).name("fido") are accessible for 'free' and are pretty reliable. ( read : you can tweak them if needed without looking a de-compiled code to understand what went wrong )
So, as much as I dislike Java and Spring. I think you are correct. A CRUD microservice can be expressed in reasonably concise way with the method you describe. ( and it's still plain old boring java code running on regular JVM )
But mostly unit testing. The commitment to isolated testable units of code is what generates abstractions of dubious value.
Don't use a class where you could use an interface: easier to mock. Don't use a static method where you could use an instance method on a collaborator: easier to inject, replace and mock. Your object graph getting too hard to compose, or other design smells? Don't worry, we have tools to handle the complexity, so you don't have to try and reduce it.
> intercept and rewrite code at execution time
I think "execution time" is one of the big issues, a lot of these patterns evolved from application server middleware like glassfish that would introspect and run packages and often provide services at runtime, sun and IBM were busy selling licenses and support for these applications servers so these abstractions were great for their bottom line.
Take away the execution time requirement and injects mocks/fakes/stubs at compile time suddenly you don't need interfaces to make mocking easier, static methods are easy to mock and you probably don't need DI/IOC.
You're right that Java had some long class names to start with, but I think it was mostly Spring that introduced the PerhapsOverlyDetailedClassNamingConvention.
Spring's introduction and subsequent wide adoption was a big driver of TDD's growth, which encouraged all those interfaces and mocks.
You can have a bean. Great. You can have a bean that's a singleton. Also great. Then you discover that you can have two different beans that are singletons, and you find that they have one thing in common: how the single instance is created and accessed.
Each such singleton bean will have about five lines of code that deal with the single instance, and these five lines will be largely similar.
So the Enterprise Java Architect sees that there is something to be abstracted, goes to town, and now you can have seven lines of code configuring the abstraction, so that you can avoid writing the five lines of boilerplate.
Unfortunately, the seven lines of configuration will be largely similar for different singleton beans...
Yeah, as a (mostly, currently) java programmer this is where the whole "framework" ecosystem loses me.
Dependency injection - absolutely. Really useful, I can have a setup method that reads config from the environment or wherever, sets up various things and passes them in (you need an SSL context associated to these CAs? Great. And you need one that needs mutual authentication? OK then, here's one I made earlier) . Makes testing easier, and makes it obvious where the setup is being done and how this stuff hangs together.
But 'full' IoC with a framework injecting stuff at runtime? Why? This just obfuscates as far as I can tell, and works around the language, making it harder to conceptualise and reason about.
(Also yes, I don't subscribe to interface-itis).
As it stands, the necessity to relax members to `public` all over the place drives a giant bulldozer through any remaining delusions that Java is a language where actual OOP can be achieved.
But I agree that Java could use better unit test and mocking support.
Mocks usually over-specify implementations by setting up expectations of a specific implementation conversation rather than an outcome. They're painful to debug after refactoring implementation - they end up write-only - and they generally inhibit refactoring of the mocked API - often mocked instances of APIs outnumber production invocations.
I try to encourage people to replace control flow with data flow where possible; messages, command objects rather than method calls; iterators, streams and consumers composed together, rather than loops. Data flow can normally be trivially redirected into a container, and if the data objects are simple inert immutable tuples, they're trivial to construct as inputs or assert against on outputs.
Fakes are good too, better than mocks in most situations, since they are easier to refactor.
Although, IME, integration tests, while slow and often brittle (especially if you have any async components and define failure conditions in terms of timeouts), have a significant upside in permitting large factoring while still being able to test a substantial amount of end result functionality (i.e. the stuff that matters, not implementation details).
Fake: A replacement module that behaves like a production module, but with certain simplifications, for example in-process vs. using rpcs, or simply cleaning up the filesystem after usage. For example, https://github.com/tk0miya/testing.postgresql. "automatically setups a postgresql instance in a temporary directory, and destroys it after testing".
Furthermore, these are external dependencies that can't be run in-process. If a dependency can be run in-process (aka library), there is no justification to ever mock it. I've even seen codebases that mock their own class B in order to test class A. Run the production code already. Ban mocking libraries.
Unit testing with mocks is another form. One does not replace the other.
The argument for static typing is often so that refactoring is easier and ide's work better.
I always thought it was funny that the core reason for needing fancy ide's and large rename and refactor capabilities was that static typing and over abstraction in the first place.
it's funny since in theory TDD should reduce API surface as it is supposed to make people think about API before implementation. But then , TDD and DDD intersect and it's not much TDD anymore when the domain model has to be implemented...
For example, C++ has a culture of caring about low-level details to squeeze out maximum performance. I once found myself writing something that wasn't performance critical in C++ due to a library I needed, and just minutes later I caught myself reading up on move semantics for maximum efficiency. I couldn't help myself getting sucked in by the language culture!
Java's culture, in my opinion, is one of abstractions. There's little in the language to force you down that path, other than decades of it being the prevalent way Java is written, which bleeds into the standard library, core frameworks, and becomes the Java Way of Thinking.
There's an architect programmer and a worker bee. Everyone wants to be the architect and to create the "enterprise" framework used by the worker bees.
Haskell's culture is also one of abstractions, but of a very different kind of abstractions than Java.
Yes exactly, that's why I personally stay away from Java. The language is totally fine, the tooling is also fine, performance is great, there's everything there in theory to build and maintain great projects, however, the culture is atrocious and overengineering is everywhere. And that's very hard to go against decades of culture momentum going this way.
Actually I rarely remember java class names or methods, I just know a heuristic shortcut on how to find them.
After 15+ years of Java, I don't even see the boilerplate code, my brain ignores it like it ignores adds on webpages.
And you don't see this as an indication that reducing code size is important?
Then, I just think, excluding some horrendous language and some corner cases, experienced programmers are limited by their brain and not the programming language, just like certain languages that are more verbose are spoken faster and terser languages are spoken slower, but the information flow is done at the same speed.
Because having to write de facto one file per class does not really allow you to write very elegant software. It forces you to write files over files and no one wants that, so they fall back to primitive types.
For the abstraction bloat it is actually similar. Here the reason is that for the longest time (and even nowadays) Java makes it painful to work with functions. So now instead of just a plain function, you have to create a class with an interface that has one method... and instantiate it. Great... not.
I don't blame the language for this.
It wasn't that hard to read/write objects using JDBC (it did usually create a lot of boilerplate)
Please name a major programming language that has built-in ORM features -- a library is the best place for this sort of stuff.
For Java this was J2EE/EJB (stupidly over-engineered) and Hibernate/JPA (decent if you avoid the footguns). For Ruby you have ActiveRecord, Python has Django and SQLAlchemy, etc.
And of course you can... purely in terms of organizing building blocks most object orientated languages are pretty much a super set. I think the problem is most people are not consciously aware that object orientated approaches are just a pattern, and like any pattern you shouldn't apply them to everything. Being built in to many languages gives the impression they are as fundamental as a loop or a function - those are technically still abstractions, but far more primitive and less subjectively applicable... Object oreanted patterns feel like they are in between domain specific and basic building blocks, the higher you go the more polarizing these become, trying to use a DSL will seem absurd for one thing extremely elegant for another - when classes fit they are beautiful, but when they are carelessly used as a default, they are annoying bags of loosely coupled functionality and data that only serves to obscure relationships.
I'd speculate that the prevalence of poorly applied OOP is more to do with legacy and culture of the language along with the focus on how to use the language for newcomers.
there's shiittton of focus on architecture, patterns, interfaces and generally testability
there's shitton of discussions about those topics in those langs
So it may be reason why people tend to over abstract stuff.
As a younger dev, it is really difficult to fight against these seniors. Most seniors that care have left, they got the whole market begging at their feet. Other seniors have given up. There's a small number of seniors that still care, and there are layers of bureaucracy between suggesting a change, incorporating it, and seeing the results. It isn't unheard of to have half a year pass by from the suggestion of a change to the first signs of its effect. For many a young dev, that's half a year away from moving jobs without falling under immense scrutiny.
Abstractions that do not fix any of PHP's problems by the way (like its chaotic std lib, its weird mix of dynamism and rigidity, bizarre alt syntax for control flow, all the error reporting cruft, adhoc variable creation as side effect of certain function calls, context-less output buffers, antiquated scoping rules...), PHP is just adding features to the language without removing the bad parts. The irony, it's so bad as being a templating language PHP devs are using another templating language on top of it(smarty,...).
Standard java libraries like java.lang, java.util, java.util.concurrect, java.net, java.io, java.nio etc, hardly suffers from that part.
This is the language - quite straightforward, hardly any getter/setters - relatively simple hierarchy and so on. Then java.awt and later swing, java.bean came with more getters/setters, deeper hierarchies. Those are defunct now of course. I'd blame it on the idea of having visual UI editors to cause the extra fluff.
Then xml, spring and all the friends came. Java got the reputation. Still I'd argue the core java packages don't suffer from deep factory-of-factory jazz. (I avoid most of DI and resent mocks in tests)
That's mostly just unpleasant though, beyond that it seems to be something deeply ingrained in the culture, even something as simple as a structure (a simple class with public fields) is a heresy in the java world, you need to make those members private and add getters/setters or you're likely to get warnings from the IDE and probably not pass code review. The enterprise world builds more and more abstractions on top of this but the problem is very deeply rooted.
A good part of the problem though is just programmers trying to solve hard meta problems instead of boring repetitive business problems.
J2EE 1.0 came out, and was both over and under engineered (baroque and unworkable, which is ironic given the 7 fallacies come from their employer). And then Design Patterns came out, and it was all over but the crying.
To me, it's not really about "seduction", it's about inviting the programmer to abstract away with tools that are heavily unfit for the job. If your only means of abstraction is a class, everything starts to look like a method call.
That's why I enjoy Common Lisp's ability to create macros on multiple stages of the language's operation: reading code, compiling code, substituting function calls. These make it possible to introduce new syntax where necessary, which in turn makes the code shorter, and to write compile-time optimizations that do not get in the way of the main function bodies.
I think the author's main idea is that your methods should carry some weight and be long enough so you can visually trace the logic. Do not write functions that do nothing but multiply 2 numbers.
> The abstraction is abstraction. It does not matter how it is implemented.
To me, it does matter. A good abstraction should not get in the programmer's way. This directly depends on making the language's syntax support the abstraction in question. This, in turn, depends on the macro system of the language or lack thereof. Finally, that is a question of the given language's expressiveness, or lack thereof.
Abstraction cost is that the system becomes more complex overall.
Good abstraction makes the system easier to reason about and leaks rarely, so you rarely have to think about the insides. But all abstractions leak under some circumstances.
If you have 100 layers of abstractions at any given moment some layer leaks and you're constantly fighting the system without being able to understand how the whole thing should work.
Language syntax only helps by removing the need for some abstractions by including them in the language.
Of course on its own more expressive languages make job easier to write and read. Not denying that part.
Maybe I'm missing something, but doesn't Java already have a function that multiplies 2 numbers, as well as a fancy infix notation for calling said function?
One of my biggest beefs with Java is that, by virtue of the one-class-per-file concept (pace inner classes), the class is elevated to the level of a big deal in the novice programmer's mind, so that a little class with a very clear and narrow scope seems to be unworthy of existence, and we get the String-for-everything syndrome. Anecdotally, I have encountered exactly this attitude on projects.
That's an interesting take, and I think you're right. In retrospect I have wondered why my more junior coworkers so rarely create new packages or new classes. Perhaps this is the reason. Will have to take that into account when doing reviews.
It was pretty nice, there was a proprietary framework to use one XML file per form with SQL and metadata to generate CRUD with SQL-backed grids, filtering, sorting, and simple actions, and everything else was added with straightforward C++ Qt code.
Almost all business logic was in PL/SQL. It was simple and quick and the divisions were obvious.
We moved it to 3-layered Java code with SQL (and some PL/SQL still) on the backend, then Jboss application server and EJB3 code, and Java Swing forms on the frontend.
The goal was to eventually move all business logic from PL/SQL to Java and avoid Oracle Tax (it was before Oracle bought Java BTW - funny how that worked out :) ).
That all would be fine, but our architects bought the cool aid and decided to overengineer the crap out of this software.
What was PL/SQL + 1 XML file + 1 Qt form became like 10 layers of XML files (1 specifying how to get the data from database, another how to transform it to DTO format, another how to represent it on the frontend, another what actions should be possible on it in the frontend, another how to transform from DTO to action object for each action, etc.
Yes you had to write less code in edge cases than in the C++ solution, but instead you had to juggle over 10 XML configuration files even for simplest cases where before 1 SQL wrapped in XML was enough.
It wasn't 3 layers of abstraction - more like 12 in practice. And all that configuration bought you was the ability to potentially write slightly less code when using several forms together and passing data between them (which we almost never did anyway).
It was developed for years and finally shipped (I think), but only to one customer out of dozens and we completely switched the technology again soon after for unrelated reasons.
function (inParam1, ..., inParamN, outParam1, ..., outParamN, outMessage): ReturnType
begin
savepoint SP_begin;
-- do stuff
outMessage := '...';
return RESULT_OK;
exception
when ...
when others then
rollback to SP_begin;
outMessage := '...' || exception details || ' ...';
return RESULT_ERROR;
end;
It's funny how I remember this after more than 10 years :)Made it very easy to add tests and to interoperate with container-managed-transactions in EJB3.
I even added an debug-only interceptor on EJB that created a savepoint before calling any PL/SQL and checked if that savepoint was still valid after the call. If no - that's a PL/SQL that isn't following the template and the developer who changed it needs to be flogged ;)
Until you're making your third one (of whatever), don't make an abstraction for it.
Perhaps the most absurd example of such a thing I can think of is when people figure out they can replace some gigantic mapreduce cluster with an awk one-liner and have it actually run faster.
Not the stream, map, reduce stuff.
But simply writing functions without side-effects and separating data from functions.
"Remember that software is at its core procedural."
That may be true in some sense. But most Java code is in fact not procedural but stateless in the sense that most often state is store in a database or another external system.
Isn't Kotlin largely scratching this itch?
This _can_ be a problem with Java, admittedly. Sometimes you've got a single meaningful implementation, but you need an interface either for using it with services that implement such interface at runtime to create proxies (e.g. Spring, Hibernate, others) that augment its functionality, or to properly mock a dependency. BUT: the IDE should be your friend. Almost every Java IDE supports "go to the implementation".
And, what is the alternative? Fewer interfaces sometimes lead to even worse patterns (for testing: monkeypatching and concrete class subclassing). I prefer the clarity of declaring an interface wherever an interface makes sense.
Using a lot of interfaces and defining APIs (which are, mostly, collections of related interfaces) makes the Java ecosystem highly pluggable. I can mix-and-match a lot of different libraries and/or implementations, as long as there's an interface they can work together in a rather good fashion.
Consider Python: there're very few (and sometimes poorly defined) interfaces and API; so, quite often you need to pick a library/framework which implements everything and works poorly with other implementations.
But are you writing a library? If not, your product does not have to be plugable. Also if you end with multiple interfaces with only one implementation you should start with classes and extract interfaces when needed.
And I think the Java abstraction root problem comes from the way of doing "unit-testing" at the class level. Your users don't care about your application architecture: they want functionalities. Method and class level unit testing is just fossilizing your code and giving you some kind of feel-good number. You should be able to drop all your code, start anew with a new stack and have your tests still usable. Refactoring should never mean having to rewrite tests.
I'm not writing a libraries, but other people are. In Java I can pick my IoC container, my ORM, my web stack, as I wish.
I can hook-in web-related functionalities through the servlet API filters.
If I pick Django in Python I get a builtin ORM and web stack. Everything is within Django, for Django; extensions are Django-specific.
In order for that to work, you'll need your interfaces to be stable. It's impossible otherwise :-)
You should not test code. You should test functionalities.
You should be able to rewrite all your app in a new language and still be confident your tests will tell you if something broke. "Unit testing" at the class level won't do that.
You're talking about integration tests and/or end-to-end tests. I write those. Usually at the façade level, not at the network endpoint (which usually means the controller level, a shallow layer that often includes some ACL and serialization) but it's just a matter of taste.
> You should be able to rewrite all your app in a new language and still be confident your tests will tell you if something broke. "Unit testing" at the class level won't do that.
Unit tests, integration tests and e2e tests are different things that solve different concerns. An e2e test will rarely expose an obscure bug within an implementation, and a unit test will rarely expose concerns in the global application flow.
Both have their place. And they have nothing to do with the "interface vs not interface", IMHO.
I don't remember how it is done exactly (it's in my muscle memory), but I think ctrl and hover over shows up a mini dropdown where you can click on what you want to navigate to.
We need to be able to bring that up at code review and put it into the style guide.
You can also refactor it back.
maybe it's a weakness of an IDE?
I am 100% sure Idea have some solution too.
Navigate -> Implementations.
Works on interfaces/classes as well as methods.
In one of my earliest jobs, I used to write j2me games. The code had 0 new allocations, mostly static stuff neatly organized in final classes. All of this due to hardware limitations. My point is, the language != culture.
But I don't really understand your point about idiomatic. Modern (new) languages, like Kotlin or Go, are less intended for teamwork or less idiomatic? That can't be true.
Theres a lot of stuff I miss about Java but the culture of layers is not one of them.
However, more closely related to the OP, has anyone else struggled with the organization of even medium-sized systems? Obviously some people manage, but I don't know how.
I used to do Go a lot more than I do now, but I think I've narrowed it down to (a) it's simplistic (boolean) visibility capabilities (b) lack of overloading and (c) lack of static functions.
Instead of `User::new()` and `Role::new()`, you end up with `newUserWithPassword` and `NewRole`. You can do a package per module - but that doesn't solve the issue with overloading, and you can get into cyclical imports. You can solve overloading with a config objects, but now it feels like you're back to over-abstracting.
A developers education, experience, and domain all impact how different abstractions are perceived. Sometimes having an abstraction can reduce cognitive load. Think of the configure vs convention arguments for frameworks. Knowing how to implement simple functionality the “right way” (within the paradigm of that abstraction) can make building software easier especially when collaborating with others.
I personally find consistency more valuable than having the perfect level of abstraction. Consistency, even if slightly more or less verbose than I might otherwise choose, means I spend less time thinking about the abstraction itself. It allows me to focus on the problem domain, which is where customer value exists. It has the added benefit of simplifying collaboration with others who are also familiar with the abstraction.
Was it worth it? Probably yes, not for the final code but for the experience of 'this is an exaggeration that you should not repeat' (although the resulting class diagram was a work of art).
As usual, extreme abstracting is bad, but extreme lack of abstracting is too. You should find a common term that make the current work simple enough, but that also allows for future changes to be simple to implement.
If you are making a list-like class because you need to keep a collection of strings which you are only comparing, don't use strings. Use <T> even if you only use it with strings, the extra cost while creating the class is almost nothing, but the cost of a future refactorization will surely be greater.
"Try not to open doors you don't need, but keep them accessible in case you need them later"
Allow me to question the "almost" in "almost nothing". Java generics are famously unreified, at runtime it's exactly the same.
I have noticed this habit of what could be called redundant genericity in my own code as well. Sometimes the motivation seems to be hardly more than a coping strategy for the notoriously verbose class names of idiomatic java. But it can also provide self-describing value: at the use site, Bag<String> reads better than BagOfStrings and on the implemention side, class Bag<S> makes it clear that nothing in there will look at e.g. individual characters. Obviously, if you do need to consider individual characters, go with BagOfStrings - or go with class Bag<S extends CharSequence>, which would notify the reader right in the headline that somewhere in the implementing it requires access to the actual text nature of its data.
As a bonus, even if production code only ever fills the angulars with one specific type, chances are that it's a type that is painful in a unit test. If you do esoteric set operations on WhateverFactoryFactories or something like that it's super-convenient if you can just play them through with strings.
On the other hand, decaying <T> to Object (rather than monomorphizing as C++ templates do) somewhat reduces the viable attack surface for optimizations.
On the gripping hand, Java is traced/JITted, so it should pick up many of these opportunities (and more) anyway.
Make it a little less readable for a future case you might never encounter?
This is less valuable the more specialized the type you're looking at -- generifying everything is just as much an antipattern as anything else -- but String is so widespread and abusable that it really is helpful to abstract over it in many circumstances.
To me it becomes effectively another variable I have to track through the code and adds a small amount of extra cognitive load. Not a lot, of course, it's not a huge thing and I do use them where appropriate.
It was - having gone through both extremes you have gained understanding of both and discretion in using either.
That is to say, you intentionally use classes and all the baggage of OOP when you know a part of a given implementation resource and throughput heavy. Otherwise, the default 'starting point' when implementing something is to start with value semantics.
It's fair to say many mainstream languages are adopting this approach. Swift is extremely functional, C# 9.0 adds record types, Go and Rust are largely functional but you can do OOP style things with them.
Java I feel has reached an evolutionary dead end and will be the 90's answer to COBOL.
> Computers do this thing, then they do that thing, then sometimes some other thing. High-level languages and object-oriented design are fanfare and pageantry distracting from our fundamentally humble jobs: pecking out if-then statements.
Then just use C and reap the speed & ram benefits!
Internal combustion engines are useful for other tasks, such as powering a generator.
We shouldn't confuse the task with the tool.
Where KISS, YAGNI and "premature optimization" are left in the corner crying.
Pointing out DRY is easy and during code review all devs know that someone will be nagging about it. Where if you build an abstraction even if not needed or plain wrong there is a high chance that no one will have any comments, because reviewers will go with low hanging fruit and point out DRY violations. It will be more work to find out how abstraction work and point out its weakness, and don't ever even try to propose that code duplication might be a better solution than broken abstraction.
Whenever I can, I try to push rule of 3 for dealing with duplicated code. Unless something is used 3 times it still is not a duplication.
This has been my biggest objection to many tutorials, videos, and books on OO, architecture, and patterns: They're useful for greenfield code, but so often in large systems, they are nearly impossible of implementation.
The trouble, of course, is that it's nearly impossible to teach how to do these things inside large codebases that are layered in ways that don't align with the intended change.
Consider that other than Michael Feathers' book, there are few, if any, good books on maintaining code. And so we all struggle, and code becomes progressively more imperative and less coherent.
This has been true of framework after framework, especially during that era, but we still have them regularly today.