Invented here syndrome (2015)
mortoray.com
mortoray.com
1. The ease of creating a library these days is resulting in a proliferation of utterly crap libraries, so bringing on a new library is more and more of a liability, both to future reusability of code and to security.
2. As I get better at programming, I more and more often realize that I can write a better version of the library, or just one that better fits my needs.
These combine to create a situation where I'm less and less likely to want to import something.
1. The proliferation of libraries enables devs to get very, very picky about their libraries - the slightest wart or lack of understanding will cause them to label a library "utterly crap". But anything they create to replace it will run afoul of another programmer's picky tastes even worse! (And to what degree they may have a point, exactly 0 devs will be previously familiar with the warts and misfeatures of this new code.)
2. As I get better at project management, I realize that myself and my fellow devs are constantly underestimating the ongoing maintenance burdens of new code. It's hard enough getting us to make sufficiently pessimistic time estimates for the initial implementation that PMs don't have to smile and multiply by large integers - nevermind accounting for the next few years of adding additional debug logic, logging, fixing edge cases, supporting legacy decisions, adding a full set of fuzzing and regression tests, etc.
I wonder if we're regressing toward the same mean, or away from it...
Ah yes, the "it's all subjective" argument. The thing is, it's not subjective. Sure, it may be hard to objectively measure the qualities of a library, but that doesn't mean we should give up and assume there's no underlying objective truth.
My "picky tastes" are for libraries that are battle-tested, performant, and scalable. If that runs afoul of someone else's tastes, frankly, I don't care.
> 2. As I get better at project management, I realize that myself and my fellow devs are constantly underestimating the ongoing maintenance burdens of new code. It's hard enough getting us to make sufficiently pessimistic time estimates for the initial implementation that PMs don't have to smile and multiply by large integers - nevermind accounting for the next few years of adding additional debug logic, logging, fixing edge cases, supporting legacy decisions, adding a full set of fuzzing and regression tests, etc.
I definitely agree that the ongoing maintenance burden of new code is consistently underestimated, but it doesn't follow from that that we should prefer third-party libraries to our own code. A new third party library is still new code, and it's not often a reliable assumption that someone else will maintain that code forever. Further, general-purpose third party libraries aren't tailored to our needs, so they often contain more code than we would write. The result? We often end up with a larger maintenance burden by using a library. And finally, even if someone else maintains your library, that doesn't mean they will maintain the integration point between your code and theirs. Keeping up with changes in a library that doesn't value reverse compatibility highly can be just as costly as maintaining a library yourself.
Some things are, some things aren't - and even some of the subjective objections can be reasonable. If your entire dev team is a bunch of crusty C programmers with no exception handling experience, introducing a new library that throws exceptions has drawbacks, no matter how reasonably they might be used. They're simply not in the right mindset for it. If it's just 1 member of your dev team, however, perhaps it's time for them to learn to deal with exception handling if the library is otherwise good.
Then I see devs getting very passionate about vanilla XML vs JSON, CamelCase vs snake_case, and making kneejerk reactions to well documented edge cases that their replacement implementations would also have (minus the documentation.)
Your tastes are a bit better, but even there I can think of scenarios where all three aren't vital to me - easily fungible dev-only debug tools come to mind.
> Sure, it may be hard to objectively measure the qualities of a library, but that doesn't mean we should give up and assume there's no underlying objective truth.
Agreed. But I find that the best way to measure is to have at least one person try - or have previous experience with - using the library.
Obscenely convoluted dependency chains that can't be simply checked into VCS? Consistently unstable? Insecure by design and defaults? Constant public interface churn? I agree that there are reasons to kill a library with fire - sometimes before you decide to integrate it in the first place, sometimes after the fact when you've realized it was a mistake.
> The result? We often end up with a larger maintenance burden by using a library.
I see this rather rarely, which does make me think we've spent time on opposite ends of the NIH spectrum. It does happen though, and those are libraries that I will still avoid.
Agreed. The big tragedy of Open Source is that after its initial wave of success it good flooded with freeloaders. If you allow me to repeat myself: The one flaw of the GPL is that it did not offer provisions to prohibit distribution in binary form. "You want free? Better you know how to unpack the tarball and build the project from scratch... all by yourself (No build scripts beyond Makefile permitted, thank you very much)".
Of course, such overzealous GPL would need to also make provisions to accept dual licensing with nice, commercial, royalty seeking licences. That way, everybody gets what they value the most: You want freedom, you pay with sweat; you just want access, you pay in cash.
Creators are going to create, they cannot help it. I do not say that you cannot provide altruistic value to others, but my opinion is that you should care about your own people first. And the big mistake of programmers as a profession is that we do not see other fellow programmers as our own people.
It does not have to be one or the other, though. The obvious solution to have the best of both camps is to find an open source library with a (mostly) sane architecture and either support it actively or fork it (depending on the bullshit-to-effective ratio in the code and user base).
If you end up forking it, you may end up outcompeting the original if you band together with like minded devs/users from the previous incarnation... instead of keeping it behind your org's firewall.
And this isn't that strange of a phenomena when you think about it. Construction workers choose hammers rather than forging them on their own, we buy food from supermarkets instead of growing it on our own, etc. One of the principle elements of humanity is our ability to stand on their shoulders of giants, which shouldn't be lost
I don't think it's really about the time you have to devote. Integrating a library can take longer than writing the code yourself, and that doesn't change with respect to the time you have.
> And this isn't that strange of a phenomena when you think about it. Construction workers choose hammers rather than forging them on their own, we buy food from supermarkets instead of growing it on our own, etc. One of the principle elements of humanity is our ability to stand on their shoulders of giants, which shouldn't be lost
This fact isn't lost on me. My point is that there are actually relatively few giants who have solved an average arbitrarily-chosen problem, and even fewer who have shared their work. If you have 10+ years of programming experience, it shouldn't be hard to find an area where you are the giant.
This goes triple if the libraries in question are commercial closed source things. One company decides to stop supporting that product or the company goes under or gets bought out and suddenly your project is anchored to something that may or may not stay working as the OS changes underneath it.
Every library provides a benefit especially on the front end with reducing development time, but also brings liabilities. So you need to strike a balance between getting enough utility out of it and not loading your project down with so many liabilities that it will be impossible to maintain or port.
The maintenance overhead of keeping your own code working is usually far less than the maintenance overhead of chasing API changes in the libraries or discovering new bugs that crop up in minor version updates. Even if you go and pull the libraries directly into your project (so you're fixed on specific versions), you still have problems with having a far more complicated build process that fails far more often thanks to problems with the library builds. Or worse, you have version conflicts with installed libraries.
In sort: Keep your list of dependencies as short as possible, but no shorter.
It's not like I'm going to decide to write a new 20k line library for my project, but if something is only a few hundred lines of code it's definitely easier to just rewrite.
That's a good sign. As you get better and more experienced as developer you start to realise the benefit of thinking things through before launching in to writing code or reaching for an existing library. Some times you want to do the first, and other times you want to do the second. There isn't a straightforward "one size fits all" rule to say which approach is best.
95% isn't "only". That's pretty darn high. I'd wish that in-house code covered 95% of requirements more often.
If we're talking about open-source (I can't see why not), you can just fork the library in question and adjust it to fulfill the missing 5%.
> Being available in a package manager, being downloaded by thousand of people, or having a fancy web page, are no indications of a good product.
1. Being used by thousands of people isn't by itself a guarantee that it's good, but it surely makes it more likely.
2. Obviously noone sane would pick third party libraries based on how fancy some web page is. That's a strawman argument.
> If we're talking about open-source (I can't see why not), you can just fork the library in question and adjust it to fulfill the missing 5%.
The percentage isn't relevant if working around the design of the library to fulfill the missing 5% takes 30x as long as just writing the code yourself. Additionally, libraries often do a lot of things you don't need, so you get a bunch of bloat along with the stuff you want.
> Obviously noone sane would pick third party libraries based on how fancy some web page is. That's a strawman argument.
We can argue about the sanity of people, but the fact remains that people on teams I've worked on have chosen libraries based on the fanciness of their webpage.
Of course, but one should be aware of a tendency to overestimate the effort required for tweaking third party code (understanding somebody else's code is always more difficult), and underestimate the effort required for writing the thing from scratch.
It's closely related to the infamous "Big Rewrite" problem.
> the fact remains that people on teams I've worked on have chosen libraries based on the fanciness of their webpage
I believe you, but I find it difficult to believe these would be the same people who tend to introspect on their profession a lot, eg. by following software engineering blogs, so the author is kind of preaching to the choir in my view.
It's not just the cost of tweaking, it's the cost of keeping it up-to-date, too. I agree with the general gist of what you're saying, just wanted to point out there's more to it.
> I believe you, but I find it difficult to believe these would be the same people who tend to introspect on their profession a lot, eg. by following software engineering blogs, so the author is kind of preaching to the choir in my view.
On the contrary, I think people who don't introspect about what they're doing are more likely to read tech blogs. How else would they find out what the latest fad libraries are so they can treat them as best practices? :)
So if you need only 5% of functionality of a huge general-purpose library, it may be still actually faster to write that 5% functionality by yourself, than to understand, tweak and later maintain the library.
Depending on the situation this can mean having to understand lots of gory details of the implementation of the library. It is often easier to write an own implementation from scratch than having to understand these details and subtile interactions.
Typically "open source" just means that the source code is available under an OSI-compatible license and not that there is also a documentation available that guides you from barely understanding the library to understanding every implementation details that is necessary to do internal changes.
You're not wrong, but you are assuming some things:
1. That the library is written in a language suitable for the task. E.g. if you're doing real-time programming you probably don't want Visual Basic and if you're doing security programming or text processing you probably don't want C.
2. That the library is written well enough to read. You'd be amazed how much truly bad code is out there. It runs, but … it's not good. OpenSSL is an example.
3. That the library is architected in such a fashion that the 5% functionality can be added without major rework. The functionality may be small, but adding it might require quite a lot of architectural change and refactoring.
Anyhow, we're talking averages here - the default mindset, if you will.
Tell some junior to mid-level developer to write any significant piece of application and the host of bugs and NullPointerExceptions will follow. Finding a library for everything and only leaving simple gluing them together as a task for developers seems like the only way of ensuring minimum required code quality.
I would tend to consider training up of any permanent staff on the team to be part of my job. Perhaps not on basic technology, but certainly on what we're doing. That way, when I leave, there's still knowledge in house.
This may seem to limit business opportunities, but I consider it part of a job well done.
Travis is definitely more appealing to me than maintaining my own Jenkins instance. But I can't help but feel that we would have caught this guy's crappiness much earlier if he had to set up his own Jenkins job instead of a travis.yml file.
Maybe not the most advisable route for commercial work, but one that keeps me satisfied.
One of the first things I do when asked to solve a problem or implement feature X is to check around to see if there is an existing solution. Its very likely that I will find something that at least is close to what I want, and from there I evaluate the library, and ask these questions:
1) how complex is the problem that this library solves? How difficult would it be to implement it in my app?
2) How close is this library to what I want? Is it easy to integrate or do I have to jump through a lot of hoops?
3) How bloated is the library? Do I have to include a bunch of stuff in the production build that I'm not using? (not always a problem depending on how the language imports libraries)
4) Is it well written? Is it actively maintained? Github stars and issue backlogs don't say everything but can provide good heuristics.
5) Do I expect that I will need to customize or optimize the underlying behavior in the future?
Ultimately, I think the biggest question here is "Is the library working for me, or am I working for it?" If the latter, maybe its worth considering writing it yourself.
It also depends on the language and ecosystem for me, not to mention individual libraries. I'm working with javascript primarily right now, and there are a lot of npm libraries that are so small it makes more sense to just copy/paste the code into a utility (after doing due diligence it of course) and iterating on top of it. But it doesn't make sense to rewrite, say, JQuery, or React.
edit: grammar
The real question is what will create the least maintenance burden in the future.
Time and again, I see projects where the developer(s) is/are unable to do anything beyond one line after another of calls into low level library primitives, or maybe filling up a bunch of directories named after design patterns which plug into a framework.
Beyond that, though, and there seems to be a widespread inability to recognize repetition and then refactor it out, or to make layers to separate domain/business logic from low level details.
Also, the fairy tale of "self documenting code" is widespread, vs say "literate programming", and even if many of your in-house staff made their own libraries, the interfaces would be hard to understand. It's amazing how the act of trying to document an interface forces you to keep it clean and orthogonal.
Mixing developers with different "styles" (inline vs abstraction) is going to be frustrating for both extremes.
I tend to agree with the author in his opinion, but much like "The Big Ball of Mud", you need to understand the forces that cause these patterns to emerge.
For code that falls outside the company's core competency, is it any surprise why companies are more inclined to use an external library?
I credit those classes with sparking my interest in programming, but they didn't actually cover it to any extent. Instead there was this attitude that you could just find a Javascript file or Perl CGI script that would do what you need, and when necessary you could shoehorn it in, despite having little to no understanding of the involved programming languages. GitHub didn't exist yet, so there wasn't a centralized, trustworthy source of free Javascript modules; instead you had to find them on sketchy ad-riddled sites and hope they did what they said they did. I remember there were even sites that would sell scripts, a concept that feels totally alien to me now in 2016.
tl;dr Invented here syndrome is a big thing for web designers who aren't developers