Why I often implement things from scratch (2006)
armstrongonsoftware.blogspot.com
armstrongonsoftware.blogspot.com
"--If you need to invert a 2x2 matrix in your program, do you write down the algebraic formula for the solution by hand or do you call a linear algebra library? (that you would not need otherwise)"
The answer to this question is obvious to most programmers. Also, it is obvious to them that the opposite answer is clearly wrong and absurd. Unfortunately, not everybody gives the same answer!
My (camp's) reasoning:
1. Your handrolled solution is likely worse than the library anyway
2. You're wasting effort reinventing the wheel
3. You risk adding unnecessary opacity by littering the program with dense math
4. Other programmers will assume there is a pertinent reason you handrolled the algorithm instead of using the library. Now everytime someone touches that code they have to reverse engineer your approach to see what it does different from the library such that you had to roll it yourself (surprise: nothing!).
The two exceptions I imagine are 1) if you really do absolutely need some slightly modified microoptimization of what the library does or 2) if you are working for a platform where there is good reason to avoid even slight bloat (embedded).
For the people who read OP's comment and believe I am (obviously) wrong, what is your camp's way of looking at this?
This assumption led to a lot of O(n^2) left-padding happening over the years, via left-pad from node.js.
Ideally libraries would be well-tested and of high quality, but at this point there may be too many of them to vet.
My own 2 cents is that if it's easier to implement the code from scratch than to evaluate and choose a library, perhaps the invent-it-here choice works best.
And honestly, even throwing those libraries out, it's sometimes not even a question of doing a "better" job than the library. The library is designed for general-use, but my needs aren't always general-use. Sometimes libraries are good, but they're wasting time on stuff that I don't need.
This is the reason why every Unity game on the web takes 4-5 seconds to load, and my custom engine takes on the order of milliseconds to load. It's not because I'm a better programmer, it's not because Unity devs are crap, it's purely because I know exactly what my engine needs to do, and I don't waste time on anything that it doesn't need to do.
I can think of tons of examples where I've started out using an existing library and then realized that getting rid of abstraction made my code faster and easier to debug. Heck, I can think of tons of times where I've privately forked libraries and deleted codepaths or rewritten algorithms to make them more efficient for my use-cases. There are very few JS libraries that I regularly use where I have not at some point needed to care about their internals.
And I am definitely not a crazy, insane, masterful programmer. If I'm occasionally circumventing d3 internals or doing stuff manually, it's not because I'm special. People have this assumption that if something has a bunch of stars on Github, there's no way they could possibly code something more appropriate or efficient, and for a lot of people, that just isn't true.
Opens source code.
Closes source code.
"Well, looks like I'm going to be a Rust developer from now on."
... but I think the majority are in the “I just need shit to work and develop fast” camp, and don’t care if it’s not optimal, so they can finish their current project and move on.
The problem is, 20% of projects are skyscrapers, and you can't build those out of pine. Sometimes, you can't build them out of standard steel and rear. You actually have to pick the correct raw materials, create a stable and maintainable architecture, adhere to strict regulatory standards, and use advanced tools to create a working edifice.
The software industry's current practices have encouraged a one-size-fits-all mentality, and that just means that it's almost time for another SLDC paradigm shift to swing the pendulum back the other way to achieve balance. It will happen, be patient.
EDITED TO ADD: In fact getting LAPACK into the build if it's not already there is more of a maintenance imposition than just writing a couple of lines of code for this.
10 lines of code is easier to maintain than a whole dependency.
1 large, complex dependency (plus the 20 lines of wrapper code to transform your program's internal data to initialise and call the library) almost never just stays as 1 large, complex dependency (plus 20 lines of ...).
But I was really just wanting to show, via irony, that the structure of the original argument is self-annihilating.
That is, "X almost never stays as just X therefore X is bad" applies to both branches reasonably, and is therefore invalid as an argument in support of one branch.
That is the precise premise of the hypothetical.
(Of course, one should always reevaluate the situation when making changes to code, but we all need a nudge sometimes when we're afraid of doing the wrong thing.)
This is only true when you implement business logic, not when you implement algorithms.
A business wants to be able to change its logic to evolve how they do things. You keep trying to conflate these things but they are polar opposites.
2. You are wasting effort evaluating libraries. And I wonder what you lose in compile time?
3. You are calling into a black box, and you talk about opacity, a simple comment: /* this inverses the matrix */ will deal with your perceived issue.
4. Other programmers will assume there is a pertinent reason you included yet another dependency. Wonder what it does that 5 lines of code couldn't have done (surprise: nothing!).
I'm not sure that's right in this case. For inverting a matrix, it makes a huge difference whether the determinant is exactly zero or just very close, and (IIRC) floating point math may give the wrong answer. A library will have worked through this issue.
If there is some situation that is somewhat likely to come up and needs special code, but it is not immediately obvious (and, going by OP's evident lack of realizing it, that's some proof to say it is not immediately obvious)? It may not be there. And even if it is the common path, it can still be buggy.
Is is less likely that a library with some actual use out there has these bugs vs. handrolling it, but, if you handroll it, you're writing it for _your_ case, specifically with what you have in mind, whereas the library is more general than that. That's usually upside, but there's also downside, in what may be obvious to your usecase may not be obvious to the library.
The central point of the top comment here, that opinions differ and that there isn't actually an obvious answer even if you think there is one, holds true here.
This is a great example of the fallacy of believing that a library automatically solves arbitrary issues.
Not if they come in as integers that have to be coerced at some point.
>What is the library supposed to do?
Compare ad to bc instead of ad - bc to zero, for one.
>This is a great example of the fallacy of believing that a library automatically solves arbitrary issues.
I didn’t say that and don’t believe it. My point only depends on he library having avoided more domain-specific rookie mistakes than I would have caught in five minutes, when it comes to computational linear algebra. Like the sibling commenter, I don’t think libraries are a panacea, and the answer depends on the specifics of the case.
In which case you can calculate the determinant exactly with your own roll, but a library will likely force you to convert everything to doubles.
>Compare ad to bc instead of ad - bc to zero, for one.
That is the same thing, a compare instruction is basically just a subtract that doesn't write the output, so you get zero/equality in exactly the same cases. Most compilers will probably turn your code into doing just one subtract and then check the flags, no matter which version you write.
My approach to deciding if using an external library or write the code goes more or less as follows:
a) Limit yourself to a few big libraries that are widely used and well supported.
Especially if you work with a language where it is trivially easy to publish a new library, let's say javascript, you will find that even if you are just an average programmer half of the libraries you have access to have worse code that you would have written (so your point 1. is not necessarily correct)
Also using a library has a number of costs with it: increase in code size, time wasted updating it when needed, learning how to use it, understanding its documentation (if there is any at all), understanding why it doesn't behave how it is supposed to, checking github for bugs not yet fixed, etc. (so the effort spent in your 2. point is not necessarily bigger)
This cost is reduced is the library is a quality one, and also if you use a good amount of it, so my second point is
b) try to avoid a library if you just you need to use little of its functionality
Not only it may be cheaper to re-implement the part that you need, but adding a whole library means that you give charte blanc to other developers to use all the other parts of that library. This may be something that you want to avoid, expecially if the library has varying quality and not all parts are well-written. On the other side, if you know that you will probably need other parts in the future, it may be better to use the library instead.
My third rule is
c) keep in mind your team
If most people in your team already know a library, the cost of using it is greatly reduced. On the other side, if it new for most of them, it may be easier for them to maintain custom code (relates to your point 3.).
You misunderstood the question. The question is not about the general case. It is about solving a linear system of two equations and two variables. The solution is not found by "an algorithm", it is a simple formula that fits in one line.
Realistically speaking, though, if you're doing one 2x2 computation, you're probably going to end up doing a bunch of other linear algebra as well. If you don't do linear algebra in your everyday work, the threshold from rolling-your-own solution to grabbing a library is going to come up very quickly.
The Armstrong example is perhaps a bit mis-titled. He's talking about "implementing from scratch" only insofar as he can use the facilities of his platform to perform exactly the task he needs to do rather than bringing in a huge dependency, in this case, an FTP server that does a lot more than the one task at hand.
The question asked about inverting a 2x2 matrix, nothing about solving a linear system of equations. The latter is certainnly one common use of inverting a matrix, sure, but from the problem as given, the use-case could just as well be "report to the end user what the inverse is".
The invertibility issue also plagues any library and while it may not blow up by a division by zero we likely still would need to handle that case. This requires to understand the libraries handling of corner cases and how it affects the results.
You triggered a conversation thread that made me understand the ecosystem dynamics a bit better. Thank you.
2. If you're Joe Armstrong the amount of time you're wasting is negligible, and you don't need to worry about working with an unneeded dependency. I mean I would be wasting a significant amount of time because I'm not Joe Armstrong but there are things I get asked to do that other people would use a library for and I think what, why?
3. Opacity is in the eye of the beholder. If you're working on a project only you and a few people you know can handle it will ever look at, then you are that beholder. If you're working for a big organization as a consultant and your contract runs down in a month, think of the worst guy you have ever worked with that was still a programmer - that guy is the beholder.
4. If you write a comment then they won't make the assumption.
Later: Why the downvote?
Even Later: Upvoted, so I guess someone didn't like or maybe just clicked wrong.
But 99% of us are not at his level. And more importantly, 99% of the people who come after to maintain this code are not either.
I think there's a balance between rolling your own and using a lib, and it boils down to the domain of the problem being solved. If you're really just flipping a 2x2 matrix, the linear algebra lib will have way too much junk you don't need, and learning to use it would be more work than just solving a math problem. On the other hand, if you need a lot of common functionality (like multitouch events on the frontend), then why reinvent the wheel when there are a handful of great libraries for it.
1. Add an "invert2x2Matrix(theMatrix)" function.
2. Implement the function myself (no library).
3. Add a comment at the top of the function to the effect of "if this ever ends up doing much more than what the name says, consider replacing it with a library."
Do it by hand. No dense math involved. One line of code (or so, depending on coding style).
If you want to implement or call LINPACK to do a 2x2, you've got other problems.
I quite doubt that, though I used to feel the same.
The one feature I need in the library is probably implemented for a specific use case which is undocumented or misrepresented; and use cases outside of that are likely to be egregiously unsupported. Fixing the issues is almost always more work than just implementing X directly and looking up gotchas.
(aside: to understand this first-hand, go pick any JavaScript[0] animation library, and then go add a library that integrates it the core framework you're using (or not)... et voila, enjoy the descent into hell.)
A simple implementation of X that doesn't lock you into patterns you don't need or will one day harm you is worth its wait in gold. Seriously, I'd rather sacrifice DRY code than fuck myself eight months down the road by using a shitty over-engineered pattern that doesn't actually fit my needs.
> 2. You're wasting effort reinventing the wheel
Probably not; it turns out most of the wheel was invented and given to you in the language and runtime. For the article, it's 100% true, BEAM is pretty much an operating system and a great one at that.
We often abstract for the sake of abstraction; I honestly think Elixir/Erlang has helped break me of quite a bit of that. Rust is also nice because it'll make things so very painful for you when you abstract for no reason.
> 3. You risk adding unnecessary opacity by littering the program with dense math
Opaque libraries are the exact same thing but are even harder to understand since the source code doesn't comply with your same coding standards, methodologies, or test suite.
> 4. Other programmers will assume there is a pertinent reason you hand-rolled the algorithm instead of using the library. Now everytime someone touches that code they have to reverse engineer your approach to see what it does different from the library such that you had to roll it yourself (surprise: nothing!).
Documentation is the answer here, it doesn't have much to do with library or DIY.
Why go through the trouble to properly evaluate a solution, decide the best solution was to implement your specific use case (or otherwise), and then not document it? You're just asking to repeat yourself or forget.
On the flip-side, its really easy to assume that library authors did something brilliant when in fact they've done the opposite. Or someone will assume you've chosen this library because you took the time to properly evaluate it and never question it...
Ex.
After eight months of having the library run slow and leak memory some junior engineer pulls you aside and softly says "I have a minor in mathematics, I noticed that library Y is only doing simple matrix multiplication... I can do the same thing with this [shows you four lines of code] and its a lot faster... I just assumed we had to use that library since its in our code base."
Adding libraries before you need them causes more headaches than not because now you've littered a dependency throughout your code which may not be sufficient for future use cases. This goes back to I'd rather have less DRY code than be bound to a shitty inflexible pattern.
> The two exceptions I imagine are 1) if you really do absolutely need some slightly modified microoptimization of what the library does or 2) if you are working for a platform where there is good reason to avoid even slight bloat (embedded).
I don't totally agree with either of those statements fully but I don't strongly disagree with them either. I'd like to add for me these would be of much high significance:
1. Your future use case are unknown.
2. You do not know if the library will be maintained; so, by using it, you're going to own it anyway. There are no free lunches.
3. You're using only a small, defined subset of some standard (i.e. RFC or specification). The existing solutions provide too much cruft for your specific use case.
4. The solution is obvious and has one correct answer. In some languages this is easier because of strong conventions and standard libraries. In others, this is probably a pipe dream.
5. The libraries are not tested adequately or are not easily testable.
I dunno, over the past few years especially, I've just started doing more things myself and yeah things do take slightly longer to start (sometimes...) they almost never take more time to maintain, expand on, or use.
Each line of code used by your project is a liability (debt), not an asset, on your balance sheet... Less code, less liability, less debt.
[0] No offense to JavaScript <3, I only choose it, because it's so ubiquitous.
I think the show Silicon Valley made two perfect archetypes of this divide:
Gilfoyle: the proudly independent coder who prefers to build everything himself rather than be subject to trusting other's code. Documentation is minimal. Code is intense, efficient, and requires a (too) high level of intelligence and effort just to make sense of. Gravitates towards back-end bare-metal systems code and hardware. Background is self-taught pragmatic hacker. Refuses management roles, preferring to do a department's work himself rather than trust underlings.
"Why would I use a library to perform elliptic curve cryptographic functions I am fully capable of writing myself? Do you think I'm a fucking moron?"
Dinesh: the extremely needy but clever programmer whose code is a massive pile of dependencies and abstractions. Unhindered by pride and happy to trust any workaround or outside source if it means him doing less "work" (he does not view learning about new libraries/codebases as work). Documentation is conversational and verbose, with huge doc blocks above the simplest functions. Code is highly abstracted and spread across many files but with an overall working logic to it - often following various (obscure) engineering standards rather than using personal pragmatism. Loves to communicate, and so has a flair for front-end-facing code and making things pretty and understandable rather than completely efficient. Background is more academic and group project based, with a deep understanding of standards and love of new technologies. In a management role, tries to befriend, integrate, and understand all his employees programming styles - and sends many memos with emoticons.
"Oooh, look, I found this neat library that animates my system font whenever I press the 'r' key! Arrrrrr! Hah. Pirates..."
(May our industry RIP)
def invert_2_by_2_matrix(matrix):
# matrix is of form, for example ((1,2), (3,4))
((a, b), (c, d)) = matrix
det = (a*d - b*c)
return ((d/det, -b/det), (-c/det, a/det))
(This doesn't do any input checking including if the determinant is 0)Which seems to me like it assumes away a big part of the difficulty -- you'd need to follow up with: for this problem domain, what am I doing with the answer? How do I need to handle a singular matrix? Should do I throw an exception, and how should the caller handle it?
import numpy as np
np.linalg.inv([[1,2],[3,4]])
and this does handle the case where the determinant is 0 (by throwing an error about it being a singular matrix).One kind of programmer will look up what it is and how the algorithm to invert it works, then implement that.
The other kind of programmer will search for libraries implementing these features.
The end result is that the first kind of programmer will know how a lot of things work, while the second kind of programmer will know a lot of libraries.
But this is really another version of the same question: When faced with a task you don't understand, do you Google "How to do X", or do you search "library to do X in Y language".
In particular, I wonder whether ill-conditioned edge cases (https://en.wikipedia.org/wiki/Condition_number) might need special-casing.
Having said that, I’m not convinced I could easily judge the quality of a third party solution, find a library that favors accuracy over speed, if I needed one, or whether existing libraries properly document how much overhead calling a routine taking nxn matrices for 2x2 matrices takes, if I needed a fast implementation.
I can't think of a similar scissor the majority of programmers both already have a solid understanding of and would be likely to be divided on - but I'd love to be out thought!
Unless they're not really an engineer. Engineering is pragmatic, not dogmatic. Your job is to do the good thing for your system, not to apply first principles to everything. Approaching it from a position of dogma is not engineering.
IMO, the best way to write the code (pulling in a lib vs doing it yourself) is to consider the context in which it’s going to be read, maintained, and used. If that’s by a bunch of applied scientists or linalg experts who can read a formula more easily than debugging and maintaining integrated libraries, then roll it yourself.
In the general case look at complexity of the options and then look at context of who has to maintain that complexity don't just assume 3rd party means it'll be easier for others to grok while bug hunting.
... but I remember it being super simple and I’d probably be in the “just write it” camp.
Schools don't teach formulas. That's the misunderstanding of a student that only sees the next exam. Schools teach you where to look when you don't have the answer.
And a good programmer should be able to learn a new domain if they want to excel in the craft, otherwise we’re back to software factories and code monkeys.
If you wanted to use equations in your code and I was reviewing it, you're going to have a LONG session explaining the finer points of floating point and how you avoided all the problems that can cause.
I can quote sections of IEEE-754 from memory, and I would use a library. I simply have zero confidence that my algebraic equations would handle massive cancellation and roundoff correctly in all cases.
And these are the WORST class of bugs to introduce into your program--a function that works right 99.9% of the time but fails catastrophically at random times.
"I might have to change it to use a THREE by THREE matrix!!"
So, how sure are you that the formulas you used handle massive cancellation and roundoff correctly?
In fact even if LAPACK does all linear algebra operations correctly 100% of the time, it also introduces a new paradigm, meaning I have to learn how to do things the LAPACK way which may not be intuitive for the class of problem I have. Or LAPACK may do things correctly for the general case at the cost of CPU time whereas my problem needs to be calculated in real-time, in which case including LAPACK is a net-negative for me.
I think the two schools of programmer divide pretty cleanly into people who assume infinite resources and no CPU time constraints and people who have to make trade-offs between resources and time constraints. Serious embedded programmers will almost never want to just include LAPACK and be done with it unless their problem domain involves a ton of linear algebra. Whereas I've yet to meet a web developer who wouldn't just pull in 400 megabytes of libraries doing god knows what to invert a 2x2 matrix with a function call. There are web developers out there who have to optimize pretty close to what the system is capable of, but most web developers don't really care that much about that because they have infinite computing resources(just spin up more instances!) but they need to release their website "like, yesterday."
So it's not about camps per se. As is everything in engineering, it's about trading a resource you have in abundance for a resource that is scarce. For embedded programmers, the real-time requirements and storage capabilities of the MCU are king and they'll gladly spend all the time in the world rolling their own 2x2 matrix inverter. And for web developers, their development time is king so it's better to just cobble together something working using LAPACK and move on because they can just reserve a bigger, faster server.
I like to implement everything from scratch (compilers, TCP/IP/arp/rarp stack, databases). But I like to use recommended libraries, not my own code. By implementing everything myself, I have a ready list of questions about how certain problems were handled, and a firmer understanding of where to look when the software misbehaves.
I’d give good odds that the person doing a 2x2 matrix didn’t even notice it was linear algebra. I wrote a trivial image rotation, zoom and crop feature a number of years ago and I bet if you looked at that code you’d tell me I fucked up by not using linear algebra. Instead I had a series of geometry diagrams on my whiteboard trying to figure out unit tests for my hand written code. Which the next guy promptly broke but that’s another story for a different day.
If I know it’s linear algebra, it will probably come up again. Now my decision is somewhat informed and also more difficult.
For example, for a one of script deployed in Pythonista on iOS, I would probably use a library if It is available/installable via pip with a single command, rolling out my own otherwise for 2x2 cases with a single google search unless I need to handle special cases.
Don't most people do this?
I think for most of us can make this mistake, and reasoning in higher order (i.e. using a function from a tested library) helps avoiding this kind of error.
But things can depend on stuff not mentioned in your questions.
The larger and more complex a system you design, the more likely it is that you end up with something breaking because two dependencies are incompatible or even your code and a dependant break due to updating the dependant.
Further, security is rarely considered. The more code you have, the more likely you are to have a security vulnerability. You have to worry not only about code you maintain, but also the code of your dependencies.
If you can write the functionality you crave rapidly, do it. You'll save yourself (and others) headaches in the future.
Moreover, it's getting quite easy to maintain dependencies today. I can think about packages lock files, Docker containers, semantic versionning, or even Github sending me alerts and pull requests when a security vulnerability happens.
I prefer to add a good dependency and having to deal with it rather than writing broken code and fixing it later. Of course I'm not talking about a 5 lines dependency, but real features.
Illustrative anecdote: this Friday I copy-pasted a little snippet of HTML/CSS code from an existing (and shelved) project. I've noticed for some time an icon in it wasn't displaying anymore, but this time I really needed to fix it. Turned out that the glyph icon where from a font used by Bootstrap, which was not available freely anymore. Bootstrap (my direct dependency) had itself that third party dependency that I thought was part of it. When they pulled the plug, my code broke.
I'm more careful about dependencies than I was a few years before. Often, if a peripheral thing can be done in an "ugly" way but allows me to avoid an external dependency, I'll wrote code myself.
“My point today is that, if we wish to count lines of code, we should not regard them as 'lines produced' but as 'lines spent': the current conventional wisdom is so foolish as to book that count on the wrong side of the ledger." - Djikstra
In an age of public bug databases and Stack Overflow, it’s pure ego trip to think that your handwritten solution is going to be better than an OSS solution or stay that way. If you reached for hand-written once, you’ll probably reach for it again and again and the older bits you wrote will be eclipsed by progress on those libraries. I’ve seen it many times before, I’m living it right now.
It’s vendor lock-in with you as the beneficiary. A sports coach would call you a Ball Hog. Let other people play.
You can run a perfectly respectable project with about 10% “special sauce”. That includes the bits that simply nobody has, and the exceptions to my advice above. You can sustain a small number of hand written replacements on par with existing libraries. And the team can absorb the Bus Number effects of that much code.
And “about” means “about”. If you are less, a competitor might steal your lunch. If you’re higher, you should be reflecting on why and asking if you really do still need all this stuff. Maybe that library you looked at before finally got the feature you needed.
I think that's as much true of external dependencies as it is of ones that you wrote for yourself. With the exception that your own solution is probably smaller, much smaller, because it only ever needed to solve your immediate needs and not everything else a third party had to deal with.
So it's quite possible that the internal solution is more approachable when the time comes to debug it.
Personally I’d rather be mad at someone I’ll never meet. It makes meetings a lot more pleasant.
Let's take the example of someone who wrote their own Perl ORM and Object Model then.
> You'll save yourself (and others) headaches in the future.
Or it'll hamper development on the system when you leave because nothing works how people expect, it's too complex for people to pick up and modify, and trying to do anything with it causes the entire structure to creak because it's EVERYWHERE.
He evaluated 3 options, then downloaded, compiled, configured and integrated the best choice into the project over the course of a full day. Its limited database integration configuration necessitated some odd quirks like duplicating a db column, but oh well.
When finally done, integration tests revealed that it didn't actually solve the entire problem. It only solved 90% of the problem. Not only that, it had a significant but elusive bug that only occurred in 5% of situations. He filed a bug report, but the project hadn't received any updates in over a year. Damn.
When it became clear neither of the other two readymade libraries would solve 100% of the problem either, the programmer wrote his own solution. It took 3 days of work but he learned a LOT about exactly how it works, and the result solved 100% of the problem with no major bugs.
All of the quirks disappeared. The solution fit the problem exactly, and remained flexible and maintainable. It eventually served as a solid foundation for future developments which had no readymade solution.
We had our own snowflake wrapper around the saucelabs client. I eventually figured out how to file a PR to do the same. Now I get to use that solution on every other project I ever work on. I have similarly used my own Stack Overflow answer to solve a problem I also had three years prior. And I recently found my SO answer to another question verbatim in our own codebase. Ironically, from a guy who often ignores my advice.
Get as many things out of your head and into the group memory as you can. Things get forgotten otherwise. Speaking of, did you guys ever consider open sourcing his solution?
When I bring in a third-party solution for something that might be feasible for me to implement myself, it must 1) be actively maintained, 2) be decoupled from my code at the function-interface boundary, and 3) fully cover my use-case, at least as it's known at the time of introduction.
In the ecosystems I work with, finding all of these qualities is not uncommon.
So now we've got some simple code to copy files between machines, because we hired someone who was too academic to figure out how to set up a simple FTP server. Here's what comes next:
1. We need to open firewall ports so these two machines can communicate. What are they? Who knows. If it were FTP, the sysadmin would know.
2. We need authentication; we can't just allow any old machines to connect. If we were using FTP, we'd have that built-in. And it would have been audited, so we can have at least baseline confidence in its safety and accuracy.
3. We can't go having the disk overflow. We need something to manage quotas, because even though we trust our users, sometimes software goes rogue. Maybe the default FTP server can't handle this, but some other product can.
4. We need logging so there's an audit trail, in case something goes wrong and we need to track it down.
And on and on and hey look you've reinvented FTP and now no one else on your team knows how it works, it's full of bugs, you've hardcoded a bunch of inefficiencies, there's no documentation, we have to explain what it is and how it works to the operations team, and hey here comes the boss asking us to make sure it works over SSL.
If you think you can do it better, make sure you think about what you're doing.
You came up with a bunch of things to complicate the scenario that do not apply to the author's problem, and those things are precisely the kind of things that make a general solution more complex.
Here, by writing their own tool, Joe could avoid all the complexity that does not apply to his scenario.
I don't think the point of the example was to demonstrate a complete and general solution to the problem of transferring files across computers, but a specific solution that solves the immediate problem in the specific environment Joe was operating in, faster than you'd get it done by trying to apply a more complex (and more general) solution.
Another point I'd like to make (that the original post doesn't) is that they could easily extend this custom solution with additional logic that doesn't exist in any FTP server off the shelf. In fact I've had to reject libraries and programs on similar grounds; they almost kinda do what we need, but then a vital piece of the puzzle is missing and forking the project & implementing our special sauce (and then maintaining it) is going to be much more painful than rolling a custom solution from scratch.
He didn't need an ftp server, he just needed to manually copy some files. Is it really wrong to do that without setting up an FTP server with authentication, disk quotas, logging etc?
- guess whether or not this is an isolated case, or whether or not this will become core functionality
- a self-assessment of the true difficulty of the problem
- a self-assessment of their own skills and knowledge in the area
- security reasoning
- API access/readabilty for other developers to use this code
- maintainability of new code
I have personally seen personal implementations that lead to bug, after bug, that have already been reasoned about in equivalent libraries.
Often for the simple fact that other devs who have to work on this code, its likely that the abstraction and readability of a third party library is probably greater than the 'quick-and-dirty' implementation.
The person who reaches for the “New File” button is often not used to asking if the thing they need already exists. They can end up duplicating or triplicating business logic which results in other engineers confidently declaring that a problem has been fixed when they only fixed the obvious occurrence.
Heavily agree. Shameless self plug, this is the philosophy behind Distilled[0]: provide a good-enough, low-abstraction tool so you can build on top of it without it getting in your way or forcing you to memorize configuration options.
Adding unnecessary black-boxes to your code can occasionally come back to bite you. I try not to be religious about this, but I am a little biased towards avoiding the abstractions of the 3rd-party libraries unless they come with a lot of benefits. A big "aha" moment for me was trying to teach interns how to use Grunt/Gulp, and eventually realizing that just teaching them Bash was easier. Then I started noticing other stuff -- for example, that the documentation for our inline documentation library was longer than some of the documentation pages we were generating with it.
There's a principle here that's true for abstractions in general[1], but I find it is especially true for some 3rd-party dependencies. I regularly need to read the source code of 3rd-party dependencies that I use, so I no longer treat them as free.
Obviously this stuff isn't a hard rule. I still use 3rd-party dependencies. But there's a balance here. You should be at least a tiny bit cautious about ready-made solutions.
[1]: https://peertube.danshumway.com/videos/watch/2daaee22-4d92-4...
Instead of downloading node.js application with numerous dependencies, I also prefer implementing it in few lines of code.
Some can claim running a 'proper' application (like nodejs/express with JWT authentication) is better. But if your business not solely depend on your solution, trivial stuff wins on the long run IMHO.
Of course security/access-control is debatable...
People who are salty about this subject have dealt with NIH situations where people think they can do better, and are either wrong immediately, or become so when their attention shifts to some other problem.
I have written and maintained internal libraries. Many times it’s the wrong choice, sometimes it’s the only choice. But I’m also more responsive to feedback than the NIH bozos my coworkers typically complain about, which may in part be why I get to hear about their grievances.
There’s an old half-joke, half-theory that we should elect a President who doesn’t want the job, instead of people who do.
You should have the individual contributor who doesn’t want snowflake code write the snowflake code. They’ll keep it no-nonsense.
I did some initial tests in Node. When I installed express, for routing, it added tens of dependencies and 2.5MB of code files. They call it "minimal".
No thanks. I'll skip the router, for now or write my own.
However, there are 10 dependencies, some of which are non-trivial. You spend 4 hours changing versions of libraries, downloading a newer compiler than the one you are using, and never get the thing running.
You could have slapped some code together to read the file, and then draw it yourself. It would have taken a half hour.
However, later on you are tasked with with representing many more objects, which have corner cases you didn't account for. This eats up hours and hours of development. Also, you're then tasked with drawing those objects in many different ways, more and more development has been added.
The two take aways are choose based on what you have to do, and make that choice based on knowing 100% what you have to do. If a customer can't make up their mind, they're hurting development whether they know it or not.
BTW, the objects were GIS shapefiles.
The easy solution was to just use Conda.
The code and steps needed to compile/bundle/start/maintain for many other languages would be far more than in erlang or other beauties.
Where are the tests? Even in erlang, unit tests are useful. Why would I trust my code? What about security? You might not want this server to allow nasty things to happen if unintended use takes place.
So even in this particular case, I think it is worth evaluating opensource solutions, their maintenance, their license, and have a proper server in place that will be far easier to maintain and trust its robustness (security and availability)
> With enough eyes, all bugs become shallow.
It's that learning curve. . .
I think his point is more it's faster to solve the problem you have ("I want to get files from machine X") rather than the problem you imagine you have ("I need an FTP server") - it's just wrapped up in "I solved it myself using Erlang".
But I very much agree that the speed of the custom solution can in large part be attributed to the fact that you can reject the cost and complexity of existing solutions like FTP (which set out to solve more than just Joe's immediate problem).
Similarly, I often transfer files by piping them to netcat.
EDIT: the point about tools still stands. For example:
host1$ nc -l 1234 > my.file
host2$ cat my.file | nc host1 1234
This is very easy to do given that unix is my toolbox. It'd be much more involved to write from scratch in (say) C.On the other hand, Joe's tool is much more powerful; it doesn't require him to manually interact with the shell on two different hosts, and his program can list files on the remote host and automatically download them under the correct name whereas I'm stuck manually redirecting nc's output to the right file (and cursing myself when I accidentally forget to change the name and overwrite a file I wasn't supposed to fetch). Those features were very easy to add to his program because Erlang is a powerful tool. On unix shell, it's much harder.
The bad ones get us stuck with that solution for a very long time. These are the experiences that prompt people to post to threads like this.