A significant amount of programming is done by superstition
utcc.utoronto.ca
utcc.utoronto.ca
The real gold are these words:
> You've really read the LSB specification for init scripts and your distribution's distro-specific documentation? If so, you're almost certainly a Debian Developer or the equivalent specialist for other distributions.
Many of us would agree, that it's true. But why is it true? Shouldn't that make us suspicious in the first place? It's almost paradoxical, even funny: making the work done is easier, than finding out, how to do that work. You see, init scripts or makefiles aren't exactly rocket science. I mean, things we are willing to achieve with them are pretty simple: just start some process, just build some project (which means "execute bunch of shell commands in the right order", and not implementing out own superoptimizing compiler).
So when we arrive to situation like one we've got today, the right message wouldn't be "come on guys, you see what we're doing because of being lazy? we should stop being so lazy!". That is, yeah, we all are lazy and maybe we should stop it, but that's not the point. What this situation actually means is that our tools for doing simple stuff are overcomplicated shit, and our documentation for these tools is shit. If anything, it's pretty obvious that a couple of lines of code, showing how to do the work is better documentation than complete list of possible flags and parameters with no examples whatsoever. So no wonder people are reading stackoverflow instead of man-pages. And so on.
the best manpage is a man page with a nice EXAMPLES section you can grep to quickly. though the permutations offered by many programs prohibit all bases being covered.
On top of that, to be able to use man page, I need to know command in the first place, while google will digest my broke english question and point me to the right stackoverflow answer most of the time.
This is an anecdote sure, but an important one. Not everyone needs to understand what's going on with everything but I can tell you the ones who seek out the information and really try to grok the systems are rare and can be well compensated for it.
Another anecdote: trying to use ZFS, and in particular de-dupe, without understanding it. De-dupe has its own trail of tears laid out from people who did not understand the memory requirements of it.
Hello, I'm Rapzid. I'm 30 years old, and I'm a man page reader. Sue me.
Progress depends on humans being lazy. We have inventions and innovations precisely because humans are so good at finding ways to make their (own) lives easier - usually by eliminating repetitive tasks via some kind of work-around or alternate approach.
There is a well known quote about insanity [0] and blind repetition that occasionally gets associated with idiocy too. If everyone was always starting from first principles, we could argue that humanity, as aggregate, is insane. (Which might still be true!)
Repeating experiments to verify results is good. Doing the same things over and over again likely means you're not learning.
0: http://en.wikiquote.org/wiki/Insanity (search for "different results")
OTOH, rest shouldn't be confused with laziness, which has a negative connotation, because some see idleness as an ill when it is assigned the highest priority before other desires/needs.
If you work in Java or Ruby or Python or C#, or even C++ all day with say, an IoC container for argument's sake, you're certainly not constantly thinking "okay this interface will be resolved to this type, which has these dependencies which will be resovled to these types, which generate this byte code, which generates this JITed assembly code, which generates this machine code, which generates these executions of instructions on the processor and accesses these memories, which invokes these manipulations of these gates and flip flops on this particular part of the ALU".
Doing so would be the mark of an unproductive madman.
The issue is that we all have our own (more or less completely undocumented) abstractions that we mentally put on top of whatever language/library we're using. And sometimes these mental abstractions are out of sync with reality.
I think a better title would be "A significant amount of (shipped, so-called production-quality) code is not fully understood by its authors".
Of course. Time spent fully understanding the code before it's shipped is time that it's not shipped.
This is sort of a Gresham's-Law situation, where "bad code drives out good code". The code that is fully tested & debugged with all situations understood isn't available to the user at first. As a result, they go with the "good enough" solution, build stuff off of it, and by the time "the right thing" comes along, all the momentum has shifted to its crappy competitor. A bunch of crappy software gets built on top of that, and as a result, we end users get the worst thing that could possibly work.
But work it does. Maybe it's worth a shift in perspective to be grateful for the stuff that works rather than worry about the stuff that doesn't quite work right.
Good Enough is our main design goal. Anything more than good enough means you're wasting resources.
For instance, you probably need to take a bit more care with defining your core database structure, than you do with positioning a button 20px to the left or right.
not when you are trying to prototype the UX - and don't care about the backend. There's no rule of thumb - it's all very subjective and intuition based. That's where experience comes in, and no amount of book studying will help you.
That's a nice quip but once you cross spans of 5 meters or so you'll find out that that is a lot harder than it seems, especially for non-trivial loads.
The joke of course refers to the fact that to build any structure that has to be both safe and economical is hard but please don't make it seem as if building bridges is easy, it's anything but.
But yes, building bridges is hard. So is building software. But the job of an engineer is solving a hard task within budget and within deadline.
Multiply this times 100 and you are just about in the same situation as web developers. Maybe in 100 years, an ecosystem will grow around to support us? I can only hope. For now, i cant even rely on xml to be marshalled the same unless i control both client and server.
ME has an advantage over CSE in this regard because in CS the requirements are usually very poorly stated. About the best and most rigorous requirements you get are from mathematicans, but most devs and managers now certainly view that as wasting resources.
To make not to fine a point of this, I always like to turn to Alan Kay's comparison: TCP/IP vs the web. When was the last time any human technology like TCP/IP scaled so well and so invisibly, it's like air we breathe without even thinking about it. The web in contrast was the work of rank amateurs.
Alas, mathematics is really the "good enough" standard we in CS should strive for just like physics is the "good enough" standard behind ME and EE. Unfortunately as CS opened to the mainstream, I think a deep fear of mathematics led us to view this as "over engineering" even when it wasn't. The results are that the majority of the web is woefully underengineered, requiring far more money and time for inferior products.
We know they are inferior, because even the simplest gui application has more consistency than a web variation of it. And that's what marketing constantly compares things too when they can't understand why the web sucks as much as it does.
"Good Enough"? Please! For the last 20 years we haven't even come close!
http://www.youtube.com/watch?v=oCZMoY3q2uM
Packet switching was "utter heresy" (20m in) when it was invented. It "wasn't a network, it was an inefficient way to use an existing network". And it almost collapsed 25 years after it was invented; the presenter is famous for inventing the modern TCP/IP congestion control that saved the Internet [1]. The TCP/IP flow congestion algorithm has been redesigned several times since [2]. It works not because it was designed well, but because it wasn't designed and instead evolved over many years with many contributions from people devoted to keeping it working.
Alan Kay is usually who I think of for great ideas that "work in theory, but not in practice". He's done some crucially important work in OOP, programming languages, and GUIs. But note that we don't actually use SmallTalk; instead we got C++ and Java. Nor do we use Dynabooks and Altos; instead, we got Microsoft Windows.
[1] http://en.wikipedia.org/wiki/Van_Jacobson#Career
[2] http://en.wikipedia.org/wiki/TCP_congestion-avoidance_algori...
Also, for those unfamiliar with the actual process of protocol development back in those days, it was a wire protocol, which meant formal modeling and testing. Sure, it doesnt catch everything, but the web is far less formal. For example, the w3c originally said it wasn't going to provide an XML parser reference implementation because any graduate student should be able to code it up in two weeks. WTH?! While I don't doubt that is true, in practice it has meant that dozens of slightly different parsers were written, leading to hundreds of slightly different incompatibilities. Anyone who has had to integrate two different XML stacks will know.
I use Ruby, which was inspired by smalltalk; modern Java is also becoming much more functional.
In some ways, it has taken the larger community 20 years to understand Kay's vision. Also, he always said that the systems he worked on were prototypes -- he's commented before that he fully expected real-world systems to have surpased his long ago.
But now we have Ruby, Node and Rust. Even Java and Spring.io have dramatically reshaped things towards a "smalltalkish" future. So I still put a lot of weight behind some of Kay's observations of the industry.
Devs are trying to fix these ecosystems: why does react use a virtual dom? Why do we need css resets? Why do we need js shims and polyfills? Because its the only way to come close to normalizing the platform.
But have you ever wondered why you expect no two browsers display the same image? Postscript met that bar and is just as old as the web. Why didn't the w3c base the web on device independent coordinates instead of this confusing and unpredicatble layering of partial scalars and "angle subtended by a pixel on a 96dpi surface at a nominal arms length from the surface" crap? No one could have made a reference implementation off those requirements, much less a consistent verification & validation suite.
And no offense to TBL, but HTTP didn't even survive first contact with netscape's vision of shopping carts. Cookies? An elegant solution? Or simple a new hell of tunneling client/server state over a supposedly stateless protocol. HTTPS everywhere requires long lived sessions as the basis?!? No wonder people are heading towards web sockets, etc. webapps are client/server apps -- HTTP was always grossly misapplied to them.
Webdev is hard, not because I'm building beatiful bridges in the sky that are "good enough" poetic balances of constraints while coming in on time and on budget... Webdev is hard because of all the underlying assumptions I constantly have to check and recheck because I can't rely on them as an ME would (or hell, even as a backend J2EE engineer would). This is why some of us lament that people don't know the stack all the way down, because we have to in order to solve real problems. Every abstraction leaks, but hell, web abstractions are flipping sieves!
No, the thing that "works incredibly well" is not the web, but whats under it that lets us make so very many mistakes and yet keep on trucking.
PostScript does not even attempt to solve this problem, so would never work on the web. Unless you mandate that everyone should have screens with the same dimensions and dpi.
Besides, windows faces a similar problem of multiple resolutions and devices. How do they v&v? They set the resolutions the same for certain tests! Even if you do this for browsers, they cant pass the test. Yes, it would be nice if we could have device independent layout constraints as well, but even the simplest most constrained test not involving layouts fails. At least now, its close. Before ACID it wasn't even close.
If webdevs can't even rely on their browser coordinate system in the most heavily constrained case, how can they hope to trust it when they try to solve challenging problems of dynamic layouts across multiple resolutions?
The fact that it was created by amateurs helped make it work incredibly well for amateurs.
The professional CS alternatives - gopher and the like - were not successful in comparison.
There's a thing in CS where solutions become so clever they become stupid - because the goal stops being task-oriented usefulness, and becomes ideological and formal purity.
It's the process that turns a plain hammer into an atomic pile driver you can only control remotely from the moon by sending it messages using catapulted owls in space suits. It's better at hammering in some abstract sense, but maybe not so much for hitting nails.
Abstraction without contextual insight is one of the most powerful and destructive of all anti-patterns.
On the web the professionals took over from the amateurs, and now web technology is another example of design-by-committee.
It still works surprisingly well because interplanetary owls are kind of fun, maybe, for some people. But is it ever a mess of half-solved problems generating recursive epicycles of complication.
If I remember the context of the "created by amateurs" quote, it was really Alan Kay complaining that the web wasn't designed by OO principles. He wanted the web to consist of objects encapsulating their own presentation logic, rater than document in declarative languages. So basically something like Java applets instead of web pages.
While OO is great for software design, I believe declarative documents have proven to be much better as a foundation for a decentralized information system. Think about how to implement Google, accessibility, readability.com and so on in a web of encapsulated objects. And it is not by accident that TBL chose declarative languages over objects, he actually though about it: http://www.w3.org/2001/tag/doc/leastPower-2006-01-23.html
This is an example of the contextual insight you talk about, and which I believe Kay lacks in this case.
EDIT: The interview is here: http://www.drdobbs.com/article/print?articleId=240003442&sit... It is not totally clear what he is arguing, but is seems he is suggesting that the only job of the browser should be to execute arbitrary code safely, but any actual features beyond this should be provided by the objects. So the browser should really be a VM or a mini-operating system executing object code in a safe sandbox. This seem to be the philosophically opposite of TBL's principle of least power.
Honestly, it seems like Kay is ranting a lot in the interview. When something like the web is not designed the way he would have done it, the only reason he can imagine is that the designers must have been ignorant amateurs.
Just because TBL is brilliant doesnt mean his work can be misapplied. Of course, i also blame the people who thought of scaling thousands of existing client server applications for a fraction of the cost: things like shopping carts and online banking. True, it drove the web to what it is today, but at great cost.
Here is another thought: if the web is so great, why are so many companies creating their own tablet/mobile app experience instead? It cant be because it requires less dev knowledge and effort?
As I have stated multiple times in the past, I think the crux of the problem is that software is an immature field that needs to stratify into a proper engineering discipline as it matures. Computer science should be the "good enough" standard behind software engineering. "Computer scientists" should not be the ones actually implementing software systems any more than physicists should be the ones designing cam shafts or laying out circuit boards.
The opening to the mainstream you refer to illustrates the problem. The people in the mainstream should not be studying computer science, and what they practice should not be called such. They are the engineers, technicians, and mechanics of software; they are not the physicists. Forcing all of these strata into the same bucket is doing more harm than good at this point and is likely hampering the field's drive to mature.
It's 2015, there's no reason that anyone should be writing Big Object Oriented Code, without practicing dependency injection, basic mocking and testing, and other modern development principles. And yet, here we are, with millions to billions of lines of terrible new code written every year.
Well duh. Not a single program shipped today, no matter how small, is "fully" understood by its authors, because to do so, one would have to understand every implementation nuance and bug of OS layers, libraries used, other parts it interacts with, ... And even for less strict definitions of "fully" - for any program that took longer than say 2 weeks to write, how can anyone claim they really know the details of everything, at the same time? Including the nuanced differences between various of OS calls, and the effect some environment variable might have on your localization code, just to name one thing?
I agree with you that a significant amount of code is not fully understood by its authors. I see similar phenomena at work in many engineering projects. Relatively few engineers work at the component or board level, rather they work at the subsystem and system levels. They piece parts of systems together to make larger systems-of-systems (sound familiar, programmers?) often with a poor understanding of when the black-box abstractions they deal with can break down and cause havoc.
Currently fighting against this in my own workplace, where it's become clear that many engineers do not understand the system we support in sufficient detail to properly troubleshoot faults with that system.
You might look at the code occasionally, and think "maybe I can clean this up?". However ,after a few minutes, you realise how much you don't understand what's going on. So you leave the code, in a sort of superstition - it works after all.
After a few years, when you need to finally modify the code, you've got a problem. The code works, but you never took the time to understand why it works. If there were any superfluous calls, or side-effects that you don't want now, you'll need to take the time to understand how it all works. If you took the time to understand it originally, and wrote a few informative comments, you might have saved yourself hours of work.
It must be said though, occasionally you do need to go deeper with your thinking. To figure out why some bug is occuring, or maybe some complex performance issue. So it is important to have some knowledge of how it all fits together, even if you don't need it all the time.
Yes, quite often that came without a perfect knowledge of every single option, certainly usually without referring to the manual for every one of them.
Perhaps this is semantics. But the definition of superstition is:
a widely held but unjustified belief in supernatural causation leading to certain consequences of an action or event, or a practice based on such a belief
Which ya, isn't quite the right thing. Anyways, agree with the sentiment of this 100%, just not the title. :)
We generally have some idea of what's in the black boxes we use, and we are always open to change our beliefs based on new evidence ("faith" is a belief that will stay the same no matter what evidence is presented: in Bayesian terms it is a belief with a prior of 0 or 1, neither of which can be changed by any evidence.)
"Hopeful" or "optimistic" programming might be better.
Do I trust my OS to have a memory allocator that works and to protect processes from reading and writing each other's memory? Yes, but not totally.
Do I trust its more modern and complex APIs? Less so.
Do I trust code I copied from stackoverflow that has 100+ points? Yes, but even less so than I trust the complex or new corners of my OS. So, I read every line I copy. On the other hand, I am lazy, so if this gets to 1000+ lines, chances are I won't read every line. I will think a bit more about the trust issue, though.
Do I trust the documentation of the libraries shipping with my OS? More than stack overflow answers with 20-ish points, but certainly not completely.
Do I make assumptions about APIs? I try not to, but it is hard. There is lots of documentation where it is hard or impossible to find out what the code claims to do in edge conditions.
For example, in https://msdn.microsoft.com/en-us/library/352y4sff(v=vs.110)...., can one pass a null transaction? I wouldn't know, and the page doesn't tell me, so this afternoon, I wrote a conditional operator calling another constructor if transaction is null.
I did consider a "let's try it and assume it always works if it works once" approach, though. I also think many people do that ('if it compiles and runs, it must be valid C')
Superstition connotes that an act is performed due to faith that it is beneficial or necessary, rather than an understanding of the underlying mechanism or controlled experiments to demonstrate the need.
I'm not sure what the details are here besides init script, but in this particular case the job is boring and difficult/slow to test. Who wants to reboot their machine all the time to test it?
... Which is why sysv init has taken so long to fix. If it's not obviously broken it won't get fixed.
I didn't initially manage to follow the right sequence of links to find the example[1], but now that I have, the context seems to be that "System V init ignores all of these" (the parts people copied without understanding), so the description still looks applicable to that example. (They serve no purpose at the point in time when they're added by cargo-cultists but cause problems much later when someone/something assumes that they are actually meaningful.)
[1] http://utcc.utoronto.ca/~cks/space/blog/linux/SystemdAndSysV...
good day.
Silly as it may be with things like git around, I can recall thinking, "maybe we will run this on AIX..." when hacking on a multi-OS shell script.
Not good practice to leave the bloat in, but at least give developers the benefit of the doubt that it isn't outright magic to them.
In Pragmatic Programmer the authors talk about "programming by coincidence", I think this would be closer to GP example.
Do I copy/paste swaths of code that I don't understand and put them in my programs? No, absolutely not, because each line is one I can reason about and want to understand.
But do I copy/paste swaths of configuration files / init scripts / etc, without understanding the implications of every single configuration? Yes, absolutely, because the depth behind each of those options is often irrelevant to me. (until I discover otherwise late at night when an alarm goes off (!!))
I think it is important to distinguish these. I don't think any dev that's done ops work would claim they aren't guilty of the latter, but at the same time I think very few good devs ever, and I mean ever, do the first.
There's a configuration language a Google that's stated goal is that it's not turing complete... and yet it is, because it was a necessity for achieving the expressiveness, and now it's almost entirely unreadable because of the convolutions needed to achieve the real-world solutions necessary in it. There's another that's python based, and widely bashed for being 'too hard to reason about', but it has none of the problems with disgusting 'standard' libraries.
Where I am now, Facebook, has adopted a python based model... and used almost none of python's features. Instead, everything is generated through convoluted lists and dictionaries, and it's treated with very little exception as a json file. This saddens me.
Almost all configuration should be done by libraries with sane defaults. So should programming, even though it's not.
There are certainly people who just copy-paste code without understanding it. As an anecdote(because I'm on a role apparently with this topic) copying C# code off of StackOverflow that does AES encryption but sets the IV to all zeroes.. And not reading the comments that say this is bad practice but a customer constraint...
I guess my point is that a LOT of people seek and implement examples and while some don't understand them some find them extremely efficient and DO understand them.. In fact every time you use a library you are, in effect, copy-pasting code that you are relying on...
That, plus the whole "ugh, I don't want to have to write all this boilerplate from scratch".
Initscripts are mostly boilerplate, and it's just plain tedious do to anything other than copy/paste what you know works.
The first principles of computation are mathematical and believing something to be correct is not the standard. Proving is the mathematical standard. I take the author's point of reference to be more formal processes than the ordinary practical programming practices they criticize.
Software development is a perfectly ordinary engineering activity, or should be (although it is mostly done by non-engineers), and the same standards should apply. No mechanical engineer or electrical engineer has ever shipped a "provably correct" machine. They have shipped machines that conform to best practices, including heuristic analyses of MTBF and so on.
Out of curiosity, if the first principles of software development are not based in mathematics, what do you believe are their basis and how does it differ from "folk wisdom" if we take "folk" to include communities of engineers rather than as merely pejorative?
Pragmatism doesn't need to draw a cause/effect link like superstition does, and superstition doesn't generally lend itself to abandoning faulty beliefs.
If you work from the specifications for a library, you'll probably find that some of the documented features don't work. If they weren't used by some "framework", or mentioned in a how-to book, they probably haven't been exercised well.
While I'm on a soap box, I'd like to point out to those in the positions where they're work is evaluated by professionals and peers how lucky they are. I'd kill for a job/client where my code is evaluated by other developers. This community sometimes forgets they are the 1% of the development world. Those who work for the googles, facebooks, and visionary startups with leaders who came from the same cloth. The rest majority of us are laboring away under management that have never produced a lick of code, or even design, for that matter.
I guess the Tl;DR is: appreciate the fact that you even get to consider the finer points, and pray for the rest of us.
The initial reasoning for them was to prevent accidental assignment, since a C/C++ compiler will complain on this: if (null = x)
but not on this: if (x = null)
In Java, neither is allowed, so the whole yoda construct is almost pointless. The exception is boolean assignment which does benefit from this idiom, but the loss in readability is a trade off to consider. In any case, I suspect most uses of this construct in languages like Java fall into the "superstition" category, where the people using it aren't really considering why.
if(0 == very_long_function_name_here(param1, param2)) {
Tells me a lot after I've read a little of the line, whereas: if(very_long_function_name_here(param1, param2) == 0) {
Makes the == 0 part easy to miss, and buried behind a lot of clutter. if (VERY_LONG_AND_SPECIFIC_ENUM_NAME_HERE == func(param1)) {
}
Makes the function called easy to miss and buried behind a lot of clutter ;)If your code has a lot of clutter it's usually telling you something about your design.
For readability a better way of handling it might be for the function could return a boolean instead of an int:
if (is_valid(param1, param2)) {
self documenting code and all that.
if (0 == number_of_students(context, classroom))Also, there are many more ways of dealing with readability: try different formatting and indentation rules, wrap lines manually where you think it makes sense and so on. As for your example, even the lowly C let's you do something like this:
int (*short_name)(int) = &veeeeery_long_and_ugly_and_unnecessary_function;
if(short_name(1) == 2) {
printf("\n\nYay, it worked!");
}
which completely solves the problem, no matter the order of compared objects. I'm not sure, but I suspect things like this are being optimized away by the compiler anyway, which would mean that you can use it anywhere you want without worrying about costs of indirection.In higher level, modern languages you have even more, much more sophisticated tools for doing this kind of things.
How long identifiers are acceptable depends on a language. With languages without namespacing or modules you obviously have to use some naming convention so that there are no name conflicts and this makes your names longer. I'd say 5 words is ok in this case.
But I'm not arguing against long names of things: on the contrary, I like having names as descriptive as possible. OTOH, too long names are also bad, because they make the code less readable and harder to work with. Like almost always it's a matter of balance: you need to know when to stop adding words to a name. I think that "name is too long when it starts making other names in the same line much harder to spot" is a good heuristic for this.
Not to mention it becomes more verbose once you make that alias a const ptr (which we do on all of our local variables).
You'd be confused once or twice, then you'd learn the technique and you'd stop being confused. Every code pattern was unfamiliar to you at first. And confusing, until you internalized it. It's unrealistic to assume that you can ever stop learning new patterns - try switching to another language and you immediately have dozens of unfamiliar, confusing patterns to learn. (it gets better after a certain amount of languages known (like https://klibert.pl/articles/programming_langs.html) because you start noticing meta-patterns)
My C is rather rusty nowadays and using function pointer here may not be the best option, but as someone else said, there are other language tools for doing this kind of aliasing, like #define. I'd go for function pointer probably, because it reveals not only a name, but also a type of function and it's guaranteed not to escape the current scope (unless explicitly returned) while #define has no knowledge of scopes at all. In languages which support real macros, and preferably lexically scoped macros (like Racket) I'd use those. In languages with closures and first-class functions I'd probably do it in yet another way. But in general, if I find myself working with a name so long that it makes it hard to spot other names on the same line I will alias it locally.
Have you read http://shop.oreilly.com/product/9780596802301.do ? It's a good, short book on the topic and it discusses exactly this issue at length in one of the chapters.
Sometimes you might not have a say in how the function was named... having some practices to deal with the unruly code that hasn't been graced by one's own perfect sense of style isn't terrible! ;)
try different formatting and indentation rules, wrap lines manually where you think it makes sense and so on.
All fine ideas, but I wouldn't discount yoda-conditions as not being in the same category.
Some languages are more opinionated than others, but one positive thing about "superstitious" programming is that code is often more consistent because of it ('pythonic' PEP 8, code patterns, skeletons/boilerplate, ...). Many times you'll end up seeing the same patterns elsewhere, for example adherence to Google C++ style guidelines on projects completely unrelated to Google -- simply because they are both practical and familiar.
In a similar vein, I'd probably religiously opt for a #define ... #undef pattern for your given example, because I've seen it more often than using a separate function pointer variable. There isn't really a technical advantage, but mostly one of familiarity.
someString.equals("someString") vs "someString".equals(someString)
They are hard to get right and many apps don't daemonize correctly or return OK before they are ready to be running - and then are actually running much later. Upstart/systemd didn't neccessarily make that more clear for people to understand either.
While I haven't tried it, I recently encountered a reference to https://github.com/jordansissel/pleaserun and it sounds promising.
I don't know how much of this applies to programming in general, really, but init scripts are.. yes... special things.
There was a discussion I recently stumbled upon about init systems in containers. Many processes spawned lots of zombies! You need a full init! Well, no; If your process is forking new processes, it should probably wait for them. If it's forking other processes that fork other processes, and they're not cleaning up zombies when they're done, that is a bug in those processes. If all else fails, your root-level process should probably implement a simple init on it's own (a wait() on loop in the background, dumping wait objects into a LRU map that other processes can do an equivalent of waitpid() on). Or just run under bash.
It is true that there are standards that should be read and followed; the ones referred to in the article are probably among these.
But other standards have a fair amount of wishful thinking in them. (How many of you JS developers have read the latest ECMAScript standard?) Sometimes, what has worked in the past is a better guide for practice than what ought to work.
In the end, the whole script usually revolves around four lines of actual code (starting a process, stopping it, querying the status, optionally sending it a SIGHUP to make it reload its configuration), but the result is a script with 50 lines of boilerplate.
Which, of course, could nearly all be automated away by a sane, declarative format.
Leaving all the political mess aside, that's the part that systemd got right. You write sometthing like
[Unit]
Description=some human-readable text here
[Service]
Type=simple
User=...
ExecStart=/path/to/deamon --with=options
KillSignal=SIGKILL
[Install]
WantedBy=multi-user.target
and that's it.Quick comparison of sysv init script lengths on Debian Wheezy vs. systemd service files on Debian Jessie: ssh: 162 vs. 15 lines, cron: 92 vs 11 lines, dbus: 122 vs 9 lines
My point being: it's a bit too much to demand that people carefully engineer their sysv init files, when about 90% of those very same init files is simply boring, repetitive boiler plate code. Also, what happened to the "don't repeat yourself" principle?
I tried writing a sysv init script from scratch. I stopped because of sheer boredom, and resorted to copy&paste to save both time and nerves.
In addition, the BSDs have much shorter rc scripts mainly by putting all the boilerplate functions into a single library file that is then sourced. Gentoo's OpenRC framework, as well.
Upstart jobs are really simple, and OpenRC and modern BSD initscripts are also much simpler than SysV (and old BSD) initscripts.
I think we're at the point where regardless of the systemd vs. everything else debate, we can all agree that SysV initscripts are shite.
> I tried writing a sysv init script from scratch. I stopped because of sheer boredom, and resorted to copy&paste to save both time and nerves.
Better yet, at my last job, I had some extra time, so I rewrote Ubuntu's qpidd initscript as an Upstart job, and it was so satisfying to see that giant script get condensed down to a few lines.
Yes, because as engineers our time is more valuable than a minimal and perfectly-standard-compliant solution. I would happily read all the RFCs, and all the significant papers in CS, and all...etc.
The reality is that we don't get paid to read, but to ship.
I've worked pretty recently as the most senior developer on a team, and I saw tons of issues caused by copy/pasted code. It results in a lot of issues. Subtle bugs around `==` versus `===` in JavaScript, unecessary variables that obfuscate the code, Python that isn't Pythonic, C that was written as an example and therefore doesn't include the error handling that is constantly necessary in C (how many C examples check the return value of malloc?).
The biggest problem, I think, is that often you end up with 10 different ways to solve the same problem in the code because instead of looking at the existing code, people Googled for a solution with subtly different keywords than the previous person and found a different blogger. This kind of duplication is very hard to identify and remove, and this does cause bugs.
Of course a more experienced developer can work around these issues, but a more experienced developer also knows off the top of their head how to write the code themselves and usually does it.
It's interesting to code in large code bases. They end up getting to the point where no single person can ever understand the entire thing. It is usually at this point that things like Hungarian and strict coding guidelines really help. If you need an array of pointers to some kind of object, you can be reasonable certain that search will be able to find it.
If your code base never gets large, you've failed as a company.
Automation helps: linters, static checkers, runtime memory checkers, automated tests, all should be run on every check in. On a C# project I even created a tool that correlated diffs with a code coverage tool and rejected diffs that weren't covered by unit tests (it misses cases where code paths were executed by preexisting unit tests, so it didn't necessarily require a new test for every diff--I consider this a bug). But you always end up making compromises because there's Xthousand lines of code written at the beginning of the project that you don't have time to go back and write tests for.
One of the reasons I'm really excited about Rust right now is that it makes it easier to set these things up at the beginning. A lot of memory checking you have to do with separate tools in C comes free with Rust's type system, while Cargo makes it very easy to get unit tests up and running. My hope is that if Rust finds wider usage, the projects I come into will be more likely to have been set up properly from square one, and it will be easier to work on larger code bases.
I dunno, I'm somewhat new at technical leadership, so I'm still working out some of this stuff.
I mean, I'm sure there are some cases, but how many C programs are capable of gracefully recovering from a malloc failure? Probably very few. If the program is going to crash anyway...
Take networking code for example: do you really want to be dying without a visible reason because you didn't check `errno` and do a retry? This is really basic stuff.
Certainty is the alchemist's stone of philosophy. People have done all kinds of interesting stuff while attempting the impossible: turning knowledge into certainty. They can't, but that doesn't mean knowledge--which is inherently uncertain--is inferior. It means certainty is inferior, because it can't actually be achieved, regardless of what any pre-Bayesian thinker might have imagined.
Knowledge is always uncertain.
Tests are always incomplete.
This does not mean "knowledge is impossible" or "you cannot say you tested the code" but rather "certainty is a chimera" and "compete test coverage is a futile goal."
The whole point if good engineering is to adopt standards that are both achievable and useful, and the kind of test coverage that people run on code they are unfamiliar with to ensure it will do what they want it to is generally adequate for that. Evidence: most shipping code actually works.
The vast majority of time spent on software development is on bug fixing and on many projects the construction/coding defects make up more than half of all the defects.
Or wait... no, maybe your characterization is wrong.
Yeah... that's probably it.
I said if you do not understand what you are testing (not the compiler, not the runtime libraries, not the OS, not the instruction decode, not the microcode, not the gates), you can't say you actually tested it. You merely observed what could be a side-effect.
We specialize and abstract for a reason. No one understands everything about everything surrounded their code, and that includes other software. Don't tell me you've personally been through the code of every single project your projects touch.
You haven't and that makes you a hypocrite.
I'm also a troll when I point out someone is trolling.
do you know how I turn female? By pointing out that someone is female.
It's amazing how making an observation suddenly means you've become the thing you're observing.
Or not, but who am I to judge, amirite?
Consider Bertrand Russell's Chicken:
The man who has fed the chicken every day throughout its life at last wrings its neck instead, showing that more refined views as to the uniformity of nature would have been useful to the chicken. [1]
Sometimes things working but not in the way we imagine they work is worse than them not working at all.
I don't think it's at all safe to assume that a comfortable life with a definite-but-unknown end point in your prime is definitively worse than a life subject to the random events of the wild. Particularly for something that's down the lower end of the food chain.
Truth of the matter is you can't just stare at a man page or textbook all day and try to understand, then code. You'll code some mess you don't fully understand, sure, but later on you'll come back and either realize it was wrong and fix it, or you'll understand why it was right. But you will eventually understand.
The people trying to understand it will have gotten nowhere in the mean time.
StackOverflow has improved my programming life dramatically, because it makes it a lot easier to "memoize" the search from what you're doing down to first principles. Whatever question you're trying to get clarity on, there is probably a StackOverflow answer out there that builds a positive case for a particular option, going all the way back to primary documents like standards.
A lot of this is surly curiosity and you can get away without it most of the time, but at least as soon as I run into the first problem it really helps me a lot to have at least a basic understanding of what is going on in the layer of abstraction I am interacting with and maybe a layer or two below. And it gives me a really comforting feeling, that I know what is going on better than describing it as some black magic.
If you examine why you're concerned about some code more, you will probably come to the realization that you're less confident in it because it is less old and mature, has fewer users, and has had fewer developer eyes on it and such.
I.e. whether explicitly or not, you're focusing your understanding where it will be needed and that coincides with where there is more risk if you don't understand.
"If you say that you don't take this relatively fast road for Linux init scripts, I'll raise my eyebrows a lot. You've really read the LSB specification for init scripts and your distribution's distro-specific documentation?"
only when i get stuck do i engage critical thinking. maybe this has been to my detriment.
very often i notice myself and colleagues making up nonsense explanations for why c++ behaves the way it does without any real basis other than intuition and truthiness.
A 'good' example is one that you've actually tested for real to make sure it works, that follows the intended principles for the system in question, that meets actual use cases developers using it will have, etc.
A significant amount of software is released without good docs -- examples being just one part of good docs. So of course developers copy from the examples they can find.
For example, creating a new init script or a Debian package description is something I don't do all the time. I don't want to learn all that stuff because by doing so I'd basically push out something else from my mind. Copying an existing configuration or script and making modifications without trying to create something original is an excellent pattern for things outside your immediate expertise target.
The "avoiding creating something original" part does indeed carry a sort of "superstition". It's like relying on a new culture by merely mimicking it, because you don't really know it well enough to break the rules correctly. You don't know why things are done the way they are but your best bet is just to copy and adapt. This creates a sphere of fuzzy knowledge where intents and black magic seem to alternate.
When I'm programming (in Python or Scala) I can and do check the language specification if something is confusing, and I don't think I'm alone in this. It may be a "significant" amount of programming, but there are definitely areas of programming with more rigour.
[1] I blame Debian, usually quixotic in its adherence to standards like the FHS - but when it came to the LSB they prefer to pretend that "alien" counts as RPM support. That set the tone for how much other distros tend to care about the spec.
http://refspecs.linuxfoundation.org/LSB_4.1.0/LSB-Core-gener...
http://refspecs.linuxfoundation.org/LSB_4.1.0/LSB-Core-gener...
A bigger problem is Debian gratuitously using different SONAME than required by LSB for some libraries - your LSB binary can't use those system libraries if it wants to run on Debian. But at least it looks like they fixed that problem in Wheezy.
Also, almost every time I go back to review code I wrote more than a year ago, my first reaction is "jesus, that's a stupid pattern/implementation. Why did I do it that way?". I've been a professional developer for about 12 years now and I suspect that will never change.
This has been especially true in programming. I've been writing programs in some capacity for 18 years now and not once in that time have I been able to simply read the reference manual of anything and following that write working code on my first try. Well, obviously I've been able to do it for trivial things, like after the Python 2->3 switch when map/filter/range started returning iterables instead of lists, simply being told "map/filter/range/etc. now return an iterable and not a list" is sufficient to explain the change in behaviour, but that's only because I know all the concepts behind it. You could say the explanation is only one level removed from my intuitive knowledge. Things start getting exponentially harder the more levels removed they get, i.e. if it's a new concept explained using new concepts previously defined, but for which I have no intuitive knowledge what they mean.
Even if it were a common ability to produce engineering work from written design principles, I'm not certain that's what we want. Experimentation is the key to science. You have to observe, change, guess, run, observe, change, observe until you can successfully predict what will happen for each change you make. The scientific method is taught to everyone, even if most ignore it and skip steps. 'Guess/check' is the correct way to do real-world engineering and design, not pure-logic simulation, because we still don't know all the first principles.
Computer science is a branch of mathematics. Here you build models, reason about them, find properties, then prove them, then rely on them. Tests - a.k.a. experiments - are usually not exhaustive and don't provide guarantees of correctness.
Having said that, I agree that experiment is of crucial importance - how else you would validate your models in the first place? However, both experimenting and reasoning can be flawed, and part of the trick is to learn how to do that properly. Computer science is difficult.
Proof is of critical importance - It's how you find those junctures, for example, how you optimize, how you design and how you reason about runtime, etc. But programming as practiced in the field is engineering first, mathematics second.
> Beware of bugs in the above code; I have only proved it correct, not tried it.
Deep thinking and mathematical proofs are tremendously useful when thinking about a well-defined and scoped problem at a more or less fixed level of abstraction.
The problem with software engineering is that in practice all abstractions are leaky, and so you inevitably find yourself dealing with issues that are too far-flung and random to be mathematically tractable. The best you can do is be rigorous about the core problem you are solving, but there is always some amount of hammering and duct tape to build it into a non-trivial real-world system. It's possible to attack this problem asymptotically with a rigorous engineering process such as NASA employs, but we don't because it's simply not cost effective for the majority of software.
That's a great point. We utilize best practices, like automated tests, to narrow the gap between pure experimentation and mathematical proof. Proofs in most code would be difficult to impossible because most of use libraries that use other libraries, etc., so we just do the best we can given the time and monetary constraints that we have.
I don't mean this to be a "mathematics from nature argument" (even infinite sets are hard to justify as coming from nature), but the course of history has usually been "solve similar problems in an ad-hoc manner many times, then generalize"
A sort of humanitarian science, since almost everything we touch have been made by other humans. But somehow STEM people are famous for looking down their nose on the non-hard sciences.
Trying to do anything in computing from first principles is like trying to code without mistakes. You're losing the value of iterative design.
Also, if you never work things out yourself, we'll only ever have one way to do things. What if a better approach exists but nobody's ever tried it?
It's good to double/sanity check what you came up with afterwards, and if the only way you can think of to do something is obviously seriously flawed, then don't go through with it and look it up instead.
Things like Shannon's description of information-as-entropy were certainly a mathematical discovery about engineered systems, and led directly to a lot of coding and data compression theory.
There wasn't a whole lot of interest in quantum channels at the time, so the fact that the theory fed something back to mainstream physics was a bonus.
I think it's reasonable that people want to create software components that have been proven to work and then just forget about their internal details and copy-pas... apply them.
It's also what Bret Victor has been getting at. Very often the component's input-output mapping is much more relevant than implementation details. And you can get a feel for that by fiddling with the components.
"A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works and cannot be patched up to make it work. You have to start over with a working simple system." – John Gall (1975)
It's a very general point - how to interact with things is often dependent on the surface, not the volume. This is the generative insight for the systems approach to engineering.
For example - style considerations. Important? Yes? Maybe? Not at the moment when it reaches the computer, but important when humans have to grok it.
So then should we "lead" development by promoting style? Will that get us somewhere? Probably not as much as some other strategy. But it's still sort of important and this blogger made a nice post the other day about style, so...
I try to guard against cargo-culting (I abhor it).
The cognition of programming languages is very much like regular human languages. It takes a tremendous amount of disciplined, repetitive practice and observation before the mastery of a language allows it to just flow effortlessly out of ones fingers.
A simple concrete example is the difference in articulative ability in constructing the same exact sentence for an essay between a 12, 18 and 25 year old post-doc. For the post-doc who has written a countless number of 10+ page essays versus the 12 year old in 6th grade who has more-than-likely never written anything beyond 1-2 pages, the ability is obvious and striking. The patterns and structures come naturally to the well-practiced individual as the Chinese parable of ZhaungZhi teaches us the the concept of achieving flow (WuWei).
The difficulty with mastering programming the way we speak english (or any other primary language for that matter) is that many projects only ever need to construct the equivalent of a proper sentence once, then that sentence is committed to source control for eternity without ever needing to be constructed again, thus the cognitive muscles creating programs tend to not achieve flow, but the muscles of identifying, locating, copying, pasting & modifying do tend to achieve flow.
To exacerbate our inability to construct things from scratch, we also look to reduce the monotony and abstract away the difficulty and reduce what actually needs to be written and kept track of from our keyboards via APIs and abstractions that segregate duty. Of course I'm not advocating that good architecture be thrown out for the sake of practice, but in a sense, programming could use the equivalent of musical scales, where the well known patterns, and language-native constructs are exercised daily (frequently enough) to the point of trivial mastery.
The argument against such kinds of practice from my peers & others on this board has been,
"duh, that's what API documentation and computers are for." "This stuff can be looked up." "This is why whiteboard programming exercises are not indicative of programming ability." "See, why learn math, when I can just type x+y into the console and get the result?"
However, I'd argue that the effortless mastery that comes from daily practice comes with tremendous benefit. The mastery allows one to focus on the architecture problem set rather than the details of what to name a method (assuming that good method names come naturally from practice). Or be encumbered by whether the first or second argument in the split method should be the separator.
Would it be acceptable as a passenger of a bus if the driver claimed that they did not need to know whether the left or right pedal was the gas or brake pedal since that could be looked up via Google in the API documentation for the bus? Of course the scope of learning how to drive can be mastered in a couple of months, but programming is more equivalent to spoken language and perhaps mastery of playing musical instruments.
I would argue that much of the arguments for or against certain interview techniques/questions/strategies really stems on what the interviewers and interviewees see as a signal for effortless mastery. Esoteric pet questions test for intimacy with a narrow topic, but also get to assume that other deep-related knowledge follows along. Whiteboard interviews are looking for problem solving ability, but also effortless mastery of a language and basic data structures that someone with 1-2+ years in a language probably should have.
But at this point, I digress.
On the other hand, you should only be repeating the same patterns so many times before you build something so you don't have to replace yourself again...
Maybe it's an 80/20 mix; mostly repetition, part novel content.
I'm pretty sure nobody actually learns from first principles,but understand first principles from things they've already seen
Same thing with languages, there are schools which approach foreign languages formally, starting with grammar and all that. Some of that schools are known for training spies, for example, which is sort of an indication of their quality.
Do not listen to those saying it's not possible. We do implement programming languages based on their specifications, we implement protocols by reading RFCs, we implement numeric algorithms by reading pseudocode in the papers, etc.
It's all trivial and mechanical. The latter notion is very important, one have to understand that there is very rarely a place for "creativity" and even thinking. You simply translate specification from one formal language into another, following simple rules. Thinking too much considered harmful.
We implement spherical-cow versions of languages, protocols, etc by reading those papers and specs. Then we spend months or years fixing it to implement the actual languages and protocols.
And what you seem to suggest, start writing a parser and then experiment with various real-world inputs until you're satisfied, is certainly not a very productive way of doing things.
Do not confuse principles (which are universal, simple and beautiful) with specs (which suck shit a metric ton per second).
> it's just a more structured way of hacking it until it works
You're confusing hacking with cargo cult coding. Hacking until it works is exactly this formal loop: collect the data, build a model, test if it's applicable. Cargo cult is "google for an answer, paste some code from stackoverflow, see if it works". The OP article is about the latter.
>This isn't something that we like to think about as programmers, because we'd really rather believe that we're always working from scratch and only writing the completely correct stuff that really has to be there;
This is the issue I have with this article - the author assumes everyone works in exactly the same way they do and so are vulnerable in exactly the same way to their vulnerabilities. I don't know anyone who actually approaches programming in such a way that they think they're writing things from scratch and it's correct. I mean, defensive coding is a concept that has existed for decades that specifically addresses how incorrect the code we interact with is. The fact that I started coding in an environment where I had to manage memory myself and now I don't makes me very aware of the fact that everything I'm using is an abstraction built upon other abstractions.
I dunno, the day I learned to program in Intro to C and Intro to Computer Science, defensive coding, not trusting user input, and the completely understandable at the time weirdness of the guy who came before you have been stressed as things to pay attention to.
I don't think I am alone in this.
The quicker your test suite can find that logical explanation... ;)
Wheels work just fine. Choose the best wheel from those available and tweak it as desired. Test it to make sure it doesn't shatter under too much weight. Apply and move on.
This pervading notion of "if you don't understand you shouldn't do it" is silly. Sure it is nice to understand, but it doesn't put food on the table. Sometimes it's just handy to use the gun as is and go shoot your dinner in the woods. You don't need to understand the formulation of gunpowder or rifling in order to use a gun properly for it's intended use.