Lost skills: What today's coders don't know and why it matters
itworld.com
itworld.com
"Today," says Rudin, there are occasionally times with you can't use the integrated debugger in your IDE (usually with some weird web application frameworks and server configurations), and younger programmers are at a loss as to what to do, and resort to hack-and-slash coding to try to randomly fix a bug, using guesswork. Me, I just calmly put in some code to display output values on the web page, find the bug, and fix it."
Oh come on. What he's describing isn't some "lost skill". If you can't figure out to use print statements for debugging you're not much of a programmer in the first place.
Most statements in that article describe really basic, common sense stuff.
Don't get me wrong: 'debugging by echo' is useful at times, but I would say that 9 times out of 10, a proper debugger is more valuable. In my case. YMMV. etc.
I can quickly run a test, grep/skim through the trace output and ignore/drill down into detailed minutia that are irrelevant/relevant to the problem being investigated.
I also tend to write my trace statements in an easily parsible format so if I need to I can write another program to analyze what happened and find the problem. Writing a full on program to process log files happens less often than chaining several unix commands to find the needle in the haystack (usually the program happens when I need to thread together several widely separated trace lines).
Of course, I didn't say nor mean to say that ALL race conditions are identifiable in this way, only that it can make the results more visible for some of the more obvious cases.
With Python, when stack traces happen now, WebError will even give you an in-browser prompt to explore the state of that Python session; so log.debug() calls are only used when the software is working but not working as intended - that is a case where debugging tools can actually come in handy because I hate cleaning up log.debug() calls littered everywhere.
For most single-threaded or simple multithreaded programs, it's easier just to throw a print statement where you need it and analyze that. Even with a large data set, a simple grep will probably get you what you need.
When you're debugging a complex, multithreaded program, on the other hand, an external debugger is much more valuable, because you can break at the exact point where things go wrong, and examine the entire program's behavior at that point. Debugging a race condition with prints can be pretty difficult, especially when the printing changes the timings of the threads.
It's a trade-off, but I usually start with traces.
In fact, I'm not exactly sure how you would debug a multi-process parallel system with an external debugger. Would you simultaneously use multiple debugger instances to attach to each process? That sounds like fun.
For a true, parallel system it's virtually impossible to "examine the entire program's behavior" at any one instance in time. Sure printing changes the timing of things, but unless you are single threaded, you are kidding yourself to think that a debugger also doesn't disrupt timings.
For example, I find printf-like debugging and using stack traces from Valgrind to be very useful for debugging my own code; however, at work, when I have no idea where the problem is, something like gdb or the debugger of Visual Studio really comes in handy.
> Bernard Hayes, PMP (PMI Project Mgmt Professional), CSM (certified Scrum Master), and CSPO (certified Scrum Product Owner).
makes me not want to continue reading much further.
To clarify: you become a CSM and a CSPO by taking a 2-day course each (well ok, there's an online exam, but you can look up the answers on wikipedia and pass). PMP is the only certification in that last that takes some effort, but it's totally unrelated to software and, thus, to most of this article's subject matter.
What, do they also expect us to cut, polish and etch our own silicon wafers as well?
Why? Because these things are awesome, and form the foundation of our culture.
Now, having said that: No, having hands-on experience with these things doesn't really help with programming. ;)
Why?
2. Foundation of our culture - no. Foundation of our culture (the engineering culture) is about looking around, identifying problems and using rational-logical thought to find solutions. It just happened that we learnt casting before making silicon chips. I have done Mechanical Workshop as a student in my undergraduate curriculum. I learnt nothing new but the fact that most of us want to cling to already established norms. Again, it's an unnecessary exercise. I really think that in a field that is evolving as fast as this, we should rethink why and what norms and parts of relatively ancient culture should be follow and where evolving and discarding old ideas is necessary.
For those interested, here's a series of books on how to build your own machine shop from scratch (by building your own foundry, casting parts, etc...): http://www.lindsaybks.com/dgjp/djgbk/series/index.html (I'm not affiliated with Lindsay Books in any way)
Tinkering? Cooking? Music? Writing? Reading? Puzzles? Even sleeping? It's all potentially good.
Ben Summers, Technical Director at ONEIS, a U.K-based
information management platform provider, points out
that "habits learned when writing web applications for
14.4kbps dial-up telephone modems come in rather handy
when dealing with modern day mobile connections. When
you only had couple of Kbytes per second, and latencies
of a few hundred milliseconds, you were very careful to
minimize the size of the pages you sent, and just as
importantly, minimize the amount of back and forth with
the server."
With today's mobile connections, says Summers, "the
latency is much worse than using a telephone modem
connection, and that's compounded by error rates in
congested areas like city centers. The fast 'broadband'
headline speeds are pretty irrelevant to web applications.
It's the latency which determines how fast the response
time will feel, and tricks learned when phone modems
ruled the world come in awfully handy. As a bonus, when
someone uses your app on a fixed connection, it'll feel
as fast as desktop software!"I'm curious as to the effect it will have, if any. So far it's resulted in two hits to our web site, neither of which explored any further than the home page.
(Two more hits since the last comment!)
And by basic I mean not understanding the different performance implications of a table scan, an index scan and an index seek in an SQL query plan. And also why "it takes ages first time, but is quick after that (when everything is already in RAM)" is usually not acceptable (every time could be the first time around if the query isn't run often or RAM is limited).
Some of the stuff that article lists is just not needed at all by a code, really. Some are strictly hardware issues. Others are oddly specific: "programming tight loops" is part of the complexity theory thing: understanding how a process will behave at relevant scales and optimising accordingly.
The biggest gap, IMHO though, is a lack of knowledge about big O and why quadratic behavior can be bad, what it is, &c... That goes hand-in-hand with a lack of knowledge in algorithms. Why is bubble-sort considered bad? What's a generator? Why does everyone keep saying to use xrange() in python? Why is it bad to use list concatenation?
I remember similar articles in the 80s and 90s bemoaning how "programmers these days" didn't know how to use a protocol analyzer or logic probe, or didn't know that xor made for a faster register clear operation (or moveq for 68k fans), or any other number of esoteric trivia that, while useful in context, did not usually contribute significantly towards a programmer's ability to get the job done.
My first debugger was an in-circuit-emulator for a Z80. It was the size of a small television set, had a crappy UI, and limited functionality. Today's debuggers can be hosted on the system itself, and have become so powerful that most people don't know how to use them to maximum effect (myself included). IDEs check your code as you type. No more writing something in VI, compiling, tracking down the cryptic error messages your compiler spat out and trying to figure out where the REAL error is because the compiler is dumb. You're shielded from the ugliness underneath, and for 99.9% of cases that's more than enough.
Do "kids these days" really need to know the sound of a hard drive dying? The last drive I heard going bad was in the 90s. Since then drives have become so quiet that you'd need a stethoscope to even hear the arm thrashing (which is why I use RAID). And how useful is the knowledge that you can open up a frozen drive and spin it up with your finger going to be as disks are replaced by SSDs?
We live in the future, where things have gotten a LOT better. Do the new batch of developers really need to know assembly language? After moving to "fluffy" languages, I only twice found need to use it (once to disassemble a stack dump from a JNI crash, and once to monkey patch a buggy device driver). Twice in all my years since using Java, Python, PHP, Objective-C, Scheme, COBOL, VB, and C#. Was it damn handy to have the right skill at an opportune time? Hell yeah. Does EVERYONE need this skill? Hell no.
How about bit packing? Memory has become so cheap and plentiful that even routers come with 16MB or more. Beyond low level networking and peripheral protocols, what use is there in packing up bits and coming up with clever encoding schemes that make for complicated (and potentially buggy) codec routines? Saving one byte in a packet header is hardly the triumph it once was.
All the "kids these days" need to know is their algorithms, profiling, debugging, multithreading issues, and how to write well structured, maintainable code in their paradigm of choice. The rest is usually industry specific, and can be learned as-you-go.
Don't worry about the kids. The kids are alright.
This is what I do every day at my job and there are other people who know how to do this, plenty of them. Most of them are EE's, though, and consider SW/FW development as their secondary job.
It isn't easy to hire embedded programmers, graduates with computer engineering, EE or CS degrees will have some knowledge but it's experience more than anything that will help you acquire these skills. When I look at resumes I look for experience building small circuits, Arduino or PIC. If I don't see anything like that or if the skill list starts with Java, PHP, ... it would be kind of unfair to expect detailed knowledge about how to use "volatile" in C or how to use a scope to find race conditions.
I don't think this is an essential difference. Programming has preferred the sequential models of computation, but more as a fashion than a necessity. Now, Turing's empire wanes as Church's empire waxes, and those parallel elements of computation that the hardware has hidden from software for so long can no longer be sequestered in silicon.
The funny thing about this is that the CS degree I got covered all of these to a fairly decent extent. I wasn't an expert on any of those topics when I graduated, but I had enough of a background working with them that I wasn't completely in over my head when I ran into all of that at work(I work on high-throughput, low-latency financial messaging APIs).
It's interesting work, but it's a very technical skillset, because you do need to understand all the various issues that can pop up.
I'm totally guilty of this. I write new code on new hardware, and have very little intuitive of how fast it should go. Is 10k ops a second good? 1M? I just don't know how fast it should go. Of course, then I pull out the profiler, and think about my algorithm, but it takes a lot of second-guessing to decide how close to the limit I am.
For example, I was writing some clojure code to write to a SQL database. I'm relatively new to the JVM stack. I was writing to the DB at 1MB/s. I thought "well, that's not great, but not bad. Maybe after network traffic and DB constraints, and writing to a laptop disk drive, I suppose that's alright". No, I replace the JDBC DB thread connection pooling driver, and the same code now writes at 8 MB/s.
It'd be nice if there were a web resource for general guidelines on what it takes to max out hardware. Basically, benchmarks for real-world tasks.
Here here!!
I had the same thought when reading the first two pages of the article. I'd love to be able to better intuit performance (or heck, troubleshoot slow systems - which I do more often). The problems I encounter are lack of accurate and understandable information about the underlying hardware and the various layers between my program and the hardware (especially important for me lately as more of my stuff runs in a VM).
It seems like you need to be lucky and find a mentor willing to teach this esoteric material.
I've met a guy who runs a company of about 6 programmers doing this, has more work than he can handle and has difficulty finding good enough programmers. I think they're mainly C++ but were recently trying to find a C# guy.
So it's in demand, but you've got to know where to look.
As an example of embedded programming, do some timing critical work with microcontrollers and you'll find all sorts of fun optimization problems. Recently I had an algorithm that took 13 microseconds to execute but needed to do it in 11 (there was another interrupt coming!). I got to have a good time with the debugger, understanding optimization levels used by GCC, reading lots of assembly, and playing with a logic analyzer. It's quite fun, actually.
Then again this is dealing with system-wide performance. Application-specific performance should be significantly less of a problem to fix.
I think the jobs that are simply "here is this function, make it 10x faster" would be pretty rare, since usually people don't know what part of the code is going slow. A lot of the times they'll guess "X, Y, Z is making it slow" but without a real performance analysis patching stuff all over the place just doesn't pan out.
Unless there are other requirements I don't see why you'd suggest sorting the elements.
I started to learn how to program in PHP. Back then there was a sentiment that PHP and similar high level scripting languages weren't real programming.
With the web so ubiquitous today, I didn't that sentiment had survived, but here it is.
I think our culture was and will always be based on exploration and innovation and this is simply moving to higher levels of abstractions today. There is nothing wrong with this.
However, I personally am not satisfied with simply being able to use an abstracted interface. I have a strong curiosity of how things work under the hood, of tinkering with something to make it do new things and even try and rewrite things in simpler forms.
I think a different kind of hacker evolves when you have a basic understanding of the entire technology stack. This breed is inevitably going to fade with the increasing complexity of this entire stack (breadth and height) paired with the current speed of innovation.
In the end, we can lament all that we are losing or work towards everything that lies ahead unexplored :)
Just like the phrase rtfm, there should be rtfs.
And I think they left out the most basic and fundamental skill or understanding: there are no silver bullets. A lot of programmers nowadays seem to religiously follow whatever new language is being hyped and try to fit their problems and tasks to the language instead of the other way around... a bit of critical thinking and seeing a bigger picture than "thisandthat is THE SH*T (right now)!!" would work wonders.
In the real world, real coders could not care less how god-like the latest scripting languages or NO-SQL-but-relax-data-maps are because chances are very good that my customers use Java and Oracle or a few other big names and since they are paying me, who am I to push religious plugs about the latest fads on them?
Bottom line is: real coders (should) just know enough about computers, hardware, software and networks to make educated decisions and guesses and typically they don't care that much which language they are getting paid to develop in... pretty much all languages "suck" in the way that they ALL have their short-comings and it is up to the engineer to understand them and work with them.
Edit: I know, I was an evangelical Java fanatic from about '95 to '00 or so.
Once relational databases were considered to be "academic" because of their mathematical underpinnings. What was considered practical was flat files, or graph databases.
I just want to point out context. This is HackerNews where a sizeable percentage of participants are either business owners, involved in a startup, or acting on side projects. As such, they have a lot of latitude in deciding what technologies they use.
Yes, if you're a J.P. Morgan specialist, you don't get to decide what tech you use. Yes, if you're a freelancer working for clients, you have to use the database of their choice, with pre-existing settings.
However, we aren't necessarily those people and we are certainly "real coders" in the "real world".
P.S.
J.P. Morgan has a group using GHC Haskell, and Gemstone Smaltalk in production so they aren't as 'uncool' as some may presume. ;)