Tests for programmers Part IV: comparing routines
solipsys.co.uk
solipsys.co.uk
This contradicts the statement that 'given an adequate spec two competent programmers will solve a problem in the same way except for stylistic differences'.
I've seen that tossed around quite a few times and if it doesn't hold water for such a simple problem it certainly won't be left standing if we start analyzing larger pieces of code.
On the whole I think this research is one of the most interesting things on HN as of late.
In e.g. Haskell everyone would just do
condense_by_removing :: Char -> String -> String
condense_by_removing c = filter (/=c)
Because there's much less going on here--the loop being more or less implicit only--there's less room for variations.Mind you, Haskell programs also show lots of variability if the tasks gets only slightly more complicated.
I find this very interesting -- a popular idiom in one language can't even be represented in another.
I've found that languages like Java make assumptions about your memory, and languages like Prolog make assumptions about your data. I'm under the impression that in Haskell in-place modification is an interpreter optimization, whereas C is increasingly employed when resources are constrained and abstraction layers are deemed too costly. Consequently, C is maligned for resulting in faulty implementations behind every corner, yet it remains one of the most used languages.
The language Clean supports this notion explicitly in its type system (and uses it for IO, instead of, say, Monads like Haskell).
However, I will investigate further the "tree edit distance" to see if I can bootstrap something quickly, just out of interest. Thanks for the reminder.
Baker's alumni web page: http://cm.bell-labs.com/cm/cs/who/bsb/index.html
It's unfortunate that Baker's "dup" and "pdiff" haven't gotten the open source treatment, or at least if they have, they're not widespread. SCO could have saved themselves a lot of lawyer's fees by running "dup" over SysV and Linux sources to see what's similar.
The two that spring to mind would are the MCL clustering algorithm that could be applied quite easily to the similarity matrix. As a more academic endeavor, it'd also be interesting to see what kinds of differences in similarity you'd get by applying Needleman-Wunsch.
http://www.micans.org/mcl/ http://en.wikipedia.org/wiki/Needleman%E2%80%93Wunsch_algori...
What's interesting is the large clusters where either a for or while loop were used. I'm curious whether if given a (limited) choice any hiring manager would have a preference between the two implementations?
TLDR: How much does/would coding style influence the hiring process?
Yes, switch. Addmittedly not as the looping structure, but for the internal decision process of the loop.
Some of these choices are indicative that the programmer isn't a native idiomatic C programmer, but some might be genuine choices for one reason or another. That's when style choice exposes internal concepts.
More later. (On current work load - much later. It's taking about 2 weeks per article here, so it won't be quick. Sorry.)
I'm so busy at the moment that important things are starting to slide and get lost.
Sorry for the delay, hope I've given you enough time to sort out travel arrangements.
I don't think 'style' should be the deciding factor in something like this unless the difference in style translates in to 'readable' versus 'unreadable', old-school coders would do it in one style, people that came to C later will probably do it slightly different.
The main criteria are:
- does it work ?
- is it reasonably efficient ? (as in within a factor
of two or so from the optimum solution)
- is it reasonably readable ?
If all of those are good then you could consider the programmer to have passed this test. Such a test by itself is not enough to disqualify someone for hiring or not hiring based on the style alone, assuming the above three conditions are met.It's just one little element in the total, one of the many 'and' gates you have to get through as a candidate when applying for a position.
Not a deciding factor, but there's a lot of subtle info in a whole code sample.
When asking someone for some OO code, I sometimes see OK-but-a-bit-naive code, which includes variable names like "theObject".
That suggests to me that the most important characteristic of this entity in the coder's head is that it is an object - i.e. that they're not very familiar with OO code.
Basically, information like the choice of variables names provides a window on the mental model the coder has of the problem. And that's useful.
I agree with you on the naming issue, but the funny thing is that that sort of information is exactly what was removed in order to make the analysis possible.
(OK, yes, for exceedingly large inputs there are issues about sizes of integers, etc., but that's not the bug. It's much, much simpler.)
Certainly there is a bug other than the one you think you've found, and I don't think the circumstances you mention demonstrate a bug at all.
I would certainly be interested in seeing a more detailed analysis.
I don't think I've found a real "bug" yet. You include, but don't use, stdlib.h; you exit with status 0, even on error; but these are nitpicks, not what you mean. I'll think a bit more.
I return 0 in all cases because my test succeeds, even if the routine it's testing fails. It's up to my shell to decide what to do about that error. As it stands it reports the error, but it has succeeded in doing so, so it hasn't failed.
But that's not the point, as you say.
And the real bug is still there.
Will you give out the answer at some point?
Are we going to see some performance data?
While pure performance should not be the overriding consideration, it would still be interesting to see how everyone's routine stacked up.
I would hazard a guess that some dissimilar looking routines will have very similar performance.
Of course, in reality, all I am looking for is the confirmation that my code did not suck too badly...
I realized my code has a bug.
It's too late for a resubmission, isn't it? :(