Using the programming language of your choice, sort the strings in this file and remove duplicates. You have 30 minutes.
sort foo | uniqUsing the programming language of your choice, sort the strings in this file and remove duplicates. You have 30 minutes.
sort foo | uniqI’ve worked with engineers who are competent programmers but just can’t handle these problems when they come up in practice. If the job actually requires it, I don’t know of a better way in the context of an interview to determine if a person is that type of programmer than with a few toy problems and a solution in code. You may miss some good people because the format is suboptimal, but you will weed out the people who just don’t have the theoretical and practical background to solve this particular type of problem.
I think there are better ways of doing things than the standard X interviews of 1 hour. One method that works well is hiring people for short term contracts (few days) and seeing how they work with the team on real problems.
The parent comment of yours identified the right need for them (a person who can come up with a solution when the time comes in the real world), but the method of toy algorithms means you could end up with the following scenario:
Person A: Spends months drilling solutions on whiteboards to memorize the answer. They pass as they can regurgitate every toy algorithm possible without thinking. They pass the interview.
Person B: doesn't deal well with interview situations and doesn't get the chance to talk in general about how they'd approach it in real life. They freeze up as they're stood up in front of a panel of people holding a whiteboard marker and being stared down. They don't progress in the interview.
Perhaps person B was the person who would sit at their desk for 20 mins in silence and come up with:
"ok if we're not deduping these on the inserts for our system, I guess the least we can be doing is maintaining one of those in-memory bloom filter things, I read about those one time as a good probabilistic data structure, might end up being a part of solution. I'll ask the tech lead if we have any memory constraints as I see that our AWS instances are compute optimised instead of memory optimised, anyway, we might get an off the shelf solution"
Person A: "which question in my 'interviewing for algorithms' text book does this belong to?"
You want Person B, you optimised your interview process for A.
The problem is that most really great people, who are in high demand, won't put up with this because they know they can get someone else to hire them with less hassle. Your approach ends up weeding out your best candidates.
I can see that there are scenarios where it might not work, for example if non-competes are in place and the industry is similar. But in most cases I think there's enough flexibility to make it practical.
I've seen the typical examples of ugly code over the years; over-complicated conditions, piles of IF statements, things that should be in a database or associative array, all because the person just couldn't solve the problem in a simple, straight-forward way.
These people get stuck on what should be a simple problem and productivity plummets. If you've spent the whole morning on something, taking it from 5 lines to 105, you're probably attacking the problem from the wrong angle.
All that for stuff which can be solved with simple 1-10 lines of code. (not these everlasting functional programming chained to eternity lines. real slim and dumb code)
sort -u fooUsing sort -u alone might actually not be as fast as sort | uniq -u.
sorted(set(foo.strip.split(' '))) #although we need to open the file first
just to be on the safe sidePerl's not too shabby either for this kind of stuff :)
By the time you say “unicode” and “collation” the programmer is going to be in serious trouble.
Fortunately the programmer can lawyer the problem. The original requirements don’t specify what the strings are sorted by. I choose “position in original source” file and complete the task with a NOP.
tr ' ' '\n' < foo | sort | uniq
> Broken. uniq alone removes only adjacent duplicates. You need uniq -u to remove all duplicates.
As benchaney (https://news.ycombinator.com/item?id=17060154) points out, the duplicates will be adjacent after `sort`.