I have prompts that test very basic concepts and nearly everyone fails. Resume fraud is rampant.
I have prompts that test very basic concepts and nearly everyone fails. Resume fraud is rampant.
So is interview fraud. The remote-interviewee-answers-questions-while-her-face-reflects-windows-popping-up-on-her-screen is tiring at this point. So, I decided to find a way to inform me if someone was being fed answers in a tech interview.
Behold, the low-tech whiteboard. Also known as a piece of paper and a pencil. With the candidates I've run into that do not pass the "smell" test -- where I think they are being fed answers -- I ask them to draw some things, on paper. It's not a true validation, but it gives me something of a clue.
I ask for a simple diagram. Different services in a network, for example. Or a mini-architecture. For their level, I'll ask for something that should be drop-dead easy.
I ask them to show me their drawing.
The responses I've received run the gamut of "I don't know" (after 5 seconds of deliberation) to "I don't understand the purpose" (after 5 minutes of silence) to "I need to shut off my screen for a while" (while refusing to explain why) to "it depends if your cloud is AWS" (not in any way remotely related to the question.) I did have a candidate follow-up with a series of questions about the drawing, which were feasibly legitimate.
This hand-written diagram is not an absolute filter (I've only used it maybe four times), but rather it can confirm some suspicions. I think I can generally gauge honesty from questions/tasks like this. And that's really what I'm after -- are you being honest with me?
It's imperfect, but it has been helpful.
The drawing approach also sounds like a good idea, though it's not like software is not going to evolve to be able to draw answers graphically which the candidate could copy down. By having them not able to input something into the machine, the only remaining option is someone listening in and feeding the answer on screen. Plausible, but that's a level of being prepared to cheat that the helper could also prepare to draw stuff out. Or they type with their feet but that's also a scenario where I'd be happy to have them come in for a final interview and demonstrate this amazing ability!
Hm. Unless the employees don't want to ask because it would be so awkward if they're wrong about the candidate being a different person from who shows up for the job?
All of this could be mitigated with in person interviews, but I’m forced to hire abroad for cost.
I wish interviews were like this, instead most I've found are either trying to read the interviewer's mind on how to approach a vague situation and answer the way they want or have to reimplenent a full library in 30 min without any resource available that normally you'd look up, solve in minutes and move on.
I wish more took your path and literally just tested for actual industry experience: general architecture, asking questions when the situation is unclear and explaining unexpected/interesting findings from a previous project. And anyway, if they end up actually being a fraud, get rid of them after the initial probation time is up.
I study hours every day for many years now. I know many complex systems however studying algorithms bore me to tears.
I've built HPC clusters, k8s clusters, Custom DL method, custom high performance file system, low level complex image analysis algorithms, firmware, UIs, custom OS work.
I've done a lot of stuff because I can't help wanting to learn it. But I fail even basic leetcode questions.
Am I a bad engineer?
There seems to be no way for me to show my abilities to companies other than passing a leetcode but at the same time stopping learning DL methods to learn leetcode feels painful. I only want to learn the systems that create the most value for a company.
I imagine if you interviewed me you would think I wrote a fraudulent resume. Not sure how I am supposed to convince someone otherwise though. Perhaps I've been dumb in not working on code that can be seen outside of a company.
Naturally it's a "low bar" but it's an awfully low bar for a job that isn't entry level.
I'm a bit reminded of a chef who claims he can cook pasta, sushi, and pastries, yet when asked to fry an egg says that it's not fair to expect him to do so on the spot. A chef who can't cook when needed is about as useful as an engineer who can't code when needed.
I do these things because I see a place in the company to grow value and do it.
I can write almost all basic coding and my true skill set is in custom complex DS pipelines at scale.
Also, re technical questions, I don't think anyone is saying that you can't ask any technical questions whatsoever, I think the concern is about giving people abstract, theoretical CS problems that will never actually come up on the job, on the very iffy assumption that their performance while being asked to dance for a committee in high-pressure job interview situation is going to be reflective of their actual skills. (And more broadly, that a good programmer must be quick on their feet with schoolboy-style CS puzzles that are basically irrelevant to most roles.)
I think de-facto fulltime jobs i've seen end up kind of like that - if you don't work out in X months, nobody will shed a tear about kicking you out - but this is still perceived as a more expensive operation that should be avoided by having a "proper" interview process. On which nobody can ever agree, but that's the sentiment.
Both parties can terminate the employment agreement within a week for (usually) the fist three month of employment.
No stigma involved.
When I do interviews (probably limited compared to you but some) I do it like I wish someone would interview me.
I focused purely on curiosity. how many things disparate things are they interested in, the things that overlap with my knowledge I probe deep. I believe in Einsteins quote.
"I have no special talents. I am only passionately curious."
If someone knows about how RDMA and GPU MIGS work they are probably pretty damn interested in how HPC clusters function. More importantly can they compress this information and explain it so that a non technical person could understand?
There are so many endless number of questions I could ask someone to prob their knowledge of a technical field it kind of upsets me that most of the time people ask the most shallow of questions.
I believe this is because most people actually study their fields a very limited amount, because most people are honestly not truly interested in what they do.
The biggest implication of this is that I may be able to tell if someone has this trait but I understand that the majority of people could not as they literally don't know the things they could ask.
Asking system designs of me if you aren't knowledgable of the field would probably be the easiest to see the complexity of systems I can build.
You have to build a repertoire of questions that defeat rote memorization, prove real experience, and show genuine ability to solve unseen problems...
They could remember the exact type signature of standard library functions.
They could define a Monad instance from memory at the speed they could type.
But you couldn't ask them a single question outside of what was presented at lectures.
They couldn't solve a single assignment. They were stuck on rote learning.
I'd blame their educational system, because it was quite consistent (sample size = 3).
The classical example: Teacher says X, the whole class repeats in choir X.
Leetcode et al. are just testing rote memory, there's no need to have candidates actually type out solutions its a waste of time. So long as they can articulate what solution they would use, why, and what other solutions they considered that's all you really need to be concerned with.
Do you struggle with work projects that feel boring or abstract?
Was originally building firmware and hardware for previous company while testing DS on their systems. They liked my work and I switched to DS and ended up a team lead.
Honestly have never had a take home exercise but would love it if I could. I basically make my own work, if I don't get work I will build other projects for them and try to sell it to the company.
I normally make good value projects and can sell it. It's how I went from hardware to DS lead of a team in a few years.
The take home exercise/work sample companies are out there. Here’s a list I helped contribute to. https://github.com/poteto/hiring-without-whiteboards
Just because it’s such a classic topic in CS, that doesn’t mean I need to remember it after decades of seeing it in uni.
What good is someone who can code the best algorithm but cannot understand the business? Unless you are working in the top 1% of the companies out there (where you may have the luxury to invent new ways of doing things), for the rest of us our main skill is: to solve business problems with zero or minimal good enough code. We (99% of the tech companies) don’t need a Messi, just an average Joe.
We don’t need you to be the best engineer using the most cutting edge tools that a YouTuber told you about. We need you to be a good colleague who people can trust, can communicate with, and isn’t going to cause a ton of stress for others through maverick behaviour or major levels of incompetence.
We were hiring for a C++ role though. I can imagine people used to higher level languages might find the very need to deal with pointers confusing
> I can imagine people used to higher level languages might find the very need to deal with pointers confusing
I'm confused by this remark. What do you mean? Is this meant to be condescending? What do you use pointers for in your C++ project that references cannot do in a higher level language?That is my hypothesis on why some people might consider inverting a binary tree an unrealistically complex problem. If you are not able to solve this problem without thinking too much you are probably not being able to write working C++ at all, but might still be able to even solve real business issues with, say, Javascript.
The business-only guy won't see the limitations and possibilities of code, while the programming-only guy won't see where the business side is flexible or where it is rigid, thus the optimal solutions can take a long time to find.
The business side might also be prone to specifying X-Y problems, which can be very difficult to spot if you're not into the business side.
At least in my experience, having significant domain experience in our dev department has been a super power of sorts, allowing us to punch well above our weight.
edit: I didn't have any domain knowledge when I started, but I was willing to learn.
Sometimes the business only has a high level goal in mind and it is up to the developer to invent a solution, keeping in mind how the end user will interact with the system in their day to day work, and foresee potential issues in how the changes would impact other systems and business rules and so on. The developer has low level visibility of systems that are not readily apparent to business people. It is necessary for us developers to hide some of that complexity from the business. Some backends are very complex webs of business rules.
The developer needs to have some bigger understanding of it all, otherwise they are reduced to a code monkey who implements spoon-fed acceptance criteria from Jira tickets, makes stupid mistakes that break other functionality and systems, and won't recognize various issues that were not apparent to the business person who originally wrote the AC's. Basically we need developers to collaborate with the business to design and develop solutions properly.
Why is the former so much more important than the latter, for 99.9% of programmer interviews?
(protip: it isn't, but I guarantee you will get a "no hire" far more often if you need a hint. It's just nerd in-group hazing, coupled with a big helping of cargo culting Google.)
Go solve a hard puzzle while talking on the phone. Does it make you better or worse? I'm just saying...if someone needs a hint for the trick, but then, post-reveal, bangs out the code under the same circumstances, isn't that telling you something interesting?
What people do instead is spray and pray CVs and hope that hiring manger uncover their brilliance under surface level incompetence.
I have no sympathy for people posting "jobless for months and reject coding test interview" in one sentence.
(hint: it rhymes with "mesmerization")
You can't trust them to figure it out under pressure of securing a job, but that doesn't mean they couldn't do it when their job is secure.
It’s certainly not being forced to recall the implementation details of a trie within 30 minutes, when I haven’t seen one in 5 years, unable to reference any docs or knowledge base, or use Google, knowing that if I fail I will remain unemployed.
You might as well be quizzing people on Calculus. That's another thing you study in a freshman-year CS course, so you must remember it, right?
(I'll say this: I've had far more practical occasion to use Calculus than binary tree manipulation in the 20+ years of my professional coding career. Particularly with AI.)
What will you say when you fail an interview over an "introductory data structures" problem?
The lack of humility makes me wonder...
I knew the solution to the first was the min-cut algorithm, but no way I could code it straight up and down let alone in ~30min. Plus these leetcode-esque questions always come with millions of test cases making the need to be 100% exact and double check the constraints.
I even had my data structures/algo book nearby on the shelf that I could've used to cheat this but that would've been a new low for me, especially considering this was for a "mid"-level position working on APIs/JS/SQL.
I can understand if you want me to be an algorithm specialist but for web development for what im assuming would be a run of the mil SaaS application... this is absurd.
The real question: If the tables were turned, could the interviewers pass their own tests? Probably not.
There is just so much misrepresentation out there it's insane.
But you don't ask them "swap the left and right field in each node". You ask them to "invert a binary tree".
Why. The Fuck. Does "invert a binary tree" mean "swapping the left and right fields"?
To me, it means inverting the relationship so child points to parent, which would end up badly-defined, there would be no way to uniquely specify a root. But if you don't know that knowledge, of what the jargon means, you're screwed. So that's what you're testing for.
In real life they would either give you an example so you can better understand the requirements or expect you to come up with an example yourself and ask clarifying questions until you're on the same page as to what "invert" means.
We don't even know what defines a "good engineer" come yearly performance review time, but claiming that it's the knowledge of graphs and recursion seems rather suspect.
Inb4 "obviously there's so much more to being a great engineer," but then why are we not testing for that iceberg, instead just scratching the surface of "CS fundamentals?" And in fact, how many times does a great engineer need to prove that they understand fundamentals? Forcing engineers to go through that every time they want to switch jobs is inefficient and, frankly, disrespectful.
> X and Y are fundamental CS concepts
In 2024, the list of X and Y are fucking HUGE. I would like to see OP sweat in an interview on some "fundamental CS concepts" they have not used in the last 5-10 years. The lack of humility in some of these comment is simply stunning.I worked with a guy who was absolutely a first class engineer. Very well paid; I guess about 300K USD. He had almost no experience in C++ for more than 20 years (mostly Java and C# for last 10 years). During a discussion, I mentioned that the original C++ std::map used a red-black tree. He was well-surprised that lookup was not O(1); instead: O(log2(n)). (My point: He knows about red-black trees, but was surprised the original C++ foundation library did not include a hash map with O(1) lookup!) Really: He would have failed an interview from this person based upon "fundamental CS concepts". Any software engineer, no matter how smart or experienced, has some weak spots in their fundamentals.
If you're hiring a C++ expert, then yes, not knowing the difference between map and unordered_map would likely be a disqualifying condition. We are not talking about C++ expert interview though.
Typical interview loop includes more than 1 question.
> how many times does a great engineer need to prove that they understand fundamentals?
This is a logical fallacy. The interviewer doesn't know if the interviewee is a great engineer or not, that's the whole point of doing an interview in the first place.
> Typical interview loop includes more than 1 question.
In my interview experiences, programming problems are "fail fast". If you fail one, you fail the whole interview. There is no "round 2".I think the most underrated skill is to know what you don’t know. To know your limits. I know that graphs exist and that they are useful in certain circumstances. I don’t know how to implement the associated algorithms. Same goes for almost any other important topic like virtual memory management , security, performance, etc.
In my experience, it's much better to give them a problem to solve and have them walk you through the thought process as they try to figure it out. Even if they fail to solve it, the actual process gives you insight into how they think and how they'll perform on the actual job.
Unless your job regularly involves inverting binary trees, that's not a great question because it's basically asking if you've encountered something before. If you have encountered it, you likely know one a half dozen ways to solve it and everyone will give you roughly the same answer. If you've never encountered it, then you just fail even if are a good fit for the job. It doesn't reveal anything except you know one specific thing. In most cases, engineering jobs are problem solving for unique instances that arise day-to-day; a good interview should reflect that day-to-day reality more.
If my boss messaged me right now and handed me something that required a similar data structure change, I'd stumble on it for about an hour or two until I understood that actual problem.
I'm sure I could "invert" a binary tree if I had to but it'd probably take a far more than reasonable time as that just isn't something I'm dealing with nor most rank and file devs.
Less than 5% of our work involves understanding computer science and math, and for those 5% I only need a couple people to really, really get it. For everything else, it’s responding to business feature development, which means asking questions when things are vague, spotting inconsistencies in business asks, and breaking up a big problem into small problems that can be tested and built independently.
It is crazy how many people will fail a question that boils down to 'write a for loop' despite going to college for 4 years in CS.
If you want people able to do stuff off the bat, hire those who did MBO or HBO (in DE, HBO=Fachhochschule, but DE doesn't have an MBO equivalent I think: that would be Ausbildung level afaict, except MBO doesn't require you to have a job at the same time). In English, my HBO translated their name to "University of Applied Sciences"; my MBO did not give a English translation of the degree
Of course the best hires would be those willing to do both, if they're capable of a 6 year commitment.
Apprenticeships and on-the-job training is also an option, but nobody seems to be willing to do that anymore.
Check out Turing School. 40 hours/ week of classroom instruction, including 6 weeks of just programming before you get introduced to a framework.
I do a fairly simple encode/decode problem (run length encoding). I describe the basic encoding concept, provide a sample input that should have byte savings with any reasonable encoding, and have the candidate come up with what the output should be. There's lots of ways to do the encoding, mostly anything works (and I'm clear with the candidate about that)... I allocate about 15 minutes for this stage; I've got lots of hinting strategies to keep clients from getting stuck here... but if it's not clicking, I'll give them a simple format to move on.
Then the candidate writes the decoder; decoding is easier; some candidates get really stuck on the encoder and I'd rather have a code sample than a stuck candidate. Some of my worst candidates have already forgotten the encoded format that they just designed, and they write a decoder that might work on some other encoded input, I guess. Hopefully this takes 10 minutes or less, but if you can't write a loop in a loop, you might get pretty stuck. I don't care about the language used, it doesn't even need to be a real language, it just needs to be self consistent and reasonable; i/o comes as easiest for the candidate.
If we've got 15-20 minutes left, the candidate can work on the encoder. The encoder trips up a lot more people than the decoder; so I stopped having people work on that first.
There's plenty of options for discussion at any point. Could you make the format better for some cases, could you make it work in a resource constrained system.
The specific problem isn't really day-to-day work, but it's approachable and doesn't require much in the way of data structures or organization or prior domain knowledge. Some candidates express that they had fun doing my problem, and especially for junior candidates, if they've never done compression exercises, I hope it is an opportunity to see that compression isn't always magic; simple compression is approachable.
I'm happy to answer more questions via email in the profile.
Example input would be something like
55 55 55 55 20 32 56 3F
3F 3F 3F 42 42 61 61 61
Values chosen "randomly" by me, not intended to have any meaning. Pattern is more intentional, although this isn't my best pattern; I haven't run my interview in a long time, I had a nicer pattern I think, but can't remember it.
All of my input happens to be bytes less than or equal to 7F (although in the interview setting, I don't mention that unless asked...) and not equal to 00. If candidates don't understand hex bytes, that's fine, even if it makes me a bit sad inside. When they get to coding, I'll give them in() that returns the next value from the input, and out(...) that outputs the next value. Etc.To set people up, I tell them we're going to do some compression today with an encoding called run length encoding; in run length encoding we encode the length of a run and its value. (If they're familiar and ask questions, I tell them today, we're focused on bytes... you can do it with bits or larger than bytes, whatever, but keeping it focused on bytes makes the problem fit into a 45-minute interview slot)
There's lots of ways to make an RLE format, if you go with the basics like (count, value), or (value, count) and we have time to explore, I'm going to ask about the sequence 20 32 56; it doubles in size with that format, so what can you do to make that better while still doing a good job overall on our example stream. Lots of options, all of them are fine as long as you can decode it correctly. The basic one is fine too, if we don't have time to explore.
When using this problem, I feel like I give more thumbs up than others on my teams; but I feel really confident in my thumbs downs, and I feel I gave everyone a fair chance. I've had a few people do really poorly that just seemed nervous, and I tried to pivot towards more of a buddy problem / build their confidence / try to get them less nervous so they can do a bit better for the rest of their day.
They might be able to solve it perfectly if encountered in the wild, but if you see it in an interview part of you is thinking “there’s got to be a trick to this that I’m not seeing”
Colleges and universities are lowering standards and accepting anyone with a pulse to get their tuition dollars.
Overall interview is "write code to solve this puzzle." But first, do this very basic thing that is needed to solve the puzzle.
80% of candidates get hung up on the basic part of the interview and never even get to the point of looking at the rest of the problem. But of those that did, we got some great people.
Here's a log file of page accesses on our server. It's a CSV. The first column is the user, the second column is the page, and the third column is the load time for that page in milliseconds. We want to know what is the most common three page path access pattern on our site. By that I mean, if the user goes to pages A -> B -> C -> A -> B -> C the most common three page path for that user is "A -> B -> C".
user, page, load time
A, B, 500
A, C, 100
A, D, 50
B, C, 100
A, E, 200
B, A, 450
etc.
So for this first question you should give an answer in the form of "A -> B -> C with a count of N".We would have two files, one simple one that is possible to read through and calculate by hand, and one too long for that. The longer file has a "gotchya" where there's actually two paths that are tied for the highest frequency. I'd point out that they'd given an incomplete answer if they don't give all paths with the highest frequency.
The second part would be to calculate the slowest three page path using the load times.
In my opinion it's a pretty good way to filter out people that can't code at all. It's more or less a fancy fizzbuzz.
Btw, if you want CSV log files, look no further, and not all my data logs have timestamps either! :D The particular timestampless case I'm thinking of, I wanted to log pageload times for a particular service so it logs the URI (anonymized) and the loading time, though I think that's not csv but just space separated, one entry per line
I feel like this is a fun thought experiment, but instead of thinking about "gotchas" I would be more open to having a discussion about edge cases, etc... The connotation of gotchas just seems to be like a trap where if you hit one, you've failed the interview.
That's probably doable in like 5 lines of pandas/numpy; a straight forward o(n) task really. The hard part is getting it right without googling and debugging, but a good interviewer would help you out and listen to the idea.
Maybe using Pandas is cheating since it gives you all the tools you'd want but I'd argue it's the right tool for the task and you could then go on how you'd profile and improve the code if performance were a concern.
Yeah, that's what bugs me about this type of question... he might be looking for that specifically, or something that can scale to exabytes of data (so some sort of map/reduce thing). I'd probably produce something like this _in an actual interview scenario_:
users = {}
count = 0
for line in open('input.txt'):
count += 1
if count == 1:
continue
(user,page,load_time) = line.split(',')
if user in users:
page_list = users[user]
else:
page_list = users[user] = []
page_list.append(page.strip())
count = {}
max_count = 0
max_seq = None
for page_list in users.values():
if len(page_list) > 2:
for i in range(len(page_list) - 2):
seq = ''.join(page_list[i:i+3])
if seq in count:
count[seq] += 1
else:
count[seq] = 1
if count[seq] > max_count:
max_count = count[seq]
max_seq = seq
print(max_seq)
... and it would really depend on whether the interviewer just liked me personally whether he'd say, "yeah, that's reasonable" or rip it apart for using too much memory, taking too much time, etc...Anyway, I'd hate to be the person to claim there's a five liner, without providing some terrible code for future AIs to train on:
n = 3 # length of path
for user in (df := pd.read_csv(io.StringIO(input)))["user"].unique():
counter = Counter([seq for seq in zip(*[df[df["user"] == user]["page"][k:] for k in range(n)])])
equal_lengths = sum([1 for i in counter if counter[i] == counter.most_common(1)[0][1]])
print(f"most common path for user {user}: {counter.most_common(equal_lengths) if len(counter) > 0 else 'input too short'}")I think this isn't quite right, though, he asked for "the most common three page path access pattern on our site", this seems to output the most common three page path per user.
(for the future AI to mull over)
import sys
unfinishedPaths = {} # [user] = [path1, path2, ...] = [[page1, page2], [page1]]
finishedPaths = {} # [path] = count
for line in sys.stdin:
user = line.split(',')[0].strip()
page = line.split(',')[1].strip()
if user not in unfinishedPaths:
unfinishedPaths[user] = []
deleteIndex = []
for pathindex, path in enumerate(unfinishedPaths[user]):
path.append(page)
if len(path) == 3:
deleteIndex.append(pathindex)
for pathindex in deleteIndex:
serializedPath = ' -> '.join(unfinishedPaths[user][pathindex])
if serializedPath in finishedPaths:
finishedPaths[serializedPath] += 1
else:
finishedPaths[serializedPath] = 1
del unfinishedPaths[user][pathindex]
unfinishedPaths[user].append([page])
for k in sorted(finishedPaths, key=lambda x: finishedPaths[x], reverse=True):
print(str(k) + ' with a count of ' + str(finishedPaths[k]))
Not tested properly because no expected output is given, but from concatenating your sample data a few times and introducing a third person, the output looks plausible. And I just noticed I failed because it says top 3, not just print all in order (guess I expect the user to use "| head -3" since it's a command-line program).I needed to look up the parameter/argument that turns out to be called "key" for sorted() so I didn't do it all by heart (used html docs on the local filesystem for that, no web search or LLM), and I had one bout of confusion where I thought I needed to have another for loop inside of the "for pathindex, path in ..." (thinking it was "for pathsindex, paths in", note the plural). Not sure I'd have figured that one out with interview stress.
This is definitely trickier than fizzbuzz or similar. Would budget at least 20 minutes for a great candidate having bad nerves and bad luck, which makes it fairly long given that you have follow-up questions and probably also want to get to other topics like team fit and compensation expectations at some point
edit: wait, now I need to know: did I get hired?
Major:
1. Sorting finishedPaths is unnecessary given it only asks for the most frequent one (not the top 3 btw)
2. Deleting from the middle of the unfinishedPaths list is slow because it needs to shift the subsequent elements
3. You're storing effectively the same information 3 times in unfinishedPaths ([A, B, C], [B, C], [C])
Minor:
1. line.split is called twice
2. Way too many repeated dict lookups that could be easily avoided (in particular the 'if key (not) in dict: do_something(dict[key])' stuff should be done using dict.get and dict.setdefault instead)
3. deleteIndex doesn't need to be a list, it's always at most 1 element
I realized at least the double-calling of line.split while writing the second instance, but figured I'm in an interview (not a take-home where you polish it before handing in) and this is more about getting a working solution (fairly quickly, since there are more questions and topics and most interviews are 1h) and from there the interviewer will steer towards what issues they care about. But then I never had to do live coding in an interview, so perhaps I'm wrong? Or overoptimizing what would take a handful of seconds to improve
That only ever one user path will hit length==3 at a time is an insight I hadn't realized, that's from minor point #3 but I guess it also shows up in major points #2 and #3 because it means you can design the whole thing differently -- each user having a rolling buffer of 3 elements and a pointer, perhaps. (I guess this is the sort of conversation to have with the interviewer)
Defaultdict, yeah I know of it, I don't remember the API by heart so I don't use it. Not sure the advantage is worth it but yep it would look cleaner
Got curious about the performance now. Downloading 1M lines of my web server logs and formatting it so that IPaddr=user and URI=page (size is now 65MB), the code runs in 3.1 seconds. I'm not displeased with 322k lines/sec for a quick/naive solution in cpython I must say. One might argue that for an average webshop, more engineering time would just be wasted :) but of course a better solution would be better
Finally, I was going to ask what you meant with major point #1 since the task does say top 3 but then I read it one more time and...... right. I should have seen that!
As for that major point though, would you rather see a solution that does not scale to N results? Like, now it can give the top 3 paths but also the top N, whereas a faster solution that keeps a separate variable for the top entry cannot do that (or it needs to keep a list, but then there's more complexity and more O(n) operations). I'm not sure I agree that sorting is not a valid trade-off given the information at hand, that is, not having specified it needs to work realtime on a billion rows, for example. (Checking just now to quantify the time it takes: sorting is about 5% of the time on this 1M lines data sample.)
For anyone curious, the top results from my access logs are
/ -> / -> / with a count of 6120
/robots.txt -> /robots.txt -> /robots.txt with a count of 4459
/ -> /404.html -> / with a count of 4300You need the list regardless, just do `max` instead of `sort` at the end, which is O(N) rather than O(N log N). Likewise, returning top 3 elements can still be done in O(N) without sorting (with heapq.nlargest or similar), although I agree that you probably shouldn't expect most interviewees to know about this.
As for the rest, as I've said, it depends on the candidate level. From a junior it's fine as-is, although I'd still want them to be able to fix at least some of those issues once I point them out. I'd expect a senior to be able to write a cleaner solution on their own, or at most with minimal prompting (eg "Can you optimize this?")
FYI, defaultdict and setdefault is not the same thing.
d = defaultdict(list)
d[key].append(value)
vs d = {}
d.setdefault(key, []).append(value)
useful when you only want the "default" behavior in one piece of code but not others > / -> / -> / with a count of 6120
> /robots.txt -> /robots.txt -> /robots.txt with a count of 4459
LOLThis is exactly what irritates us about these questions. There's no possible answer that will ever be correct "enough".
Here's my solution in TS.
const parseLog = (input: string) => {
const userToHistory: {[user: string]: string[] } = {}
const pageListToFrequencyCount: { [pages: string]: number } = {}
for (const [user, page, ] of input.trim().split("\n").map(row => row.split(", "))) {
userToHistory[user] = (userToHistory[user] ?? []).concat(page);
if (userToHistory[user].length >= 3) {
const path = userToHistory[user].slice(-3).join(" -> ")
pageListToFrequencyCount[path] = (pageListToFrequencyCount[path] ?? 0) + 1;
}
}
return Object.entries(pageListToFrequencyCount).sort(([, a], [, b]) => a - b);
}
It could be slow on large log files because it keeps the whole log in memory. You could speed it up significantly by doing a `.shift()` at the point when you `.slice(-3)` so that you only track the last 3 pages for any user.I still struggle with this. I don’t find it a blocker, though. The bottleneck is usually to understand and parse business requirements. If you know about good code practices as well, then the least of your problems is to know whether you can use ‘in’ or ‘[]’, ‘var’ or ‘let’, ‘foreach’ or ‘for’, ‘def’ or ‘fun’, etc.
Unless youre talking about using tail / grepping some linux logs I've never done this once let alone on a daily basis...
But if needed frankly I would look up any needed regex / pattern for this on the job and on the fly. The topic of "log parsing" is massive anyway. What types of logs? Is this something like request logs from NGINX? Custom stderr / stdout from some cron job? Some sort of horrible XML based system dump? I could go on and on...
When you know the solution already of course its obvious.
The company has very excellent coding and test interviews but the codebase is just shit.
It is also possible that this is because of all the fraud developers they've hired. A "company" is not a person. It doesn't produce shit code itself. While management can also be a source of bad codebase, it is usually the developers.
Obviously it is a different story when management hires good developers but can't maintain a good workplace culture, apply unnecessary time pressure, don't pay them enough etc.
Plot twist: the devs who are good at leet code quizzes are frequently crap product developers on teams.
If you can't, or can just barely, complete fizzbuzz in the allowed interview time with a lot of coaching in your language of choice, then you definitely aren't ready to work as a SWE. If you breeze through all my extra sections in half the time, then you're great. Partway through, and you're probably a decent junior to senior engineer.
Which would be how long? I haven't practiced it specifically so it would be a good test I guess
Thanks for the answer!
I use it all the time, especially when logging the progress of a long running process at every x records.
I'd also note that one of the reasons I prefer doing such things in person is (in addition to not taking up too much of the candidate's free time) that if they misinterpreted something or are going off in the wrong direction or something, I can correct them right away instead of letting them waste a bunch of time. This also lets you see how they tend to interpret instructions and take criticism and corrections.
It's also IMO a failure mode for a candidate to get a simple task that should be done in 5-20 lines of simple code and try to build some over-complex super-modular and extensible thing for it, and resisting correction on it not needing to be that over-engineered.
That's still different compare to interview though because you need to explain, code, solve a question under 20-30 minutes. It's not hard if given enough time. But most people would need to practice some questions before going into those questions so that they can solve those under the time limit.
The amount of insane answers I've seen to that one alone...
Then if they pass, I test proficiency by having them loop over the dict and update each value in-place.
So, I wouldn’t pass these kind of interviews. In over a decade I’m never being asked these kind of questions though (I have done take home assignments and leetcode, but always with google opened)
I have used Python to solve average business problems, yet I cannot produce non trivial code without looking at the documentation. Same for the other dozen programming languages I have used in the past.
hello = {1:1, 2:2, 3:3}
is about as trivial an ask as someone can make.That can usually be solved by a quick read of the reference documentation (2-5 mn?).
For my part, I've written enough python that I doubt the literal syntax will ever be far from my fingers.
I needed a small perl script recently (perl 5’s feature set & stability plus availability in the environment made it the right fit) and realized after 15+ years of no perl much of the specific syntax was fuzzy to outright gone from my head even though I’d contributed to large perl projects for years.
Python work is much more recent, but I’d bet I would accidentally mix in some JS or even PHP syntax doing the dictionary assignment, at least w/o a cursory lookup. I’d like to think it’d come through that I know what a dict is and what it means to set one up and operate over it, but who knows, I might be interviewing with someone who is evaluating skill on the basis of immediacy of syntactic recall.
So, the equivalent of creating a dictionary, yeah, sure. But there's loads and loads of stuff that I only use maybe once a week (and someone else maybe uses daily) and that I'd have to awkwardly Google (I use Kagi btw) even during an interview.
Your interviewer is probably also questioning if it's (a, e) or (e, a), but you passed the fraud filter.
If you have just started your career and are full invested in 1 or 2 programming languages, sure this may sound alien to you.
If your interviewer finds that problematic, well, that's on them.
In interviews I've done, we only looked for culture fit because the technical part was a coding assignment they had already done. Honestly too big an assignment since it's uncompensated (not my decision), but to my surprise nobody turned it down -- and everyone got it wrong. Only n=3 or n=4 iirc but those applying for a coding position could not loop through a JSON-lines file too big to fit in RAM (each line was a ~1kb JSON object, but there's lots of lines) and sum some column from each JSON object into a total value. The solutions all worked to some degree, but they all chopped up the file, loaded the first chunk into RAM, and gave an answer for that partial dataset only.
EDIT: Another thing: about 80% of the candidates I interview wouldn't be able to pass our Product Manager SQL interview. It's basic shit, but not as basic as the stuff I ask. All the PMs in my current job have better skill than 90% of the backend engineers I interviewed in the last two years. Resume fraud is rampant.
I might recommend, when asking this question, to give the definition with a few example inputs and outputs. That should avoid these types of issues where people are perfectly capable of coding the requested algorithm but aren't mathematicians / toy problem experts
You know you can ask the interviewer about this, right?
We do give the definition of Fibonacci, and we steer the candidate in the right direction if they make math errors.
Factorial is just a for loop and with Fibonacci you might want to talk a bit about recursion and caching. That's it.
They're dumb.
People still fail.
OP here. As I said, "I give the solution away".
I basically tell the definition/formula at the beginning, and give examples and test cases for you to check your results against.
I also help people along, like I would in pair programming.
A lot of people still fail. Some with allegedly 25 years of experience. Luckily this is at the beginning, so I don't have to spend the whole interview with them.
Your what now? (Well, unless you're a database company.)
In all my time as a PM I've never had to know, let alone prove competence in, my SQL knowledge.
In all my time as a PM, the vast majority of my "SQL" usage is crafting slightly more advanced filters and queries in Jira or similar.
There are analytics platforms, telemetry, metrics, you name it.
I've worked as a PM on fairly complex products (integrating and manipulating healthcare data, managed platforms atop Kubernetes, etc.) and never needed SQL knowledge.
As a general rule in terms of how I allocate my PMs and their efforts now? "PMs have far more valuable things to be doing than lovingly crafting handwritten SQL for god-knows-what reasons."
But it's super basic SQL. Select and joins, mostly.
Developers with several years of experience in their resume still fail it.