How do you spot it? I've had candidates with stellar performance on system design interviews, and I've never had any suspect of them (and the resumé also seemed pretty strong in those cases). But now I'm curious if I missed something.
How do you spot it? I've had candidates with stellar performance on system design interviews, and I've never had any suspect of them (and the resumé also seemed pretty strong in those cases). But now I'm curious if I missed something.
The main tells are a mismatch between apparent practical experience and apparent knowledge of the design space. I always dig into a few specific technical aspects of the design as far as possible to see how deep the candidate can go.
For example, if there's a queue I'll ask what would happen if in production the queue starts to fill, and how to mitigate. Most candidates reflexively suggest making the queue bigger, some suggest adding more downstream capacity to drain faster. People with real experience will have encountered this problem and know that you have to spend time investigating to find the root cause of the backpressure, otherwise you could add capacity to random things without reducing the queue occupancy. Are your queue consumers getting throttled by some downstream service? Are your workers CPU bound? The "right" answer here is basically "we should figure out why the queue is filling up before changing anything", but very few candidates answer that way.
Another kind of mismatch is the red flag I listed, where candidates will basically list buzzwords or ideas regardless of whether or not they are related to the question I asked. For example if I ask a question where the user-facing API is very simple (ie, write opaque data to a log) then they start talking about the tradeoffs between graphql vs REST, that's a bad sign.
It’d be fine to just up the queue capacity if the input is just inconsistent/spikey but averages to the input, but under steady state (or steadily increasing growth) higher input, you must increase the output rate or you will outgrow whatever capacity the system has
This is the opposite of what "people with real experience" do, at least in production-down situations. The first step is to mitigate, the second step is to root cause. If there's some no-brainer step that has a chance of alleviating the issue while you root-cause, and is unlikely to make things worse, you should take it.
You and OP are actually agreeing with "the correct answer is it depends and now let's discuss context". This echoes my experience as interviewer too. It's a red flag when the candidate responds with "the correct answer". That's what OP is calling out.
I'm replying here because I get the impression you're looking for "the right answer" as you see it: "the first step is to mitigate then do root cause". You're right! But it also could be too adversarial.
Most interviewers are adversarial.
Let's be honest - if an interviewer wants you to pass a system design interview, they'll make it work. I see this with particular candidates all the time. If we want the person to make it through - we'll let them get through. If we don't want them to get through - no amount of correct and behaviorally appropriate answers are gonna make them get through.
Also making a queue bigger is a bad reflexive response to queues being full. I was very involved in the SEV review (postmortem) process at Facebook and witnessed lots of cases where bad situations were made much worse by misguided hasty responses.
I agree though, hasty response is the mark of a junior eng
Depends on the function of the queue. If it's a dead-letter queue, there's literally no downside besides a likely-trivial amount of cost. If the problem is upstream of the queue, and the queue is actively feeding well-functioning consumers, yeah, it could make things way worse. This is where being an experienced engineer comes in. Also having 2-person approval like you would for any code changes. Point still stands that you should take low-risk actions to mitigate if they're available to you before root-causing.
I take your word for it that they’re trying to trick you or something - it’s not my story. But it kinda comes off like you’re denigrating people for not knowing something.
The key is that the response should reflect the actual design we're discussing, it should be more specific than something that could be said in any interview.
The leetcode grinder would almost certainly ask clarifying questions since it’s heavily emphasized as part of the rubric in pretty much all system design interview prep material.
Of course, the non leetcode grinder might ask those questions, but even the social cues that you’re allowed to ask clarifying questions might not be understood.
In my experience, I’ve seen a lot of people say that systems design test the seniority of a candidate, but then stick to a rubric that the leetcode grinder has memorized and the senior engineer might miss several important points on despite having the experience because they didn’t understand all the unwritten rules of the song and dance.
But TLDR: the hollowness of each step of the approach does become apparent in my experience.
An interview system with a high false-negative rate is not seem as problematic if the false-positive rate is also low. The cost of not hiring a very good candidate is way lower than the cost of hiring a bad one (even considering the opportunity cost in most cases).