Agreed. ‘sendOnce’ implies something very specific in most async settings and, in this interview question, is being used to mean something rather different.
Then we send all traffic to proxy.service instead of canon.service, and implement queuing, rate limiting, caching, etc on the proxy.
This shields the weak part of the system from the clients.
I guess I hadn't considered the existence of generic TCP proxies And I'd probably have concerns around how much latency it would introduce in an environment where we have some requirements on the rate of data collection.
That said our service Does act as a kind of proxy for a few protocols.
We’d get on calls with them and they’d be like “you can’t do multithreading!” we eventually parsed out that what they literally meant was that we could only make a single request to their API at a time. We’d had to integrate with them, and they weren’t going to fix it on their side.
(Our solve ended being a lot more complicated than this, as we had multiple processes across multiple machines that were potentially making concurrent requests.)
Far easier than the original single threaded solution - and has fault tolerance baked in cause you can run it on multiple clients
Redis is just another SPOF, and so is Postgres without fiddly third party extensions (that are pretty unreliable in practice, IME). I'm talking about something truly distributed.
Plus it's just good practice that I'd want to be following anyway. Once you get in the habit of doing it it doesn't really cost much to design the dataflow right up-front, and it can save you from getting trapped down the line when it's much harder to fix things. Especially for an interview-type situation, why not design it right?
To be honest, for me, in an interview-type situation, if you insist that Redis is the problem in that scenario - you would have failed the interview (the interview is never one-way, interviewers can fail it too).
If you literally just drop in etcd or Zookeeper rather than Redis and then develop in the same way then I'd say there's no additional cost to doing that. (I mean sure if you dig hard enough you can always find a way in which solution A is worse than solution B - e.g. most things have worse latency than Redis - but in this scenario the latency of the external API is going to make that irrelevant). Of course if you're just running those in single-node mode and developing against them without thinking about the distributed issues then you've still got plenty of ways to shoot yourself in the foot, but it's a small step in the right direction.
Developing more fully distributed from day 1 requires discipline that takes time to learn, but I'm not convinced that it's actually slower - I'd compare it to e.g. using a strongly typed language, where initially you spend a lot of time bouncing off the guardrails, but over time you adapt yourself and can be productive very rapidly on new projects.
> To be honest, for me, in an interview-type situation, if you insist that Redis is the problem in that scenario - you would have failed the interview (the interview is never one-way, interviewers can fail it too).
Interesting - to me Redis in a system design is very often a case of over-architecting. It's easy to use and programmers enjoy working with it, but very often it isn't letting you do anything you couldn't do without it, and while it can speed things up, I see a lot of cases where the thing it speeds up is something that was already fast enough.
1 etcd pod doesn't give you "no SPOF", you need 3, and then you need them on multiple VMs (or physical machines if you're not on the cloud/not in k8s), and then the cluster needs to be multi-AZ, and if you're really serious about the "no spof" that may mean geo-redundancy too... come on, just the deployment costs alone are significant.
But if you're in the habit of using HA-capable systems then whatever you have to hand will be HA-capable, and so there won't really be any additional cost to using that.
And again, I think there's a real antipattern where people take a single-server application and then claim they've made it fault tolerant by making it run on multiple hosts, but it's still relying on a single DB server. In my experience that doesn't actually improve reliability any (at least not if you've got a good deployment process for your single-server application) and it complicates your architecture to no real benefit. (Indeed, frankly, I think a lot of developers reach for an external database because they have no other idea how to store data from their application, when using embedded sqlite/hbase or - shudder! - the local filesystem, would let them use much simpler architecture and not really reduce the actual reliability of the system).
> 1 etcd pod doesn't give you "no SPOF"
No, but it gives you a clear path to removing your SPOF when the need arises. Which is much harder if you've built your system on Redis.
> using HA-capable systems
All the systems mentioned in this discussion are HA-capable (even Redis, for some usecases, is perfectly HA-capable; typically for a distributed lock it isn't appropriate, but then again, for the scenario under discussion, you don't need a perfectly safe distributed lock so it would work just fine).
The more interesting question is not whether a system is HA-capable, it's whether the system is appropriate for the job that's required of it (given said system weaknesses & strengths, plus the specific job needs). And my argument was that both Redis and Postgres were fine, for the job that was described. In an interview situation I want to see that my interviewer is capable of thinking through particular situations and having a good honest debate about strengths and weaknesses of a proposed solution _for a proposed problem_ - not just pushing their preferred solution as dogma. In many business scenarios it's fine & correct to architect systems as "HA by default" but in interview situations we're debating hypotheticals, and I am going to judge you based on the hypothetical at hand, not based on your day-to-day job, because I don't know what your day-to-day job is (and it's not what's being discussed).
It really isn't, outside of some stretched definition. Nor is Postgres without third-party extensions (that come with significant issues in my experience).
> The more interesting question is not whether a system is HA-capable, it's whether the system is appropriate for the job that's required of it (given said system weaknesses & strengths, plus the specific job needs).
I used to believe this kind of thing, but I've come around to the opposite; actually rather than carefully considering the strengths and weaknesses of any given system in the context of a given job, it's a lot more efficient to have some simple heuristics that are easy to evaluate for which systems are good or bad, and avoid even considering bad systems. Of course occasionally you do need to dive into a full evaluation and pick your poison, but if a task doesn't have very specific requirements you avoid a lot of headache by just dismissing most of the possibilities out of hand.
> And my argument was that both Redis and Postgres were fine, for the job that was described.
But they're not contributing anything to the job that's described! Adding an extra moving part to the system that doesn't actually achieve anything is a much worse error than choosing the wrong system IMO.
> Adding an extra moving part
I specifically mentioned I considered those as good solutions for the problem at hand only if you already have them/ don't need to add them, that's their strength (lots of systems already use Redis or a SQL database, e.g. Postgres - but anything really would work just fine for the task at hand).
What happens when the unstoppable force meets the immovable object? The unstoppable force works over the weekend to implement a store-and-forward solution.
When the client(s) can send more work than the server can handle there are three options:
1 - Do nothing; server drops requests.
2 - Server notifies the clients (429 in HTTP) and client backs-off (exponential, jitter).
3 - Put the client requests in a queue.
Interview question/solution does 2 in a poor way (just adding a pause), it's part of the client and does 3 in the client, when usually this is done in an intermediate component (RMQ/Kafka/Redis/Db/whatever).It’s not the ability to communicate effectively that’s at play here, it’s your ability to read your interviewer’s thoughts. Sure thing, if you work with stakeholders, you need some of that as well, but you typically can iterate with them as needed, whereas you have a single shot in the interview.
Plenty of times, at the end of the interview, I do have a better mental picture of the problem and can come up with a way better solution, but “hey, 1h has already passed so get the fuck out of here. Next!”
I've tried that approach in a couple of interviews and, no surprise, I did not get those jobs because interviewers really seem to hate it if you dare step outside the little cocoon of leet they've constructed.
It doesn't matter why you can't fix the server, you can't fix the server. It doesn't matter why it can only handle one request at a time, it can only handle one request at a time. That information doesn't change the solution to the task. Make up whatever you please - maybe it's a third party system so you can't access the code. Now try to come up with some useful questions that actually help you solve the task instead of wasting time asking pointless questions.
No wonder people reject you if this is how you approach interviews. Maybe it would make sense to ask these kinds of questions in a real work scenario, but it does not make sense in an interview where you are given a made up task. Just accept the constraints and solve the problem. You sound like a high school student being intentionally obtuse, like you came into the interview thinking you're too good to be evaluated that way or maybe you just don't have a clue how to solve the actual problem you're given so you try to stall to avoid having to admit that you can't do it.
That being said, the time spent on these questions should be minimal. It should go something like:
Candidate: It seems like we're solving the wrong problem. Why can't the server handle multiple simultaneous requests from a single client?
Interviewer: Its a 3rd-party server we don't control. (or even just a "because those are the artificial constraints for this puzzle" if the interviewer is feeling particularly lazy)
Candidate: Ah okay. <proceeds to work on the problem>I also think they're unnecessary in the first place. I wouldn't hold it against anyone in a scenario like you described but I also wouldn't expect it. Solving the task as it's posed is good enough.
I might just mention that my first instinct would be to fix the root of the problem instead of asking about it, I wouldn't expect the interviewer to have spent time world building for their made up task.
As it stands, we still don't know why the server was broken in this way and why they created a work around in the client instead of fixing the server.
what is the delay actually doing? does it actually introduce bugs into that backend? how do we check that?