Am I missing something?
CPU, RAM storage are incredibly cheap nowadays. But so much can be going on at a time that _certainty_ is now what's pricey.
I attribute a lot of my performance to initially targeting very low powered hardware. I ran it on a raspberry pi cluster until the index was around 1 million documents (I think). Having those limitations early on excluded a lot of design mistakes that would have been a lot harder to correct later.
Serving search requests is a embarrassingly parallel operation that could be very easily load balanced to different instances. If so Max Schrems hectored every other search engine out of Europe, and Ursula von der Leyen, taking pity on my living room search operation, sent me a bunch of money, I'd probably be able to deal with 1000x the traffic about as fast as I'd be able to lease space in a datacenter.
I am a bit choked in terms of how big I can grow my index, though. I could probably double, maybe even a couple of times, but not too many times if I can't find serious optimizations.
- strong read-after-write consistency (or their users could receive inaccurate information)
- as-high-as-possible availability (because they lose money whenever the site goes down)
- low-latency operations that are difficult to optimize--such as recommendations, complex full-text search, or large graph traversal
And engineering teams don't notice that their solutions are suboptimal until their product hits scale.
A lot of engineering teams scale prematurely. Maybe their complicated engineering playground will indeed scale to 100M users, at the expense of being extremely complex and costly from an infrastructure point of view, where as a more "boring" stack may not scale to 100M but will handle up to 10M in a much simpler & cheaper way.
The more complex solution makes sense if you're sure you'll hit 100M, but megalomaniac founders' dreams aside, this rarely happens and they end up dealing with the unnecessary complexity & expense of a stack whose scalability isn't actually required.
A lot of startups' intentions is to look "serious". If everyone has an engineering playground full of microservices with 5 different programming languages all hosted on AWS, you have to be the same if you want to join the "big boys" club.
Sometimes startups don't even care about solving the business problem. The objective is instead to get the "startup CEO" lifestyle for the connections and reputation regardless of whether you end up solving your business problem in the end. Same for a lot of engineers - they'll happily join and build the engineering playground while enjoying their salary regardless of whether it's the right option from an engineering point of view. If anything, an engineering playground gives you the opportunity to manage lots of people and solve (self-inflicted) problems that will look good on your resume in a way that a boring solution with a handful of engineers won't.
On the other hand, if you live and die by solving the business problem the most cost-effective way, all the "scalability" goes out the window and a simple stack on top of bare-metal makes much more sense, because the only thing that pays you is solving the business problem (in the cheapest way possible) as opposed to bragging about your self-inflicted technical complexity at AWS conferences.
There seems to be many people who simply fall for the marketing and think that kubernetes and hundreds of microservices are somehow necessary to run a small to medium sized web service. They can't comprehend how a technology stack with a 20 year track record could be even nearly as stable and reliable as a grab bag of 2 weeks old javascript microframeworks. It's absolutely inconceivable how something that performed well a decade ago on hardware that was a hundred times slower could keep up on modern servers.
I have 4 load balanced VMs and literally just need to pull the logs into a central location sorted by time. There's a shocking amount of complexity around all the popular tools in this space.
The neat thing about simple tools is that they're often relatively simple to build.
RSYSLOG has existed for a long time and is probably the simplest solution for sending logs somewhere.
* 100k users updating their gas / electricity meter readings every second
My first question was what unit was the meter reading in, he said it was the current total (called a totalizer in the biz). I'd worked for a Flow Metering company before. So I said:
"There's no need to upload the total every second reduce the frequency and the problem becomes easier"
Didn't accept that and I said ok a few optimized servers behind a HAProxy load balancer would the trick as the processing is very simple here. The database is the harder part as you can't have a simple approach here is 100k requests / second is going to cause contention for most DB's. You'd need to partition the database in a clever way to reduce locking / contention.
This answer was unacceptable he ignored it and asked me to start building some insane processing system where you use Azure Functions behind an api gateway to process every request into a queue. Then you have a scheduled job in Azure of which the name has changed by now. This job looks at the queue in batches of say up to 1000 records at a time and writes these in bulk into the DB Once it's finished it then immediately reinvokes itself to create a "job loop".
You would create multiple of the "job loops" that run in parallel. Then you would need to scale up these jobs to drain the queue at a rate that meant we could process 100k requests per second.
You would also need something to handle errors, for example requests that are broken somehow. These would go into a broken request queue and could be retried by the system automatically a certain number of times if they went over the number of retries then they would go into a dead letter queue to be looked at manually.
I'm pretty sure this was the actual process he would take in designing this type of system. Use every option the cloud gives him to make some overly complex thing which is probably quite unreliable. I also suspect that's the reason I wasn't hired as I was just looking at the thing sensibly rather than in the "cloud native" way.
He doesn't work for that company anymore and is now back to a normal (non-senior) software developer.
In general, this comment is exactly the sort of discussion I'd be interested in hearing if I interviewed people.
To give him the benefit of the doubt maybe it was an artificial question and the real question is how do you make a large scale "cloud native" system say something like Netflix / Facebook.
Of course the problem with those systems is that it's very difficult to explain it in detail of every part in an interview so his example was just a small part of that system. Also most companies do not require that level of scalability but I concede they are fun to build just not very fun to maintain/operate.
I have no doubt that the architecture I described was something he had personally built as part of what he was doing at the company. Me questioning that was probably taken as an attack.
I never got the job or even a single word of feedback.