You Are Not Google (2017)
blog.bradfieldcs.com
blog.bradfieldcs.com
The former examples are all managed! That's amazing for scaling teams.
(SQL can be managed with, say, RDS. Sure. But it's not the same level of managed as Dynamo (or Firebase or something like that). It still requires maintenance and tuning and upkeep. Maybe that's fine for you (remember: the point of this article was to tell you to THINK, not to just ignore any big tech products that come out). But don't discount the advantage of true serverless.)
My goal is to be totally unable to SSH into everything that powers my app. I'm not saying that I want a stack where I don't have to. I'm saying that I literally cannot, even if I wanted to real bad. That's why serverless is the future; not because of the massive scale it enables, but because fuck maintenance, fuck operations, fuck worrying about buffer overflow bugs in OpenSSL, I'll pay Amazon $N/month to do that for me, all that matters is the product.
> I'll pay Amazon $N/month to do that for me
Until you pay Amazon $N/month to provide the service, and then another $M/month to a human to manage it for you.
In this case you're only shifting the complexity from "maintaining" to "orchestrating". "Maintaining" means you build (in a semi-automated way) once and most of your work is spent keeping the services running. In the latter, you spend most of your time building the "orchestration" and little time maintaining.
If your product is still small, it makes sense to keep most of your infrastructure in "maintaining" since the number of services is small. As the product grows (and your company starts hiring ops people), you can slowly migrate to "orchestrating".
I have noticed that too. With some managed services you are trading a set of generally understood problems with a lot of quirky behavior of the service that's very hard to debug.
I've seen this with many AWS deployments, which on a pure hardware cost is 3-5 times more, but with the way it makes it 'easy' to scale instead of optimise instead costs 10-20 times more. When the investor cash starts drying up and the focus is going to be on competing for profits in a market that is much smaller in general, many organisations are going to find themselves locked into AWS and for them it's going to feel like IE6 and asp.net all over again.
We make a decent business out of doing just this, at scale for clients today.
I built a system selling and fulfilling 15k tshirts/day on Google App Engine using (what is now called) the Cloud Datastore. The programming limitations are annoying (it's useless for analytics) but it was rock solid reliable, autoscaled, and completely automated. Nobody wore a pager, and sometimes the entire team would go camping. You simply cannot get that kind of stress-free life out of a traditional RDBMS stack.
DynamoDB vs RDS is a perfect example. Most of that boils down to the typical document store and lack of transactions challenges. God forbid you start with DynamoDB and then discover you really need a transnational unit of work, or you got your indexes wrong the first time around. If you didn't need the benefits of DynamoDB in the first place, you will be wishing you just went with a traditional RDBMS vs RDS to start with.
Lambda can be another mixed bag. It can be a PITA to troubleshoot, and their are a lot of gotchas like execution time limits, storage limits, cold start latency, and so on. But once you invested all the time getting everything setup, wired up...
In for a penny, in for a pound.
The default config for a LAMP stack will easily handle 100 requests per second. 10 if you app isn't optimized.
Run apt upgrade once a month and enable automatic security updates on ubuntu.
That is neither hard nor "vodoo".
I've used managed services and I don't see the point until you hit massive scale, at which point you can afford to hire your engineers to do it.
Now that person I pay to operate the service can focus on tuning, not backups and other work that's been automated away.
Sounds like a massive win to me.
That is, the shift to managed services does largely remove a large portion, and just changes another.
That's nice in theory, but my experience with paying big companies for promises like that is I still end up having to debug their broken software; and it's a lot easier to do that if I can ssh into the box. I've got patches into OpenSSL that fixed DHE negotiation from certain versions of windows (especially windows mobile), that I really don't think any one's support team would have been able to fix for me / my users [1], unless I had some really detailed explanation -- at that point, I may as well be running it myself.
[1] And as proof that nobody would fix it, I offer that nobody fixed this, even though there were public posts about it for a year before my patches went in; so some people had noticed it was broken.
I've been informally tracking time spent on systems my team manages and managed tools we use. (I manage infra where I work.)
There is very little difference.
And workarounds and troubleshooting spent on someone else's system mean we only learn about some proprietary system. That's bad for institutional knowledge and flexibility, and for individual careers.
Our needs don't mesh well with typical cloud offerings, so we don't use them for much. When we have, there has yet to be a cost savings - thus far they've always cost us more than we pay owning everything.
I mean, I personally like not dealing with hardware or having to spend time in data centers. But I can't justify it from a cost or a service perspective.
That sounds like a nightmare to me, and I'm not even a server guy, I'm a backend guy only managing my personal projects.
> I'll pay Amazon $N/month to do that for me, all that matters is the product.
I don't want to pay Amazon anything, I want to pay Linode/Digital Ocean 5 or 10 or 20 dollars per month and I can do the rest myself pretty well. My personal projects will never ever going to reach Google's scale and seeing that they don't bring me anything (they're not intended to) I'm not that eager to pay an order a magnitude more to a company like Amazon or Google in order to host those projects.
but companies throw money, hardware and vendor lock-in at problems way before commiting to sound engineering practices.
i have a friend who works at an online sub-prime lender, says their ruby backend does hundreds of queries per request and takes seconds to load. but they have money to burn so they just throw hardware at it. they spend mid-five figures per month on their infrastructure and third party devops stuff).
meanwhile, we run a 20k/day pageview ecommerce site on a $40/mo linode at 25% peak load. it's bonkers how much waste there is.
i think offloading this stuff to "someone else" and micro-services just makes you care even less about how big a pile of shit your code is.
Between that and the dogshit documentation, it's truly thrilling to be paying them a ridiculous amount of money for the privilege of having a managed service go undebuggably AWOL on the regular, instead of being able to resolve the issues locally.
The article isn’t asking why use Dynamo vs SQL in the context of the traditional NoSQL vs SQL way
It’s why Dynamo vs QLDB.
Or why Cloud Firestone vs Cloud Spanner.
Or Firehose + some lambda functions vs EMR
It’s wholly separate of the managed angle and focuses on the actual complexity of the products involved and the trade offs therein.
In the surface the two sound similar, but there’s much more nuance to this article's point (take the example of Dynamo, the author isn’t opposed to it for being NoSQL, they’re opposed to it for low level technical design choices that are a poor fit for the problem their client had)
You're gaining an expertise, but instead of it being with the underlying technology, it's with the management services for that technology. I think the premise of your argument is correct that understanding these services is simpler than the technology itself, but I seem to inevitably find myself in some edge use case of the service that is poorly documented or does not work as described, and trolling stack overflow or AWS developer forums is the only solution other than trial and (no) error.
Let me tell you: it's not fun when you're entrenched into a proprietary service that is now outdated and kept only for legacy reasons, and the version you are on doesn't allow you to do things like update your programming language to a more recent version. Now you get to move to another managed service and lose all that expert knowledge you had with the old one.
If the 0day in your familiar pastures dwindles, despair not! Rather, bestir yourself to where programmers are led astray from the sacred Assembly, neither understanding what their programming languages compile to, nor asking to see how their data is stored or transmitted in the true bits of the wire. For those who follow their computation through the layers shall gain 0day and pwn, and those who say “we trust in our APIs, in our proofs, and in our memory models and need not burden ourselves with confusing engineering detail that has no value anyhow” shall surely provide an abundance of 0day and pwnage sufficient for all of us.
Thus preacheth Pastor M. Laphroaig.
In that way, cloud/serverless migration is really the whole worldwide IT workforce progressively willingly migrating IT (design, development, operations & applications) out of companies, to just a few big providers.
Just like energy.
And who gets to decide destinies in the world? Those who control, have and provide energy.
What a time to be alive.
End of XIX century was the rush to oil-based energy to control the world. End of XX was the rush to data-based capture and control.
Such centralized systems also become significant points of failure. What happens to your small business if one of the many services you rely on disappears or changes their API?
"All that matters is the product" sounds very much like "all that matters is the bottom line" and we've seen the travesties that occur when profit is put above all else.
Also having granular IAM for services and data is very helpful for security. You have a single source of truth for all your devs, and extremely granular permissions that can be changed in a central location. Contrast to building out all of that auditing and automation on your own. Granted, IAM permissions give us tons of headaches on the regular, but on balance I still think it's better when done well.
If you're concerned about AWS/GCP/Azure looking at your private business's data, I think 1.) that's explicitly not allowed in the contracts you sign 2.) many services are HIPAA compliant, hence again by law they can't 3.) They'd for sure suffer massively both legally and through loss of business if they ever did that.
Depending on the problem, choose the right tools. For example so far everything I've seen about React and GraphQL tells me that maybe GQL isn't necessarily the best solution for a 5-person team, but a 100 person team may have a hard time living without it.
Kubernates / Docker is significant work. And when we had 4 developers making that effort was wasteful and we literally didn't have the time. Now at 20 we're much more staffed and the things Kubernates / Docker solves is useful.
Meanwhile we have a single PostgreSQL server. No nosql caching layer. We're at a point where we aren't sure how far we gotta go before we hit limits of a PSQL server, but at least 3-4 years before we hit it.
Point is look at the tools. See where these tools win it for you. Don't just blindly pick because it can scale super well. Pick it because it gives you things you need and a simplicity you need for the current and near-term context, and have a long-term strategy how to move away from it when it is no longer useful.
In reality that means that when something doesn't work as expected, instead of logging in and solving the problem, you'll be spending your time writing tickets and fighting through the tiers of support and their templated replies, waiting for them to eventually escalate your issue to some actual engineer. Wish you a lot of luck with that, I prefer being able to fix my own stuff (or at least give it a try first).
If you don't believe that this counts as "managed," then for all your words at the end, you believe at the end of the day that making a solution more complex so that failures are expected and humans are required is superior to doing a simple and all-in-code thing. For all that you claim to not like operations, you think that a system requiring operational work makes it better. You are living the philosophy of the sysadmins who run everything out of .bash_history and Perl in their homedir and think "infrastructure as code" is a fad.
The Raspberry Pi is managed by you, as opposed to managed by a different company / different people so you can focus on your business problem.
(Sure, managed solutions can satisfy all of those. But not without understanding a few orders of magnitude more complexity than a simpler solution, and you'll never get there if your attitude is "I'll pay Amazon to do that for me." Amazon doesn't care whether you succeed.)
Wait, what? Why wouldn't they? Successful companies will spend money with them; defunct companies can't.
So you want technology that prevents you from doing something you might want to do, real badly?
Now when many sql databases support json I don’t quite understand why you would use a documentdb for anything else then for actual unstructured data, or key value pairs that needs specific performance.
I used mongodb for a while now I’m 100% back in sql. Document dbs really increased my love for sql. But I’m also long from being a nosql expert and I never have to deal with data above a TB.
It's tough to get the correct level because you can ansible > docker > install an rpm or just change the file in place. Both have their place and "hacky" solutions can work just fine for many years.
Future of what exactly?
The choice of technology should first be based on the problem and not whether it is serverless. Choosing DynamoDB when your application demands strict serializability would be totally wrong. I hope more people think about the semantic specifications of their applications (especially under concurrency).
Serverless sounds great for "some" kind of systems but in other cases its a complete missfire. Our job is to know which tool to use.
On the other hand, I find it frustrating how many developers either don't know basic SQL or try to get cute and use the flat side of an axe to smack down nails instead of just using the well-known, industry-standard beauty of a hammer that SQL.
It's managed at every shared hosting in the world.
Great goal. I guess you never had any use case to login to a production server to investigate a business critical issue.
This was the most enlightening piece of the article for me. Their alexa rank today is 48 (globally, 38 in U.S.), so whatever your site is, you are probably not dealing with as heavy a load as them. What techniques do you have to employ to serve this many requests from a single database server?
https://nickcraver.com/blog/2016/02/17/stack-overflow-the-ar...
Lots of caching (redis, CloudFlare) and trying not to use the database unless absolutely necessary, I would expect.
Lots of caching, lots of RAM, SSD storage, and a low-level ORM for SQL.
but consider their stack there are better "starting" alternatives (they also used them before dapper), like Linq2Sql, EF Core, etc...
Also compared to other languages Linq and EF Core are way better than most alternatives (besides that some things are just not working, like lateral joins)
How much do you think stack overflow is reading vs writing? How often do you even click a link from a stack overflow page once there?
They probably also save a lot on ops team - running a single sql server with a failover is well-understood, error-resistant and covered in vendor's guides/best practices.
https://www.ageofascent.com/2019/02/04/asp-net-core-saturati...
hipness has increased
By the way, 200M is less than 2500 requests/second which is not much at all. In terms of actual throughput, they are not that big.
They are probably not spready uniformly throughout the day...
At a large scale (e.g. hundreds or especially millions of hardware nodes) the most common faults will be due to individual nodes / services / whatever failing, so you want a complex fault tolerant system to deal with those faults.
At a small scale (e.g. stuff that can fit on one or several servers) the most common faults are from the system itself, not from individual nodes. Here, using a complex system will drive up the likelihood of failures, especially when you don’t have a large team to manage the system.
That's pretty scary for anyone having at least three nines SLA
That is, tech performance rarely killed a startup, but programmers not understanding product kills them all the time.
Which of the following two developers has the better chance of breaking six figures next year:
"Used Hadoop, MapReduce, and GCP for fraud detection on..."
...or...
"Used MySQL and some straightforward statistics for fraud detection on..."
This is a big part of why all these things exist in places where they shouldn't. As a dev that always goes for the simplest solution first and has yet to break a hundred k at 40 ... I'm spending my evenings now trying to figure out how to deploy the latest technologies where they're totally not needed.
I’ve interviewed a lot of incredibly bright people that didn’t know any technologies more modern than C++.
I've done a lot of interviews at BigCo and definitely not looking much at the technology.
tldr: software developers often can't measure the impact of their code, so they fall back on describing the technologies they use, which drives a counter-productive desire to employ "sexy" technologies.
I jest partly, because every time I write something in C# or PHP (even with the half-baked "strong" typing they are introducing), I constantly curse and think how easy it'd be in C++
But then they didn't have the context as to why those tools were chosen and so often they aren't suited. But I've never met anyone in the last 20+ years with thousands of engineers who was doing it for their resume. In fact the best thing for your resume is for the project to be a success anyway.
Looking at new tech is way more interesting than doing the same thing you've been doing for 10 years. Its good to have some mechanism in place to vet that since I know I don't trust myself to always make a good choice (and in my opinion I have excellent taste).
The end result is that not only is the routine stuff boring but it's also career-limiting.
I don't find it insulting at all. Reality is that many, if not most, recruiters are looking for candidates that have experience in whatever technologies are hot at the time. Expertise in a sought after technology can be worth tens of thousands of dollars extra in annual compensation. It's natural that some people will try to use that technology to further their career if possible. And I think to a degree, a lot of people base their job decisions on what they'll get to work on. 'Is it a stack that is in demand and growing?' 'Is it a stack that might be difficult to learn but pays really well?'
Now the sensible thing to do is to not shoehorn in some tech if it doesn't make sense for a given application. And beyond that, the a big reason people are tempted to chase the shiny new object is because the IT recruiting process is broken.
I just see it as people trying to do the best they can for their career, usually there's no malice involved.
Maybe, or maybe it's just realistic in a world where common advice is not to stay at any company for more than 2 years.
The common advice and wisdom is actually not to do this.
And even if it is just for the resume, it's smart and understandable. Consider the difference between, "Eh, I evaluated some NoSQL options here and there, but it seemed safer to stick with Postgresql since we didn't have a compelling reason to experiment. Postgresql works fine at the scale of our product and we're very familiar with it," versus, "Yeah, we used Mongo for a project and it was a shit show, switched to DynamoDB which has been solid. The product is still 80% on Postgresql, but we use Cassandra for analytics, and of course we run InfluxDB for Grafana, which we're going to replace with Prometheus for better scalability." These could be two people facing exactly the same serious of engineering choices, with the first guy making the better decision every single time, but the second guy sounds like the kind of curious and hardworking person you'd want to hire, while the first guy sounds maybe... stuck? Counting the days to retirement? Maybe boring SQL guy has saved his company a ton of engineering work that they were then able to invest in product work instead of engineering, but when he and NoSQL dilettante guy are both interviewing at a new company that wants people who are "curious" and "passionate" about "dedicated to learning and growing their skills" and "ready to meet new challenges head-on" he has to be worried about sounding like a dud.
But I've never met anyone in the last 20+ years with thousands of engineers who was doing it for their resume.
It's not something you can distinguish from being eager to learn and overenthusiastic about new tech, which we all are to some extent.
I have seen this happen before, and when the interests and ambitions of an individual are counterproductive to the direct needs of the company, that can be deeply problematic.
Also, we like playing with new toys. Solving the same old problems the same old way we were doing it 20 years ago just isn't sexy.
Uh? That's just the norm. I've met very, very few people who actually think about introducing a shiny new technology as a last resort. Usually it's the complete opposite, they start with the shiny new thing they've read about and build systems and even features around it. I've often heard the CV mentioned explicitly as a reason, but there is also a lack of imagination- they just can't think of ways to adapt the current system to their requirements.
Ah, as for the project to be a success... "I've started as the lead developer in a team of two, and after X years the project has grown and I was leading a team of 20"- this is a common success story. Nobody cares about the fact that the team of 20 was needed because the tech choices made were so poor you needed to explode the size of the team.
This is not a real criteria for choosing anything. You can say the exact same thing about ASP.NET WebForms, MySQL, MSMQ or pretty much anything that was popular at any point in time. If you're only looking for "success" stories, you're not doing due diligence as an engineer. Who defines "success" anyway? The same person who chose React, Spark and Kafka in the first place and whose salary is directly proportional to how hyped-up those technologies currently are?
Asking questions is not assuming. I think you're letting yourself be triggered by a flippant remark at the end of an otherwise solid comment.
Keeps me sharp and gives me an opportunity to explore tech that isn't in my day-to-day wheelhouse.
If you are still running on Tomcat, all the best getting the market standard genius talent to work for you.
You would also like to learn about what happens to workplaces where smart people don't work or even like to work at.
Often there's an aspiration to reach Google-level operation. Invariably some director/VP insists on building for that future. Then it turns out sales can't break into the market and adoption is low.
Now you have a neglected minimum viable product because you're scaling that skeleton instead of adding features your existing customers want. Or you're delivering those features at a slow pace because you're integrating them into two versions of the software: the working one and the castle-in-the-sky one. Then there are all sorts of other things that throw a monkey wrench into your barely moving gears: new regulations, competitors nipping at your heels, pivots, demands from management for "transparency and visibility" into why it's taking you forever to deliver their desired magnum opus.
Products that reach Google scale probably get there without even noticing it because they're correctly iterating for the marginal growth.
More importantly, the world changes and software has to change with it. An accounting system will need to be fitted to changing regulatory frameworks in the future. That knowledge is the basis of good application design. Schemaamnesia databases with highly-scalable vendor-locked PaaS peripherals that are constantly changing also require constant developer attention, but do not rise to that challenge.
A technology that's been around for decades and is still considered a viable candidate will be much more reliable, documented, mature, bugfree... than some new library that is untested hasn't been worn smooth over time.
Also personally love the callout to Joe who while being a professor is consistently practical on when approaches do or don't make sense and at what scale they do.
But a basic stack can only be used for basic systems and for basic problems. And there just isn't many of these going around anymore as they've all been solved. Or more commonly nobody is interested (including from the business) in delivering something basic. They want to innovate for their users.
No, your company isn't special, and "innovating for your users" is probably the worst thing you could do, compared to delivering a good product using established practices that are tried and tested to deliver results.
1. I know the risk of only being able to peer through the fence at the distant piece of software that is running your business, unable to gain any insight while your production application is limping badly and customers are running away like water. On top of that, thrashing it out with Mr Clippy is better than average support. If my business depends on it I want experts at hand who have the tools and the access they need to do the job they do.
2. The insane pace of serverless is entirely fad driven and lacks quality engineering which is required for critical pieces of software. The tools are universally poor quality, unstable, unreliable, poorly documented and built on gaining mindshare, making IPO and selling conference seats. Best practice never materialise as the rate of change does not allow an ecosystem to settle and work the bugs out. The friction is absolutely terrible but no one speaks of this in fear of their cloud project being labelled a failure. Every person I have spoken to for months is hiding little pockets of pain under a banner of success. Some people clearly will never deliver and burn mountains of cash hiding this.
3. Once you enter an ecosystem, you are at the mercy of the ecosystem entirely, be that a service provider or a tool. Portability is always valuable. It has cost, scalability, redundancy and risk benefits far beyond the short term gain of a single vendor decision. I'm currently laughing at an AWS quote for a SQL Server instance with half the grunt, no redundancy, no insight possible for only 2x the capex and opex combined of dedicated hardware including staff. But can't move to Azure because everything is tied into SQS.
I can never be behind anything but IaaS myself. This is contrarian especially in this arena but I will put my 25 years of experience on the line every time and say that it is the right thing to do. IaaS is choice, flexibility, allows you to gain deep insights and protects you from serfdom, fad technology and pick and choose mature products rather than what the vendor sees fit.
This is just another rehash of buying a mainframe. It's just bigger and you pay hourly to write COBOL.
2. Give it time; this will evolve and it is still in its infancy. Like all things it is buggy in the beginning but as adoption hockey-sticks so too will the stability, documentation, etc.
3. Portability is traded for breadth of services and depth is gained through vendor lock-in and the one-size-fits-all package. Of your concerns I would say this one will be around for a long time or at least until a "conversion kit" is built to shoehorn all of your stuff into another provider allowing you to jump ship or test the waters elsewhere.
I am a veteran like you (20 years) and while you choose IaaS because you like the control most want to punt their problem to someone else and to pay for that.
If our goal is to get from LAX to JFK we could fly our own plane, charting our own course, looking up weather, doing engine checks, refueling, dealing with air traffic control. and we'll get there and be in full control the whole way. However most will pay to be shuttled in to a commercial airplane. Then of course there are some that are willing to pay more to have their own personal pilot get them there in a chartered aircraft where they are afforded a more tailored experience.
There are some really great pilots you can hire out there to do it all for you (or if you are in fact one yourself) but if the goal is to simply go from point A to point B in as little time as possible we don't have time to find, train, and rely on a single pilot or ourself to get there. We will take the hit and pay others to get us there.
That is what all this serverless non-sense is about if you ask me. The tradeoffs of simplicity and handing the busywork off to someone else is more enticing than the control we have in the process. Also, isn't it nice to be able to say, when you arrive late, that it was the airlines fault? :)
Or other people are wondering if we could replace all of our relational databases with kafka. While complaining about inconsistent data sets. Well let's talk about advantages of a good relational schema first.
Maybe I'm turning into a grumpy old admin/DBA... but mariadb or postgres used well just solve so many problems.
No you're not; you're saving the company tens or hundreds of thousands of dollars over the coming years.
Unless you're using your DB as a data transport, dear God why. And if you are, dear God why. I work with Kafka and it's very good for a certain subset of problems. Replacing RDBMSes is not one of them.
Sometimes I am wondering, how people will rationalize the layer bloat in 5-10 years, when hardware went through a few more upgrades.
The problem is that maybe cost benefit decisions are really hard. To understand the cost of map reduce, for example, I probably have to use it at least once. High cost just to understand something. And I have to understand the alternatives, so I'm trying those out too.
You know what's faster? Signaling. When large companies say "this approach works for us" it's really cheap and easy to say "Well that's probably good enough then".
Do you lose out on efficiencies by not fully understanding the problem? Duh. But probably not as much as understanding an entire domain for every technical decision you make.
The steps outlined, particularly 1,2,3,4, are really expensive for every single technical solution. Reading multiple technical papers for a db is probably more costly than just acquiring customers, hitting scaling limits for your uninformed choice, and moving to a new db based on a somewhat more in depth approach or hiring a consultant.
Let's assume that if you guessed a technical solution, the cost of 'guessing' is 0, and the cost of picking the wrong solution is 'low'. Should I waste time doing anything other than guessing? If the cost of 'do what google does' is 0, or near 0, and better average chance of working, why not pick it?
Unless the stakes are high, putting that kind of investment into every technical decision just seems needless. And they arguably have to be very high to offset the cost, and I would state that determining those costs upfront itself may be difficult and error prone.
For every mis-guess listed in that article, how many new customers did they acquire due to making the right 'guess' based on signaling from other companies? The companies clearly hadn't gone under from those bad decisions, so would it really have been the right call to have invested in such deep thinking about the problem early on?
Not sure I believe what I wrote entirely, just thinking out loud.
That makes more sense if you s/cost/risk/. The answer to why you shouldn't just do whatever Google is doing is because trying to copycat the leader is the one thing that will guarantee your _failure_ in the market place.
Same with AWS. AWS Lambda was built to solve existing problems. Everybody will use it to solve the same problems it's designed for, but only a tiny number of people will actually succeed in the marketplace because it just becomes a numbers game. If you want to win you have to play a different game--look for problems where AWS Lambda is a poor fit.
If you're doing the same thing as everybody else you're by definition just reinventing the wheel. Why bother?
> The answer to why you shouldn't just do whatever Google is doing is because trying to copycat the leader is the one thing that will guarantee your _failure_ in the market place.
The article seems to demonstrate the opposite - decisions following the leader, and paying for it later, but surviving in the interim.
> If you're doing the same thing as everybody else you're by definition just reinventing the wheel. Why bother?
I don't really agree. When it comes to something like a DB or architecture, that's probably not where you need to be inventing new solutions - products are rarely sold based on their technical implementation, and more about the results they drive.
If you were building a new Google search, and you used all of the same tech, that's probably a bad idea. But if I'm building a product for a completely isolated domain, who cares if it's solved using the same tools that Google solves search?
Regarding lambda, there's likely tons of room for competitors even within a domain to implement off of the same technology - again, it's just a tool, and not likely to be the differentiator.
In other words, if a Postgres based solution takes 1 month to build, and a CloudSpanner/Aurora/pick-cloud-managed-sql-DB solution takes 1 month to build, just pick one. Most of the time, if you make a mistake, you'll find out when you have more customers/funding/engineers/knowledge of what actually matters.
I've seen lots of time wasted overthinking solutions only to find out that the requirements were changed a month before release, or customers would use the system completely differently than expect.
I don't know why it is but people seem much more interested in the tech than in the value they can create. Which is odd. Why? Because it's the second part where you get to exercise the most creativity and do the most interesting work. How do you think LinkedIn arrived at Kafka, for example?
What'a possibly more baffling is it's not just the devs who buy into the technology hype: often it's the so-called business types, and those in management and leadership positions, advocating and egging them on.
Well, Franz Kafka wrote about incomprehensible, oppressive bureaucracy as a way of life, and LinkedIn did a decent job of porting that to Web 2.0.
Well, for some engineers who don't have much say in product development, the amount of value they can create is bounded by the business specification/deliverables they have to produce (and maybe equity or lack thereof plays a role here). However, if the architecture is considered their responsibility, then there's a lot of fun in using _shiny new toys_ even if it's overkill and will incur unnecessary technical debt. It's fun to overengineer things.
Not saying it's a good thing or even professional behaviour, but I understand it.
There would be no Internet, Web, VR, ML, AI etc and we would all be writing assembler or using punch cards.
The Internet, Web, ML, and AI at least were all originally invented using research funding, not business funding. The Internet (specifically the TCP/IP suite) was funded by DARPA, the research arm of the US Department of Defense. So you're right that some people need to focus on something other than business value, but usually that is someone paid to do research, not paid to implement a specific solution that needs to be widely deployed within 1-3 years.
But if you're picking a product to use for a particular task, you aren't doing that kind of research. You are instead doing engineering. That is, you are determining how to solve a particular problem using the knowledge and resources (including reusable code) at hand. So you need to think to determine what collection of resources will actually do the job well, at reasonable cost and time, etc. Different problem.
The issue is the quirks. Every managed service has at least a dozen quirks that you're not going to know about when you visit the flashy sales page on the cloud provider's website. And for the vast majority of users, they're not going to have access to the source code to understand how these quirks work on the backend. So you end up in a situation where yes, it does take way less time to get 95% of the functionality done, but getting that last 5% can still take a considerable amount of work.
As an example, I am using Azure Event Hubs lately. It is supposed to provide something like a simple consumer api like Kafka does, but with consumer group load balancing across partitions. Awesome, there is a system that automatically handles leasing across partitions in a way that abstracts this all away from the client! Except, well actually the load balancing is accomplished via "stealing leases" (meaning, they are not true leases) so if you use the api you are meant to use, you will get double reads - potentially very many if you want to commit reads after doing more-than-light processing which can take time. So you need to use the much more poorly documented, barebones low level api and probably still end up writing a bunch of logic to dedupe.
Except, you use this kind of tool to begin with because you want to set up a distributed consumer group to read from a stream... so now you have a non-trivial engineering problem figuring out a way to get a distributed system to manage deduping in a light-weight way across hundreds of processes and machines...
Several good managed offerings exist if you do not want to run it yourself.
We appeared to have a persistent bug though, processing left running overnight always seemed to crash. So after a couple of nights I stayed in the lab to try and see what the cause was. At 7:30pm the lab door opened and in walked the cleaner who reached down and unplugged the computer so that he could plug in the vacuum cleaner. Problem solved.
https://www.reddit.com/r/talesfromtechsupport/comments/5yrs1...
Let me know when your command line tools run multiple large scale processing jobs on petabyte datasets.
His article was him constructing a straw man about why people use Hadoop and attacking that.
How many people have used Hadoop for a project? How many Petabyte+ datasets do you think are out there?
Unless you truly believe the answer to those 2 questions are the same, I think you can see why that article had to be written.
I also appreciate quality of the writing and step by step journey.
When I first read the article, I found it timely as I felt bombarded by IaaS & PaaS offerings with “web scale” solutions for problems I couldn’t relate to, basic examples, and few case studies.
BUT. I have also grown a bit skeptical of the "Just use a RDBMS!" mantra. The same advice applies. Think about your use case. Even if your data isn't exoscale, a relational DB might not be the best choice.
My current project has data on the scale of hundreds of billions of rows. Nothing crazy, easily handled by a good Postgres box (by the numbers.) And for certain access patterns, it would be.
Unfortunately, it turns out that a lot of the analytics queries our business users need end up being joins and aggregations involving full or nearly-full table scans. PG is not particularly well optimized for this; it gets CPU and disk-access bound. Queries take tens of minutes, causing timeout errors in our BA tools.
Loading the data into a Presto (Facebook product) cluster instead dropped query times for the same aggregation into the tens of seconds range. Sure, our data doesn't even really count as "big" and Presto is probably overkill if you're only looking at size, but it is optimized and built for highly partitioned parallel aggregations.
I think that's the better way of saying "just use a RDBMS" - "If you're not sure, use an RDBMS to delay a decision until you have more experience and knowledge".
Postgres has to deal with a transaction modifying 100,000 rows in a table, needing that transaction to then see its changes but then aborting, leaving the data exactly as it was before. Athena/Presto don't have to worry about any transactions, table files bloated by changes that need vacuuming, etc. etc.
So I don't think it's fair to claim PG does not scale, when you are comparing a transactional and analytical database.
With that in mind, it's my experience that every tool or framework has a different window of scale in which it's ideal, which has both lower and upper bounds, and one simply needs to find one's own project along that axis when choosing a technology. Hadoop may be the best solution as N approaches infinity - and we as programmers love thinking in terms of infinity - but it may not start being the best solution until 10x your actual range for N.
A lot of the time design meetings in a startup revolve around "Is this scalable?" or "What happens when we have 10,000 users or 1 million users?".
I think the problem isn't choosing overkill technologies, the problem is trying to solve a problem that doesn't exist yet and probably never will.
I think it would take a very severe recession, in the 2009-2010 style, to hammer venture capital down considerably. Instead I expect Japan stagnation to continue to envelop the US economy, and for the exact same debt-laden reason. Ever slower growth, low traditional inflation (higher real inflation from currency debasement QE), debt taking up an ever larger share of capital available for investment, stagnant productivity, stagnant wages. With the enormous amount of wealth in the US, trillions of dollars will always be looking at the venture capital market in a given decade. ~2012-2035 will probably be the best years for venture capital in US tech history, and broadly the best context for start-ups, that we'll ever see. It's early in the loose money sloshing period from perma low rates, but not yet late such that you're eating an always-on QE real-value debasement hammer constantly (which sucks down your real value creation, as you run uphill against the Fed while it tries to debase government debt to keep the US Govt solvent).
Just thought of a naming convention, how many types of microservices do you run:
10? deciservices, 100? centiservices, 1000? milliservices.Assuming they aren't randomly changing unversioned APIs, of course...
The number of projects I've had where the team has had one arm tied behind their back because they "have" to use Hadoop or NiFi or Lambdas or something else a manager is determined is the thing everyone's using is just lunacy. They're all tools which have their place, but you really have to know them to know when to use them. And as importantly when not to.
I mostly do consulting gigs where projects are 6 weeks to 6 months long and it's been years since I've seen that not hamper a team.
It's overkill for anything I need to do, but my home server is six 8-core ODroid devices, orchestrated with Docker Swarm, with most of my services on there being load-balanced, and glued together with Kafka if they need to talk to each other.
Do I need my internal video streaming server to be able to scale horizontally? Of course I don't, there's only ever three people max watching things from there at any point in time. However, I find that, overall, my brain thinks more-or-less in terms of these microservices, and it doesn't really hurt much to do it that way.
If I find it pleasant enough to do, why the hell not make it hyper-scalable and able to reach Google levels?
EDIT: Just a note, I know that Docker Swarm probably isn't quite capable of Google-size. Still, moving to Kubernetes or something wouldn't be terribly hard (the reason I didn't use it was because a few years ago I had some issues with Kubernetes on ARM and Swarm worked outta the box).
I guess that backs up my original point even more then; my stupid video streaming server might actually be able to scale to Google size some day :)
I'm not even sure it is engineers at bigger companies who choose to jump on these bandwagons. I suspect it is often wannabe-technical managers who read that MegacorpX is using tech Y that sells the idea to upper management, along with some unfounded beneficial promises, that causes some of the trends we see.
My point is: these generalizations in either direction are not helpful. Many people are operating at Google scale, and many people aren't. Be aware of that when you read about potential technologies.
So if you're the audience of this article, you're not Google.
If you're doing large scale ML, having to respond to requests in very low millisecond numbers or having to crunch through significant amounts of data (often more than Google produces) then you need to use similar approaches to them.
Even doing ML at all can be quite demanding.
Well, duh.
The corollary is: if you're using one of these tools, and you HAVE thought through your use case and reasons, don't get defensive about it. If challenged by those with lesser knowledge after reading an article like this, calmly and rationally explain your reasoning. And like any technology choice, be prepared for someone else to offer another option you weren't aware of, with better reasons.
Again, duh.
For example, one might like using MapReduce or Kafka Streams for their programming paradigms, not for the redundancy or scale they provide.
Another example, Kubernetes, which makes it trivial to run containers and attach storage to them.
You'd think so, right? But what happens to me all the time is I try to use "small-scale" tools to handle some small problem, and then I find myself writing glue code that's already implemented in, say, Kafka Streams. So I might as well just run one Kafka broker locally and write a Kafka Streams application. Or the same with the Django ORM, which I reach for all the time just because I don't like writing database access code when I can write up some models and be done with it. Every time I reach for "small" tools I end up writing tons of code that's not actually solving my immediate problem or question.
Don't we just. Which makes us vulnerable to every shiny gewgaw and every personal and cultural bias out there.
That's not necessarily entirely true. In the early 2000s, I worked for Cadence Design Systems, who at the time needed to build and test the plethora of tools they'd built or acquired on a large variety of systems. I worked on "GridMatrix", sort of like make but using a large, heterogeneous cluster of machines, built on top of the Condor and LFS batch scheduling systems---the same sort of thing underlying MapReduce.
On the other hand, I get where the author's going: cargo-cult design is rampant in enterprise software development. But it's not just cargo-culting Amazon or Google; it also involves fashion and resume padding.
And then, there's "eNumerate multiple candidate solutions. Don’t just start prodding at your favorite!"
Bwahahahahah. <- Unamused laughter.
Our first and only response, as an engineering discipline (if you want to call it that) is to pick the first idea that comes to mind and beat it to death with a stick.
Maybe Hadoop isn't necessary for everyone who needs it. They'll learn. And if they don't, more job security for you.
Otoh there are some crazy advantages of running certain "large scale" softwares. Not everyone needs the "scale" Kubernetes offers, but managing tasks declaratively as containers that get scheduled on boxes in a hermetic fashion? Or maybe continuous delivery - not everyone needs CD, but it certainly offers many advantages to those who do it.
Most people could probably live off of shell scripts running under Cron on a box with CGI scripts written in Perl without any issues, that does Not imply there are no advantages to new technology because not everything is about scale and fault tolerance.
My old laptop could do something like 150K transactions per second in PostgreSQL. You can scale really far before your database becomes the issue.
Or, for the youngsters: Portland giveth and Seattle taketh away.
Some things never change.
Oh wait. They dont say that. They fund, and hope and pray!
Most programmers try to do that. And they keep trying. They're wasting their lives.
http://highscalability.com/blog/2011/1/26/google-pro-tip-use...
https://people.eecs.berkeley.edu/~rcs/research/interactive_l...
It might seem expensive to do things the "wrong" way first every time, but I really do think it's necessary.
This is the crux of the issue, psychologically. When making an important decision in the face of uncertainty, having someone else with a lot of clout (or a group of people whose collective wisdom you trust) simply tell what you do feels amazing. That feeling is hard to ignore in favor of constructing your own cost-benefit analysis.
https://www.jwz.org/doc/backups.html
In particular Addendum B:
RAID is a waste of your goddamned time and money. Is your personal computer a high-availability server with hot-swappable drives? No? Then you don't need RAID. RAID is not a backup solution. Even if you use RAID, you still need backups.
Side note: the linked URL redirects to a questionable image when navigated from HN. Interesting choice by the site owner.
I have more than one friend that has overengineered an industrial solution for basic computing needs. (and when things break, complexity goes exponential)
as to URL redirect: I don't see the image/redirect you refer to, even clicking on the link here on hacker news. Can you elaborate? I take the links I recommend seriously.
Or to put it another way, form follows function.
It's not about scalability or about anything technical. It's about domain understanding and data isolation.
There's NO problem here to understand. It's just one way to manage our data. Or another way speaking, design for failure.
Sure, but there are tradeoffs to any approach. More services == more devops == more complexity in fault tolerance, communication, data storage/consistency, etc.
I'm not saying it's a bad model _for you_, obviously I have no idea, but the point of the article is that some tend to jump toward complexity when they shouldn't.
If i need to change one part of data, i don't migrate the whole data. I just need to change that only changed part.
Secondly, extreme YAGNI and extreme future-proofing are two ends of one of the considerations on how to approach writing a software. Neither extremes is good. In reality, the best location on that spectrum would depend on the situation at hand.
There is no silver bullet. And I'm sorry but I didn't learn anything new from this blog post, something that is already not part of the software engineering vast body of knowledge.
edit: I know UNPHAT is not an "exact synonym" of YAGNI. But what you're describing is a problem solving approach. That's probably even less new. Please look at Polya 'How to Solve it' if interested.
You can read through the transcript too ~> https://changelog.com/podcast/260#transcript
But there's also Elixir which can do away with most of these concerns.
So what do I do? Go with Rails to ship stuff? Or struggle a little bit more with Elixir to get that much more legroom?
That said, clamoring toward things like React has yielded a net positive result. So cults can be a double-edged sword.
I think we maxxed at somewhere around 60 concurrent before I left. They gave up on the idea a few months later.
I'm acutely aware I'm not Google.
Should we stop using MySQL too? That handles more data than I have.
I figure, in a way we should be thankful for this cluelesness. If people had any clue whatsoever, IT employment would shrink by a factor of at least 3, and salaries would drop pretty massively as well.
I'm reminded of an issue my old boss had at his house. His power bill was raising at an alarming rate, he assumed it was the central hvac, spent a few hundred getting it tuned up, no effect. Spent a bunch of weekends replacing weather stripping, adding sealant, etc. minor effect.
I came in with an app that lets you estimate current power usage by looking at the power meter[1], and killed circit breakers one by one while measuring current usage.
The cause turned out to be a malfunctioning swamp pump he forgot the house even had.
He assumed he knew the problem, by focusing on the common cause of such problems, he assumed wrong, and because of this, every solution was addressing a different problem than the one he was trying to solve.
-
Another example is the story[2] of the old as fuck ibm mainframe that was powering a website that would take on the order of 10s of seconds to return data. Everybody assumed all of the delays was to be blamed by the "legacy hardware". Every solution pitched involved migrating off of it, a daunting task no manager wanted to green light. Finally they brought a consultant in to figure out what was going on. Turned out it returned data in 6ms or less, the cause of the lag was the java app that would read the output from the mainframe to transform and send to the browser.
They could have solved the problem years ago if they had just actually tried to understand the problem.
-
[1] (power meters have a thing that happens every n watts, along with a decal that tells you what n is, so you can determine current power usage by measuring time between these n watts, there is a handy dandy app https://play.google.com/store/apps/details?id=com.sam.instan...)
[2] 7074 says Hello World
"No that cant be the issue because it shouldn't be affecting it"
Yes under ideal situations this wouldn't be an issue but if the system is not working as expected then how can you be so sure that this bit is working as expected.
Things that don't scale or be AA Quality will show up.
Go viral and fail to service. 15 minutes of fame squandered on an avoidable tech mishap.
One time a top comment on a thread was about a misspelling.
Deal with problems when you have them.