What I Wish I Had Known Before Scaling Uber [video]
youtube.com
youtube.com
When did we loose our heads and think such an architecture is sane? The UNIX philosophy is do one thing and do it well, but that doesn't mean be foolish about the size of one said thing. Doing one thing means solving a problem, and limited the scope of said problem so as to have a cap on cognitive overhead, not having a notch in your "I have a service belt".
We don't see the LS command divided into 50 separate commands and git repo's.....
I work at Airbnb and while we have many fewer engineers and services than Uber does, most of the issues he talked about in the talk resonated with me to one degree or another.
One thing to possibly consider is that once you have set up tooling for a platform; logging, request tracing, alerting, orchestration, common libraries, standards and deployment. Deploying and operating new services becomes relatively straight forward.
Saying that, 1700 is a lot, which makes me intrigued to see inside uber.
I'm currently working with a very (very) large UK-based retailer, with various outlets globally, and I can tell you from first hand experience that their selection process is ruthless. There are no, or at least very few, lamers in the tech team.
And yet, when you look at the systems, and the way things are built, some of it just seems crazy. Only it's not, and for many reasons, of which here are only a couple:
- The business has grown and evolved over time - with significant restructuring occurring either gradually, or more quickly and in a more intentionally managed way; systems have grown and evolved with it, often organically (which is something I see at every company I've worked with).
- Legacy systems that make it really hard to change things, and migrate to newer and better ways of doing things because they're still in production and depended upon by a host of other systems and services.
- The degree of interconnectedness between systems and services is high across the board; this isn't a reflection of bad design, so much as a reflection of the complexity of the business.
In any large organisation these kinds of things are apparent. I would imagine that in a large organisation that has become large very quickly, like Uber, if anything the problems would be magnified.
To say that they have poor engineers is therefore very unfair: these are people who would have needed to ship RIGHT NOW on a regular basis in order to facilitate the growth of the business. That's going to lead to technical compromises along the way, and there's little to be done to avoid it. There's also a difference between being aware of a problem and having time, or it being a priority, to fix it (e.g., iterating and retrieving).
It's interesting because this kind of technical debt will eventually open the door to competitors once Uber is under more financial pressure...i.e. they have burned through all the investor money.
It probably doesn't help that canonical REST as usually applied to APIs doesn't really have a single well-known pattern for this.
It's not sane. That's the entire purpose of the talk.
It's led to my personal conclusion that most production issues are caused by people, not errant hardware or systems.
Not really. DevOps is simply a modernish label for system administration that we have for dozens years already.
Sysadmins have tools to automate their work for a long, long time (cfengine, bcfg2, even Puppet and Chef predates DevOps hype). DevOps didn't bring anything new to the table.
It is also not about hype (or not). I do agree that the name is posterior to the beginning of the practice, but it is the name that we have.
It's still no progress, hacky scripts written by sysadmins or similarly hacky scripts written by programmers, and systematic way of automating tasks was available to sysadmins and was used by them for a long time. DevOps brought nothing new to the table.
We create ourselves so many of the problems we are paid to solve.
Sounds lucrative ;)
I'll bite. :D
It's a bit more nuanced than that IMO, the "deploy often" mantra is only as good as the process around release + deploy. If you half-ass testing and push to production without without a process for verification -- or if, say, your deploy process is half-baked, or your staging environment is worthless -- you can probably expect "interesting" production deploys on a pretty regular basis.
As much as we'd like to pretend we're all good engineers, this happens more often than you'd expect -- even with good engineers: at some point a company transitions from scrappy startup to a shambling beast, and the things that used to work for a scrappy startup (like skimping on testing and dealing with failures in production) are insufficient when you've got more eyes on the product. Further, the engineering culture remains stuck in "scrappy startup" mode long after the shambling has begun.
And all that's ignoring the fact that less frequent deploys with more changes have their own set of problems. We actually got to a point with deploys of a certain distributed system such that we were terrified if we had more than a few days worth of changes queued up. So many things that could go wrong! :)
> We create ourselves so many of the problems we are paid to solve.
This, on the other hand, I completely 100% agree with: if not us, then who? :)
EDIT: minor formatting change
Occasionally we'll have a problem where we cannot deploy to production for several days (normally it's once a day). A massive inventory of ready-to-deploy features builds up. When we do finally deploy, this deluge of features and fixes creates new issues that force a rollback, delaying the deploy even longer...
It's a well known fact, both for systems and networks.
Edited out consideration: traffic in your area might be a good proxy for Uber volume. Traffic late at night is generally low.
From personal experience, having driven for Lyft and Uber I'm two cities (Sf, sd) surge in SF does get acutely high in the morning, more broadly in the evening (work departure times are spread out over a wide range of hours), and acute around 2am. In California, bar closures statewide are 2am at the latest, so that accounts for unified departure times in multiple markets.
Man, I'm gonna save so much money.
Most of the other alerts are due to network issues, moving servers due to hardware intervention that change the topology of the network, etc etc.
A system running just fine will usually continue to run just fine. Until it's changed.
Literally lol'd. It's "chaos monkey".
Some events (or "changes") I've seen in last year that caused a system to stop running fine:
- analyst ran query that locked reporting DB, which had ripple effects
- integration partner "tweaked" schema of data uploaded to us
- NIC started dropping packets
- growing data made backup software exceed expected window, locking files
- critical security patch released (didn't cause a problem exactly, but did change the system from having no known security holes to having a big one)
On and on. So, yeah, I'm not disputing the idea that changes are often the source of problems. I'm just saying that any moderately complex system is constantly encountering new events whether or not the developers are making changes.
This isn't a rebuttal of your statement so much, as it is rebuttal of a common attitude I see in business folks. A lot of non-technical executives seem to have the mentality that software is "done" after it's built, which is naive IMO.
Active systems require active maintenance. You can avoid a lot of problems with intelligent architecture and robust instrumentation/monitoring, but at the end of the day systems rot and will eventually stop running fine.
And even if your perfectly planned and architected system runs in total isolation on a private server where security isn't a major concern, you'll still build up technical debt if you aren't routinely upgrading major libraries, etc. You'll get to a point where you want to implement some feature, and while there are 3 excellent OSS libraries for your platform to accomplish it, none of them is compatible with your 5 year old version.
You probably know all of this and might even agree with most of it. I just had a visceral reaction to the idea that "a system running just fine will usually continue to run just fine," and felt compelled to respond. :)
Note: most of my experience is with public or enterprise web applications. I imagine other types of systems have other problems.
http://www.darkcoding.net/software/facebooks-code-quality-pr... - "Exhibit C: Our site works when the engineers go on holiday"
Compare: "bomb explosions tend to increase around the time when we send in the bomb squad. I guess bomb defusing is pointless."
IOW: Surgeries and bomb defusing tend to move forward a lot of the death-probability-mass (while, in theory, destroying some of it).
At minimum, you've got keeping up with security patches, library/framework deprecations etc. Software which has literally not been changed in years is often an insecure timebomb, with unpatched vulnerabilities, and as soon as you do need to change it, you're stuck because all its dependencies are deprecated. Besides, you also need to keep up with the market, competitors, customers' usage patterns changing, etc.
It would be nice if we as customers had the freedom to pick up which upgrades to install, and to limit ourselves to security patches and (maybe, some) library deprecations. Windows upgrades let you do just that, even if it is a pain to manage. In my previous job, we held monthly meetings to discuss which windows patches to install in our dozens of business critical Win2k3 servers.
For most software, upgrades are opaque. In the best case, you are given a binary choice: install now or install later. In many cases you dont even get that, either the product stops working until you apply the patch or the patch gets automatically applied without your explicit consent.
Which would itself not be a problem if most upgrades did what they are ment to do. Instead, as the GP suggest, patches introduce new (sometimes critical) bugs all the time. With pure repair patches, the balance is mostly positive siince they fix more bug than they break, but with every new feature (which usually is required by a minority of very vocal customers) comes a risk that some other more important feature is going to stop working as expected.
Industry wide, we lack the discipline to change things in a responsible way. Change management is today what source version control was back in the 90s: something that most people has heard good things about, but most people is not doing correctly, and a significan minority not at all.
If it is 1000 microservices as in different apps, then they must have at least 2000 running apps (at least 2 instances per app for HA).
Maybe uber only have 200 "active" microservice app running at the same time where each microservices have N running instances.
I just cant imagine running 1000 different microservices (e.g different docker images, not docker instances) at the same time.
Anecdotally, I've worked on services that ran tens of thousands of instances across the world. You build the tools to manage them and it works very well.
Microservices can go to far. I'm very thankful for this video.
People talk about interpreted languages being slow sometimes, now your program is divided over 1700 separate servers and instead of keeping variables in memory you have to serialize and send them over the network all the time.
But since asking the question I've realized that if your application already needs a huge amount of servers because it simply gets that much traffic, then putting something like this in its own docker instance is probably the simplest way (it might even use postgres inside it), if those boundaries change now and then.
But most companies aren't near that scale.
- Select(db, somekey, someparameters) [return some db object]
- http_get_query(http://service.com/somekey/someparameters [return some JSON]
They are external (micro)service:
- they both need the target system to be available.
- they both may fail in weird and unexpected ways.
- they both need to handle failure gracefully.
Their usage have different properties:
- A database call need to have a permanent connection pool to the database, usually requiring db user and password.
- A http call is just call and forget. It's a lot easier to use, in any applications, at any time.
A better question would be why not write a module or class? There a pros and cons to either, but advantages include: better monitoring and alerting, easier deployments and rollbacks, callers can timeout and use fallback logic, you can add load balancing and scale each service separately, it's easy to figure out which service uses what resources, it makes it easy to write some parts of application in other programming languages, different teams can work on different parts of the application independently as long as they agree on an API.
E.g. I can't think of a REST service doing this; something like a direct socket connection (over a HA / load-balancing level) with a zero-copy serialization format like Cap'nProto might work.
Time to forbid adding more stuff and start cleaning.
(12) [...] perfection has been reached not when there
is nothing left to add, but when there is nothing
left to take away.
https://tools.ietf.org/html/rfc1925What number is that?
I would love to hear a breakdown. This sounds like a nightmare to maintain and test.
left-pad-as-a microservice
You can also have another service to monitor the left-pad service.And yet another one to ensure that the log-aggregation-monitoring-service is working.
It'll be interesting to know the number of people oncall at any given time and the number of prod alerts per hour/day/week.
I don't know if they provide popcorn at the meetings where the project managers explain why they deserve the next sprint.
(Snark exists in this comment.)
It's because they hire smart and ambitious people, but give them a tiny vertical to work on. It's a person level problem, not a technical one in my opinion.
I think you solve this by designing your stack and then hiring meticulously, instead of HIRE ALL THE ENGINEERS and then figure it out (which is quite obviously uber's problem)
Microservices address that gap.
And in the process the field is transformed from one of software developers to software operators. This generation is witnessing the transfer of the IT crew from the "back office" to the "boiler room".
Personally, I vastly prefer debugging and building on non-microservice architectures than on something split willy-nilly into dozens of services (because most implementations I've seen don't do microservices with clean conceptual boundaries - it's more political/organizational divisions that determine what lands where, not architectural concerns).
That's one way to scale development projects.
Have multiple teams of dev who work on separate stuff. They can develop in their corner however they want and that makes them happy. The final thing is a clusterfuck of services with little cohesion [YET they ALWAYS manage to put the shit together in production in a mostly working state].
The alternative implies to have a consistent and cohesive view of the components. For that, you'll need to have people with experience in architecturing scalable and sane production systems [so they understand what are the consequences and tradeoffs of ALL their decisions]. I know of very few people on the planet who can design systems. (We're talking past unicorn-rarity here). Plus, the developers must actually listen AND care about the long term maintenance AND the people involved (i.e. not want to do shit that will hit either them in 6 months or the team in the next office).
The amount of collaboration AND communication AND skills required to operate this strategy is beyond the realm of most people and organisation. There are very few individuals who can execute at this level.
I honestly think that the rise of microservices as "the technical cure for cancer and everything else" is a very interesting surfacing of the various systemic disfunctions of our beloved software industry. The disfunctions have been present from day 1 but the environment has somewhat radically changed.
Personally I much prefer microservices for debugging. You can quickly identify which one is the problem then test the APIs in isolation pretty quickly. Sure beats having to wait 20 minutes for a monolithic app to build.
That's the sort of thing I was including in operational complexity, not just the "are the packets making it through" stuff.
This sounds a lot like the code coverage fallacy. (to which I usually answer "call me when you have 100% data coverage").
https://gotocon.com/dl/goto-chicago-2016/slides/MattRanney_W...
Any speculation as to why Uber doesn't just want to use something like Netflix Eureka / Hystrix instead?
That's also why that makes it an interesting place to work at and helped them achieve this growth.
Personnaly I think this should be made into a global app with no geo-fencing (e.g. Available everywhere basically).
I used to work in startups, and overall was impressed with velocity. Then I joined a big valley tech company, and now I understand.
It's because they hire smart and ambitious people, but give them a tiny vertical to work on. On a personal level, you WANT to build something, so you force it.
I think you solve this by designing your stack and then hiring meticulously with rules (like templates for each micro service), instead of HIRE ALL THE ENGINEERS and then figure it out (which is quite obviously uber's problem)
Good talk, will likely watch again.
Know your data. Are you serving ~1000 requests per second peak and have room to grow? You're not going to gain much efficiency by introducing engineering complexity, latency, and failure modes.
Best case scenario and your business performs better than expected... does that mean you have a theoretical upper bound in 100k rps? Still not going to gain much.
There are so many well-known strategies for coping with scale that I think the main take-away here for non-Uber companies is to start up-front with some performance characteristics to design for. Set the upper bound on your response times to X ms, over-fill data in order to keep the bound on API queries to 1-2 requests, etc.
Know your data and the program will reveal itself is the rule of thumb I use.
Basically if you're building a hypermedia REST API you return an entity or collection of entities whose identifiers allow you to fetch them from the service like so:
{"result": "ok!"
"users": ["/users/123", "/users/234"]
}
The client, if interested, can use those URLs to fetch the entities from the collection that it is interested in. This poses a problem for mobile clients where you want to minimize network traffic... so you over-fill your data collection by returning the full entity in the collection. {"result": "ok!"
"users": [{"id": 123, "name": "Foo"}, {"id": 234, "name": "Bar"}]
}
The trade off is that you have to fetch the data for every entity in the collection, the entities they own, etc; and ship one really large response. The client would receive this giant string of bytes even if the client was only interested in a subset of the properties in the collection.GraphQL does away with this problem on the client side rather elegantly by allowing the client to query for the specific properties they are interested in. You don't end up shipping more data than is necessary to fulfill a query. Nice!
... but the tradeoff there is that you lose the domain representation in your URLs since there are none.
Microservices is all about taking a big problem and breaking it down into smaller components, defining the contracts between the components (which an API is), testing the components in isolation and most importantly deploying and running the components independently.
It's the fact that you can make a change to say the ShippingService and provided that you maintain the same API contract with the rest of your app you can deploy it as frequently as you wish knowing full well that you won't break anything else.
It also aligns better with the trend towards smaller container based deployment methods.
It's an aspect. It's often the beginning of a micro-service migration story in talks I've heard.
> Microservices is all about taking a big problem and breaking it down into smaller components
... and putting the network between them. It's all well and good but the tradeoffs are not obvious there either. Most engineers I know who claim to be experts in distributed systems don't even know how to formally model their systems. This is manageable at a certain small scale as some computations take 35 or more steps to reveal themselves. But even the AWS team has realized that this architecture comes with the added burden of complex failure modes[0]. Obscenely complex failure modes that aren't detectable without formal models.
Even the presenter mentioned... why even HTTP? Why not a typed, serialized format on a robust message bus? Even then... race conditions in a highly distributed environment are terrible beasts.
[0] http://research.microsoft.com/en-us/um/people/lamport/tla/am...
You just have to know your data. The architecture comes after... and will change over time.
You don't need to be an expert in distributed systems to use microservices. It's literally replacing a function call with an RPC call. That's it. If you want to make tracing easier you tag the user request with an ID and pass it through your API calls or use something like Zipkin. But needing formal verification in order to test your architecture ? Bizarre. And I've worked on 2 of the world's top 5 ecommerce sites which both use microservices.
And HTTP is used for many reasons namely that it is easy to inspect over the wire, is fully supported by all firewalls and API gateways e.g. Apigee. And nothing is stopping you using a typed, serialized format over HTTP. Likewise nothing is stopping you using an message bus with a microservices architecture.
Until you want to make five function calls, in a transactional manner.
Architecting correct transaction semantics in a monolithic application is often much easier then doing so across five microservices.
No basis at all? I knew I was unhinged...
> You don't need to be an expert in distributed systems to use microservices.
True. Hooray for abstractions. You don't need to understand how the V8 engine allocates and garbage collects memory either... well until you do.
> It's literally replacing a function call with an RPC call.
You're not wrong.
Which is the point. Whether for architectural or performance reasons I think you need to understand your domain and model your data first. For domains that map really well to the microservice architecture you're not going to have many problems.
And a formal specification is overkill for many, many scenarios. That doesn't mean they're useless. They're just not useful, perhaps, for e-commerce sites.
But anywhere you have an RPC call that depends on external state, ordering, consensus... the point is that the tradeoffs are not always immediately apparent unless you know your data really well.
> And I've worked on 2 of the world's top 5 ecommerce sites which both use microservices.
And I've worked on public and private clouds! Cool.
The point was and still is the same whether performance or architecture... think about your data! The rest falls out from that.
It absolutely helps with people and org scalability. I haven't seen it help with technical load scalability (assuming you were already doing monolithic development relatively "right"; we ran a >$1BB company for the overwhelming majority on a single SQL server).
I wrote about this a couple of years ago in more detail, so I'll just reference that: https://www.chrisstucchio.com/blog/2014/microservices_for_th...
HN discussion: https://news.ycombinator.com/item?id=8227721
With micro services you could have different versions of those utilities in use which would not be possible in a monolithic app.
You have no idea...:)
Alternatively, you can deploy the new version to many machines, test it, then make your load balancer direct the new traffic to the new instances, until all old instances are idle, and then take down these.
code > function/method > module > class > ... > microservice
So yes it is heavyweight. But it also works.
Of course, this is context sensitive and not everyone has easily decouplable code bases, so YMMV. But not recognizing the myriad purposes isn't useful for criticism.
At perhaps 15, breaking up into groups of 5 (or whatever) lets small groups build their features without stomping on anyone else's work. It cuts, not only coupling in the code, but coupling in what engineers are talking about.
The two sort of obvious risks, those teams probably should be reshuffled from time to time, so code stays consistent across the organization. If that's done with care, specific projects can get just the right people. If it's done poorly, you sort of wander aimlessly from meaningless project to meaningless project.
The other one is overall performance and reliability. When something fans out to 40 different microservices, tracking down the slow component is a real pain. Figuring out who is best to fix that can be even worse.
Do the complexities introduced not have a cost, i.e. are those marginal returns offset by the choice in the first place? Call it "platform debt."
It's probably not worth discussing unless you're actively feeling either perf or coupling debt pressure.
I've thought about this problem before, both for a related problem space and with friends working in this specific space. The short version is to create a grid where each square holds car data (id, status, type, x, y, ...) in-memory. Any write-lock of concern is only needed when a car changes grid, and then only on the two grids in questions. This can be layered multiple levels, and your final car-holding structure could be an r-tree or something.
The grid for a city can be sharded across multiple servers. And, if you told me that was necessary, fine..but as-is, I'm suspicious that a pretty basic server can't handle a tens of thousands of cars sending updates every second.
Friends tell me the heavy processing is in routing / map stuff, but this is relatively stateless and can be sent off to a pool of workers to handle.
The routing is very complex too but as you note scales well, until you want to start routing/pickups based on the realtime location of other cars.
- How do you distribute the car actors on nodes, assuming the number of nodes is variable? (I think riak_core looks interesting, but it does not seem to have a way to guarantee unicity of something - it's rather built to replicate the data on multiple nodes for redundancy)
- What happens if a node fails? What mechanism is going to respawn the car actors on a different node? How do you ensure minimal downtime for the cars involved?
- What happens if there's a netsplit, e.g. how do you ensure no split brain where two nodes think they're responsible for a car?
It feels to me like the traditional erlang-process-per-request architecture coupled with a distributed store (riak or w/e) makes it possible to avoid the very difficult problem of ensuring one and only one actor per car in a distributed environment.
I think you are right about the riak core stuff - you could probably keep track of cars using some sort of distributed hash and kill multiple cars if they were to ever spawn.
In fact a way of instantiating processes and finding them based on a CRDT is probably a pretty cool little project...
That describes no system, ever, including Erlang. TANSTAFL
Although... Ethernet latency would probably make it tough to stick to 2ms.
EDIT: to downvoters: why shoot the messenger? :)
Is that strange? I've got 6 repos at work for various utilities and "personal" projects I'm working on in addition to the 4 repos for team wide projects.
Edit: I remembered more personal projects, make that almost 3 dozen personal Git repos.
Or maybe some instruction into how to assemble the organization.
10 git repositories per microservice
I'm reminded of http://danluu.com/monorepo/On-call shifts are going to be interesting with an average of 2.5 engineers per service, not to mention handling people switching teams or leaving the company.
They did that; the reason is job security.
But I can't imagine how you could possibly need 8000 git repositories unless you're massively over engineering your problem. Project structure tends to reflect the organizations that build them.
Analytics can run in batch mode. Also it is read-only operation. So it is far simpler than designing a distributed database.
As an aside, why would Uber want to reinvent a distributed database for storing user information?
The things you highlighted are very light load comparatively. Sign in, load account info, then dispatch to location specific shard.
REST/JSON with strong schemas (and loose where that makes sense), using Postgres and C++.
Sure, in the long term things should be refactored, structured, and optimized. But if you do that too soon you risk locking yourself out of potential value, as well as gold-plating things that aren't critical.
How often does that really happen though? Once you've amassed enough technical/data debt, resistance to refactoring increases until it never happens at all. Having well defined, coherent data models and schemas from the start will pay off in the long run. Applications begin and end with data, so why half-ass this from the get go?
I believe that if you aren't extremely certain about what the future holds it may be best to work with a more flexible technology first and transition to a more structured setup once you have solved for your problems and identified intended future features. And if you are extremely certain about what the future holds you're either insanely good at your job or just insane.
Of course there are other considerations. A more "planned" structure always makes sense if you're talking about systems or components that are life-critical or that deal with large flows of money. The "fast and loose" approach makes the most sense when you can tolerate occasional failures, but you have to have fast iterations to be quick-to-market with new features.
500KLOC JVM/.NET application? No big deal.
50KLOC JS/HTML-based SPA? Pfooooh. That could take a while...do we really need to?
Oh, and my servers are in Ruby.
Sometimes it's easier for me to alter the JSON payload a certain way in the frontend, then the python backend handles it and saves it to the DB.
The RDB helps not to stray too far into crazy-land, while python and JSON gives you the flexibility to prototype and experiment.
Granted, I am neither a database nor ORM savant, but I find that it makes explicit almost as easy as implicit - but with more safety! I haven't seen that elsewhere, but I haven't looked very hard either. I have heard claims that Groovy/Hibernate do this just as well as well, but it isn't clear to me that this is completely true.
That will probably confuse my procedural mind, much like declarative "stuff" often does! :-)
I suppose folks in this camp would overlap significantly with those in camp 2, though.
I like schemas ... most of the time! I like Avro ... some of the time! And JSON some of the time. And I write mostly in Python.... and Scala.
The world is not so black and white.
Yes, and no.
Yes: REST / JSON is nice. I've used them widely as a kind of cross-platform compatibility layer. i.e. instead of exposing SQL or something similar, the API is all REST / JSON. That lets everyone use tools they're familiar with, without learning about implementation details.
The REST / JSON system ends up being a thin shim layer over the underlying database. Which is usually SQL.
No: databases should NOT be flexible, and "NoSQL" has a very limited place.
SQL databases should be conservative in what they accept. Once you've inserted crap into the DB, it's hard to fix it.
"NoSQL" solutions are great for situations where you don't care about the data. Using NoSQL as a fast cache means you (mostly) have disk persistence when you need to reboot the server or application. If the data gets lost, you don't care, it's just a cache.
You can make your schema very light and accepting almost like NoSQL which is how you get into the situation you described; the solution is to use stricter schema. That and it helps to hire a full time data engineer/administrator.
> Using NoSQL as a fast cache
I'd rather use caching technology, specifically designed for caching, like Redis or Varnish or Squid.
Agreed, it is hard to fix. NoSQL databases can be really hard to fix when they are full of crap, too.
Properly encoding rigidity through type systems, SQL checks/triggers/conditions, etc is hard. It takes a really long time to really iron out all the string and integer/double/long typing out of your system, let alone do it in a way which matches up properly with your backing datastore. Once you've got it set up and nailed down with tests, then you're golden, but that's a cost that is usually not worth paying until long-term need is determined.
Rest/JSON is a well understood, broadly adopted, low friction RPC format.
NoSQL is not always MongoDB (for example, Google Datastore is ACID compliant), and schema enforcement via an ORM layer I would argue is actually a good thing, as it provides schema validation at compile time.
Initiatives like JSONSchema go some way to restoring some constraints to the format and prevents unchecked deviation over time.
1. new business evolving quickly to meet and discover the product that fits the market they are chasing.
2. established and optimizing for a possibly still growing market but very well established set of features and use cases. They can take longer to deliver new features and can save lots of money by optimizing.
What I really liked about Thrift is that all I needed to know how to use the service was the thrift definition file. It was self-documenting.
My current project uses HTTP and JSON to communicate from the public API to the backend services. There is significantly more overhead (latency and bandwidth) and no enforced document structure (moving toward Swagger to help with that).
HTTP+JSON is great for the front-end where you need a more universally parsable response, but when you control the communication between two systems, something like Thrift/Protobuf solves a lot of problems that a common with REST-ish services.
'NoSQL' is just a broad term for datastores that do not normally use standard Structured Query Language to retrieve data. Most NoSQL do allow for for very structured data as well as some query languages that are similar to SQL.
BigTable (Hadoop, Cassandra, Dynamo) and block stores (Amazon S3, Redis, MemcacheD) are absolutely critical to cloud services. Json tuple document DBs are needed for mobile and messaging apps. Graph is for relationships, and Marklogic has an awesome NoSQL Db focused on XML.
Full disclosure: I am the founder of NoSQL.Org - but I also use multiple relational SQL databases every day.
Eg rolling out a public API in Thrift/protobuf will severely difficult its adoption, whereas Rest/JSON is pretty much the standard - but then building microservices that communicate to each other in Rest/JSON quickly leads to a costly, hard to maintain, inconsistent mess.
We're obsessed with "one fits all" absolutes in tech. We should have more "it depends" imho