[1] http://techblog.netflix.com/2012/02/fault-tolerance-in-high-...
If you are making a big system, you probably would not want to try to put all of the components in one monolithic place. I argued that this makes testing and release more complicated (and slower). But, I suspect it also affects runtime speed by making the stack deeper. You would be doing more switching to get your requests to the right components.
So if n=3, it doesn't seem to bad. How much can you do in parallel? Stuff like search, it seems like you could partition quite a bit hit a bunch of services (10?)in the first 5ms, then spend the remaining 10ms sorting results, just send your best effort.
I'm (of course) assuming a customer request means there's a customer sitting there waiting and 30ms is good enough. Won't work for a game. Won't work for HFT.
I guess you'd use the standard stuff, many layers of caching, efficient implementation of the specific services, good algorithms. Heck, you could play a game with docker and run the whole infrastructure on 1 big machine, i think those local sockets would be a lot quicker that the 500,000 ns of a datacenter read.
I mean, if an l2 cache miss is a disaster for you, clearly this won't do anything for you. But how fast is fast enough? do you need to hit that target every single time(real time)? Can you be late 1 in a million requests?
Are you sort of dismissing the idea out of hand or are you (aside from the silly title) really thinking through the consequences of implementing an application this way?
Since each service follows a request/response format, it becomes easier to see where the majority of a request lifespan spends its time to help eliminate bottle necks. You'd be surprised how low the overall latency for a request can be in large companies like Amazon (< 100ms)