Towards Modern Development of Cloud Applications (2023)
dl.acm.org
dl.acm.org
> The call to hello.Greet looks like a regular method call
That’s a departure from how components interact in boq — an internal and widely used production platform that has _some_ of the features from the paper. There component interfaces _are_ RPC interfaces (e.g., Stubby / gRPC + protocol buffers), and interaction between them is possible exclusively through the component interfaces. Hence it’s very explicit at the call site that an RPC is being made (which could happen to execute locally with all the standard RPC functionality — context and deadline propagation, etc.).
RPCs looking like regular method calls sound a bit scary (easy to miss in code reviews); I wonder if enforced naming conventions + IDE + code review tool support would be enough.
Edit: it seems to require to pass a context object, so the readers won't confuse it with a local call (from https://serviceweaver.dev/):
sum, err := adder.Add(ctx, 1, 2)
---Also, the paper claims that most benefits come from a non-versioned serialization format:
> Most of the performance benefits of our prototype come from its use of a custom serialization format designed for non-versioned data exchange [...]
However, I don’t understand why local RPC calls have to serialize protocol buffer messages — can’t they already pass them as-is to the local handler?
(disclaimer: a googler, no internal knowledge on ServiceWeaver)
I didn't read the paper in enough detail to know the answer to this, but mightn't this enable different implementation languages for different components? In my experience, it's difficult to accomplish reliably that without using a language-agnostic serialization format (like proto).
Even if that's the goal, it seems like a handler could determine whether it could elide the serialization depending on the implementation details of the components.
However, now that I think again about the serialization format choice, it may result in a limitation on the size of monoliths (in terms of the number of people / teams contributing to it). When the number of contributors grow, the likelihood of bugs in a binary grows, and teams adopt more elaborate qualification processes, and also become much more sensitive to binary rollbacks as a remedy to discovering bugs in prod. Then they could institute policies like all changes should be protected by a feature flag (aka an experiment).
If non-versioned serialization format is used, that means that the platform cannot possibly rollback a single component. However, using versioned serialization won't be enough on its own to support per-component rollbacks — it at least requires independent component qualification (where each component is tested against "stable" versions of other components) + rollback testing to make rollbacks A2 -> B2 to A2 -> B1 safe.
I wonder if it's an explicit design choice — i.e., whether Service Weaver supports monoliths up to a certain organizational size (and then you should split into separate service weaver apps)?
Let's say you start from a codebase with three portions (call them services, modules, whatever): A, B, C, D. A sends a (synchronous) remote procedure call to B, which sends a message over a message bus which is also used by C and D. C and D do not talk to each other except over the bus.
It sounds like this approach would identify the remote call dependency between A and B (which could be split into different deployment units), but not the message bus usage. Or, at least, it can't identify who is subscribing to a topic where B pushes its events.
As a result, you would get two deployment modules:
- A
- B and "everything else"
Which doesn't sound right.
Am I missing something?
They don't have to, they artificially added that constraint to make the benchmarks fairer.
5ms protobuf, 2ms custom, 0.4ms in-process
CORBA had the same issue. The call could take 10us or 10s and no way to tell by the user. This was ofc widely considered as huge design flaw.
It wouldn't be too surprising if this is the sort of thing that an optimizing compiler or query planner could do better than a human.
If not, you're probably at the scale where performance regressions are caught and rolled back at early phases of rollout.
Like most magic, it's either going to make things 100x better or 100x worse, depending on how leaky the abstraction is at its current state of maturity.
Always the afterthought. At a previous job we had a bunch of microservices developed on peoples' computers, and they only place they all ran together was the handful of integration environments.
In previous projects I defined systems in a single code base, and parts could be deployed separately by providing different configuration files.
It was a very productive approach: one could run the whole system in a single process during development, which made writing integration tests a lot easier than spinning up dozens of docker images.
It still required some manual work, and deployments were still to static for my liking. Ideally it should be possible to split off and scale subparts dynamically.
In the Clojure world there are several projects now that explore splitting up services in a transparent way.
For example Electric Clojure splits up your code into frontend and backend parts, making the frontend-backend split transparent. Another project is Rama, which does something similar but for distributed steam processing and partitioning.
I'd love to explore something like this but for enterprisey service meshes: the programmer just defines services, and a compiler decides how to split these over different machines, and all the RPC/serialization/deserialization is done for you.
What does it all mean. Well it's a technology built for Google scale. It may have merits in other place as a lot of tech has done, but at the same time, for 90% of teams this doesn't matter. You have a monolithic code base in a single repo and you can deploy and vertically or horizontally scale quite easily depending on your requirements. For companies that are 200+ engineers split across 15-20 teams this might matter. They already be doing some sort of microservices or service splitting while still using a monorepo. Being able to remove a lot of platform level code that you manage versus it being an open source thing is advantageous because you can go back to focusing on the business case not the glue code.
That is only people getting the point parsing JSON and XML all over the place doesn't scale and there is a reason why SUN-RPC, DCE, CORBA, DCOM, Java RMI, and .NET Remoting existed in first place.
What is old is new again
The paper focuses on microservices, and then tries to avoid claims of "they just don't like microservices" by describing the ways in which microservices are improperly used. Do they go back and compare this to monoliths or other architectures? Nope; it's really just "hey I have another microservices idea", heavily gilded. They mention "monolithic applications divided into logically distinct components", but you could just claim your microservices are divided into logically distinct components.
They also seem to completely ignore the problem that a logical separation doesn't mean your components are better off. In a complex system, often completely separate components still need to be integrated together in order for the system to function at all, much less operate efficiently. It's not a design flaw to combine different things. It depends on the application. So just separating things logically isn't some scientific computing advancement, it's just categorization.
In reality, their solution (a "single binary business logic application" and "an interface that can combine them") is literally a description of shell scripting with Unix tools. Don't get me wrong, that obviously works great, since it's been popular for 44 years (older than IPv4). But if you want to come up with some kind of modern paradigm for distributed computing, maybe we should flush it out a bit more. What we have here is a Google engineer's attempt to make a paper suggesting we make shell scripting for the web, without much to show for it.
(Personally, I think the more people try to control the interface, the worse things get. The best and most long-lived solutions in all of computing have had almost no interface at all; a raw TCP stream, 3 raw file descriptors, a set of random arguments, and a set of random key=value pairs, have enabled all modern computing paradigms to flourish)
- Making remote calls seem like they are local resulted in poor design decisions, the benefit of SOAP/REST was that people considered what the interface a useful service should be.
- Why not flip it and look to move groups of microservices onto the same machine, updating how the app communicate. - if component boundaries are fine grained, the combinations of local/remote services relative to each other increases, along with the testing burden; just because the system hides remote deploying, it still should be tested for.
- Incorporating this with storage, eg dynamic shard rebalancing would be super cool
Because it is a bit buried in the paper, this is the prototype implementation they talk about.
From the article:
> The Deployment Model is a Detail.
> If the code of the components can be written so that the communications mechanisms, and process separation mechanisms are irrelevant, then those mechanisms are details. And details are never part of an architecture.
> That means that there is no such thing as a micro-service architecture. Micro-services are a deployment option, not an architecture.
It is an interesting idea, but I'm not sure I'm fully convinced.
Sure, you can parse remote calls and package imports to find dependencies. But this assumes applications are integrated using remote procedure calls. What if applications talk to each other using a message bus? Asynchronous patterns are meant precisely to decouple publishers and subscribers - in number (how many processes will read a message?), identity (who are those processes?) and time. It looks like this system would not be able to decouple portions of code talking to each other using a message queue.
Any thoughts?
I don't think it would be very difficult to replace the poll on the database table with a pull from SQS instead.
Performance is evaluated against a single example implementation (section 6.1) and 9x improvement was achieved when co-locating into a single process. To be fair this seems a reasonable thing to do if you can get the application to be efficient enough to allow it - but there may be very good reasons not to do it in many applications.
It seems to reason that if you can co-locate your calls into a single process, you'd gain at least 9x.
This intuition is what encouraged us to run with SQLite in production back when most developers didn't think it was a good idea.
My main critique is that it's one example, and it would be good to see the technique exercised across a number so that we can see the strengths and weaknesses.
Maybe it is explained somewhere in the paper but how do atomic updates work if at least one of the components has to be a singleton that does not support rolling upgrades?
Towards Modern Development of Cloud Applications [pdf] - https://news.ycombinator.com/item?id=38146809 - Nov 2023 (10 comments)
what a lovely time !
For example, FoundationDB does exactly this (via dynamic assignment), as do many other databases. All of the HashiCorp runtime tools also do it. I’m sure there are also much earlier examples.
Reading the paper two thoughts come to mind:
- "What's old is new again"
- "those who do not learn from history are doomed to repeat it"
There were several attempts in history to implement transparent RPC (new again). All failed due to https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu... (learn from history)
Looks like any abstraction trying to hide distribution is inherently too leaky.
or as mark-twain would (probably) say
"History Does Not Repeat Itself, But It Rhymes"
Need to get that VC money into the hot WASM space.
But you're right - today's world of YAML programming is no better.
Nix and Guix are steps in the right direction (but both have their fair share of issues)
The paper unfortunately hides that in reality you have to pass a context object in your RPC calls, hence there is no ambiguity whether you are calling a potentially remote object.
It's in the example on the project home page: https://serviceweaver.dev/
// The "RPC" handler
func (adder) Add(_ context.Context, x, y int) (int, error) {
return x + y, nil
}
// The call-site
var adder Adder = ... // See documentation
sum, err := adder.Add(ctx, 1, 2)> Java RMI use a programming model similar to ours but suffered from a number of technical and organizational issues [58] and don’t fully address C1-C5 either
Apart from that, it looks like Java RMI allowed remote objects returning other remote objects, rather than only immutable values. With that you could abuse it by making a call to one java.rmi.Remote object, getting another java.rmi.Remote object in response, then passing it around, and then finding a totally different subsystem suddenly make RPCs (however, such abuse probably would be easy to spot in a code review, as it requires a modification to the remote object interface).
---
The authors also acknowledge that it doesn’t solve the distributed computing challenges:
> our proposal does not solve fundamental challenges of distributed systems [53, 68, 76]. Application developers still need to be aware that components may fail or experience high latency
I think at least in terms of latencies their platform can occasionally inject latencies into some percentage of the tasks, then verifying if any alerts fire, if there is a fear the components become dependent on a certain deployment shape (within a cluster).
Actually this is a very powerful concept as it allows one to achieve high level of reuse. Jini (https://jan.newmarch.name/java/jini/tutorial/Jini.html) made mobile objects the core idea of its architecture.
How powerful the approach is one can see when looking for example at Jini concept of a Lease and LeaseRevenewalService:
a server program can register an object (client side implementation of a service) in ServiceRegistrar (to make it discoverable and downloadable). Registration is lease based so the server has to renew it periodically. But it can be delegated to a LeaseRenewalService that can do it on its behalf so that the server can go to sleep (ie. not use any server machine resources).
All of the above happens without any party a-priori knowledge about any code that needs to be present at use site - code is downloaded automatically on-demand - the only thing common to client and service is a Java interface.
Basically your code has no concept of "network boundary", there's only package import each other, there's no "microservice".
But when deploying, i can choose which package to be deployed as a service.