Everything Is an X
lukeplant.me.uk
lukeplant.me.uk
If you think about it, everything can be described as a system with an input, some buffer (queue), a bunch of workers doing work in parallel, then output: * webserver (e.g nginx) * database server * network interface in your computer * all the networking hardware in between * the whole disk IO stack in your OS * cpu (instructions queue)
I wrote about it [1] at some point in more detail if anyone's interested.
[1] http://blog.dfilimonov.com/2020/04/24/everything-is-a-queuei...
* State between communication. If service (worker/function) A needs to talk to B, A usually needs to keep some state until B answers. Maintaining this state (what if the message gets lost, what if we redeploy A, what if the storage fails) adds a new dimension IMO.
* Circular dependencies/recursion. It's astonishing how things like A communicating with A, directly or indirectly seems to be implicitly missing from that typical architecture.
* Message growth. What happens when you have an O(n^2) or worse growth in messages? How do you track and manage that?
Generally speaking, it looks like microservices behave a lot like the early procedural programming languages by not considering the actually complicated stuff.
The alternative to asynchronous interfaces (which is basically what GP describes) are blocking interfaces - like system calls, or just waits for a specific event. And these are never an answer, unless it is _guaranteed_ that the blocking call will return withing a given timeframe and you also know that you will have absolutely nothing else to do meanwhile.
> State between communication.
There are two ways to keep state - locally on the stack or in an explicit data structure. As always in programming, you have to clean up when you destroy an object/process/stateful thing.
> what if the message gets lost
This absolutely should not happen, unless the sender doesn't expect an answer and the message can clean up itself. The latter is the case for example when sending a message means just copying it towards the destination, like in computer networks. The former is the case in particular in UDP connections.
> what if we redeploy A, what if the storage fails
You absolutely need to answer all messages that require an answer (called "IO completion" elsewhere). Of course, the answer can be "cancelled" or "failed".
Before destruction or reset, you need to synchronize with all users that hold a direct handle to the object being destroyed. That could just be done by having only a single owner who is responsible to wait for a "cleaned up" event and to then destroy the object. Think Unix processes - processes that exited still appear in the process table until their parent has wait()ed for them.
> Message growth. What happens when you have an O(n^2) or worse growth in messages? How do you track and manage that?
In general, asynchronous IO is achieved with queues (which are what GP discussed). With queues you have the choice to limit their size right in the queue (if the queue is full, block sending, or reject it temporarily and notify when there is progress). Or you can allow unlimited queues and push responsibility for memory management (and allocation policies) to the users of the queue. For example, users can allocate messages on their own and just link them into the queue - not additional memory allocation needed.
+1 In my experience “modern distributed systems” is codeword for handwaving away the complex questions like transactionality and consistency and thinking mostly about the happy path behaviour.
and Neil Gunther's other works.
[1] https://engineering.linkedin.com/distributed-systems/log-wha...
They tend not to care about the collateral damage or nuances, they want to order the software according to their beliefs. Once you recognise this you can see the same pattern repeating in a lot of other non-software disciplines. I guess it's just a personality type.
https://en.wiktionary.org/wiki/if_all_you_have_is_a_hammer,_...
In big companies this is often less "everything is an <X>" and more "By strongly delineating boundaries, I can prevent you from screwing me over".
If you have to communicate with my pieces as a service, you have to tell me what you need and you can't blame me for not delivering it when you didn't tell me. You can't say I'm blocking you if I can point that my service is up and answering. etc.
Love how the article can be interpreted as making statements about itself. The article states everything on this list is an implementation of 'everything is an X'. Or put differently, everything is an X, where X = 'everything is an X'.
Or in my own house, now that I have kids of my own: "everything is sticky".
But (a little) more back on topic, I find a lot of open source software to be: "everything is a missing dependency hunt".
Some days, C++ feels like: "everything is undefined behavior".
Everything is being reinvented again.
That's interesting. Has anyone heard of such a meta-relational database?
Schema-as-data is a central part of the relational model, and is implemented (usually with some limitations) in most RDBMSs.
OTOH, DDL manipulation is usually more convenient than DML against catalog tables, even when you can in theory acheive the same result either way.