Data Parallel, Task Parallel, and Agent Actor Architectures
bytewax.io
bytewax.io
Difficult patterns to optimize for is my experience if any IO or network boundaries are in play. Even that and it sometimes seems the more knowledge on how threading works on one particular cpu architecture.
So for me at least, on a high level these patterns seem easy enough and seems to play well in to the popculcural small service fad and large data processing. But be aware of the underlying cpu architecture and threading and how thread pools work on your particular OS.
Oh yes and then comes the debugging and reading the code part, which we all know are where the real efforts of time comes to play.
Use these when absolutely no other options are available. Just like multithreading.
Noob question: Can 90%+ of the "effect" from the above not be replicated or be done with Go-Lang's CSP (or another language that also uses CSP) ?
Is Golang or CSP model perfect ? Of course not, but It it is "relatively" easy to reason about compared to "catalogue of other options" ?
var task2 = service.DispatchSecond(param2);
var final = service.DispatchThird(await task1, await task2);
or
var queries = users.Select(user => FetchPurchases(user.Id));
var results = await Task.WhenAll(queries);
Very easy.
As long as you don't touch System.Threading.Tasks.Dataflow namespace, everything will be good.
https://pragprog.com/titles/pb7con/seven-concurrency-models-...
What I have never used, but is inferred from https://en.wikipedia.org/wiki/List_of_concurrent_and_paralle... and in special the original CSP is how all this lacks total control in build your network patterns.
In special, is my dream to have a language that allows to express better what is explained starting here:
https://zguide.zeromq.org/docs/chapter2/
There is not (AFAIK) languages that make nice to properly wire this AND keep the semantic understanding/composability of it.
For example, is hard to know if a function if is running in a thread/process, what is their priority, etc.
I imagine that a language with orchestration (like Erlang/Elixir) in a explicit way will be far easier, but is hard to figure how do it with good composability
Most simplistically, ports are designated as "input" or "output" (from point of view of the node). But, these designations can be expanded so that ports can carry designations matching the ZeroMQ socket patterns. The implementation can then leverage the transport agnosticity of ZeroMQ when supporting adaptation to MT vs MP but with the cost that all edges require data serialization instead of pointer passing.
If you need to use Go, there is also ergo: https://github.com/ergo-services/ergo
Yet another cool tool is lunatic: https://github.com/lunatic-solutions/lunatic
You can build cyclic data-flows [2] - not just DAGs - which seems a key ingredient.
You can also write contract-based tests that verify the end-to-end behaviors of your "agents" [3].
Today, you're limited to SQLite or TypeScript. Looking forward, we aim to open it up to any program packaged in a docker container.
[1] https://docs.estuary.dev/concepts/derivations/
[2] https://docs.estuary.dev/concepts/derivations/#approving-tra...
IIRC, Akka based their implementation on Erlang (possibly OTP), but it's been a long time since I did anything with either Erlang or Scala, and I never used Akka in Java.
The same could be said for splitting a bunch of threads off with opemMP. It isn't about divergence, it is about data dependencies, data owership and synchronization, which are all the same in both scenarios.
When the average dev team finally gets the simple queue they asked ops for, the thing is probably late and only present in 2/3 of dev/qa/prod (or some similar SNAFU). Flink works with AWS EMR, so we can't explain everything by just considering whether hosted services are available, but admin/setup/general conceptual overhead is on a different level. Partly because they've been burned in the past with stuff like this, and partly because they are lazy.. devs mostly want simple infrastructure mirroring simple data-structures where they can test/develop discrete code locally without thinking about system-level stuff.
Consider the problem of a) developing a new agent in an actor system vs b) developing a new "step" in a simpler batch-oriented pipeline. For (a), to actually run my code experimentally locally I need to simulate the appropriate message(s) locally and maybe other aspects of system-runtime. Whereas for (b) one digs up appropriate input and starts hacking on a docker-container that one trusts can be jammed into an airflow DAG later. Both approaches need a bit of dev/test harness kinda setup, but (a) is potentially more involved and more importantly it usually crosses team-boundaries (requiring both devops and devs). This will probably create pain unless the org has a talented "platform team" already in place.
Another aspect involving team boundaries is that besides dev teams disliking ops/infra, they also don't trust each other! So after effort is sunk into things like the "smarter-queue" that might support messaging and actors making actors, it turns out every team wants their own queues, routing, codebases, underlying storage, etc. For better or worse, attempting to provide features along the lines of flexibility/interoperability are thus undermined. Out of necessity, teams often want to provide some limited access to data which other teams can consume. But they resist providing any kind of access to code APIs / runtime.
Typical scenario: Ever work at an org that's supposed to have a data-lake, but every single new app or new feature inside an existing app generates requests for new buckets that literally nothing in the existing system can access? Some manager or "senior" dev is having a knee-jerk reaction that they want to build a kingdom. In the end they won't actually enjoy answering access-requests to their walled-garden, but they think they can push those over to support requests. With only data as a deliverable, they don't need to think about any interop and, bonus, no one will even know if their messy scripts are in version control.