Flow-Based Programming
jpaulm.github.io
jpaulm.github.io
Instead of losing yourself in ever more complex syntax convolutions as has happened in a lot of functional programming, you make the components (even long running stateful ones) self contained with ports for data input and output as the sole means of communicating with them, over buffered channels to allow asynchronous computation, and most importantly keep the network definition separate.
Just this idea in itself is just brilliant. Hats off to Mr Morrison for that!
It allows to decouple complex software into reusable components, without clever FP syntax.
(Though, FP is perfect for implementing the components themselves. It just doesn't really scale all too well for whole program architecture, in my experience).
One point to note: The visual component of many fbp systems is completely optional, as is the idea of using novel DSLs. You can as well define your networks and components in pure code. See GoFlow https://github.com/trustmaster/goflow) and my own little experiment FlowBase (https://flowbase.org) for examples of that, in Go.
I successfully built a rather complex little app to convert from Semantic web RDF format to (semantic) mediawiki XML dump format, in two weeks straight, of linear development time: for each component (of ca 7), implement, test, go to the next component (See: https://github.com/rdfio/rdf2smw)
The same implementation in procedural PHP took months, and still doesn't have all bugs and strange behaviours filed out.
This concept sounds exactly like actor systems like Erlang/OTP and Akka, only with a different set of terminology.
The submitted site and your comment don't mention those anywhere.
Are there appreciable differences between actor systems and FBP?
The FBP model provides implicit backpressure, because of the bounded buffers on the channels. The actor model on the other hand is more loosely coupled and flexible.
This actually means FBP systems are slightly more optimal for efficient within-one-node, in-memory parallellisation, whereas actor systems shine more for distributed systems.
They shine on different scales, in other words.
I wrote a post many years ago on this, that seems to have aligned with the experience of a number of people, based on the reactions in the comments: https://rillabs.com/posts/flowbased-vs-erlang-message-passin...
[0]: https://doc.akka.io/docs/akka/current/stream/stream-graphs.h...
Note that this is why people recommend using GenServer's "call" instead of "cast" in Elixir/Erlang by default - it applies a sort of backpressure because call waits for the process to return a result rather than being fire and forget.
It's not exactly the same thing, or as configurable as Go channels, but it is an option.
> Programs in the original CSP were written as a parallel composition of a fixed number of sequential processes communicating with each other strictly through synchronous message-passing. In contrast to later versions of CSP, each process was assigned an explicit name, and the source or destination of a message was defined by specifying the name of the intended sending or receiving process
I wonder what gives? Why this return to an older form?
https://en.m.wikipedia.org/wiki/Unit_record_equipment
What you are programming is the graph of how these machines connect and flow data(keypunch cards) through each other, while the machines themselves and their function are abstracted out. Data does not stay "at rest", it's presumed to reach a terminating point where it exits the graph. A buildup of unprocessed data in a machine's inbox results in an overflow.
It's a very useful model for making a debuggable asynchronous system since it imposes static constraints everywhere that you can map to your real hardware resources, versus the emphasis on dynamism seen in the actor model(actors hold private data, modify their state, create new actors - all explicitly disallowed in FBP).
I have yet to see an FP concept for composition though, that is as simple and generic in its implementation, as the FBP principles.
I have sometimes thought FBP networks provide roughly the same function (pun not intended) as a monad, although I never seem to fully grasp what a monad is, so I can't tell for sure :o)
Also, FP algorithms developed for one evaluation context cannot easily be ported to another. They are implicitly entangled with assumptions. For example, I cannot evaluate a lazy algorithm in an eager evaluation context without paying a huge price. Adding back-pressure for list processing where a program was not designed for it would easily result in deadlock. Adding exceptions without adding unwind (or bracket, try/finally, etc.) to the existing program expressions will easily result in buggy code.
I think your assertion might be true for some specific integrations of FP nodes into FBP. Is this what you mean?
But it wouldn't be a big difference if the host language was heavily imperative, so long as it favors immutable data structures. The FP aspect for this exploration of DSLs is much more social and cultural than technical.
I think that if we really want to decouple AST from evaluation context, the main tool we'd need is to deconflate module 'import' into separate steps to 'load' a module's AST and 'integrate' definitions, allowing for intermediate processing. This requires some host language changes, elimination of module-global state, etc..
My guess is that you forgot to put the definition of the domain sans www in your webserver config.
If my guess is correct then this link should work:
Edit: doesn’t work either. Perhaps you don’t have a TLS cert for that domain? In which case maybe this link will work instead:
Edit 2: Without https it works. And like the person responding to this comment said it redirects to the GH repo.
So it's useful to design with FBP in mind but with a linear interface as the entry point.
Another aspect of this is that FBP graphs are static but you may have a need to reconfigure them frequently; that is, you may want to have a graph compilation step drawn from a source language, rather than manually wiring it up.
A way of making that graph compilation more than a syntax is to include a formal constraint solver: Excel, for example, flows the data after determining a solution for how cells relate to each other. The power of the spreadsheet paradigm really lies in these combinations of concepts.
Lastly, there isn't really magic in the algorithmic/implementation aspects of FBP. It grew out of 1960's mainframe types of problems, and so it can be implemented in a low level way with static pools of memory and pieces of assembly code. But it remains conceptually just as relevant to today's massive distributed systems.
For same reasons we start looking at FBP and functional reactive programing to simplify the design and maintenance of complex interactive UIs. End up implementing Kelp (https://kelp.app) with visual FBP editor and reactive framework (https://kefirjs.github.io/kefir/).
In Factorio, you build a larger and larger factory out of pre-established functional components (assemblers, labs, chemical plants, etc) that take in a limited set of inputs and produce (usually) a single output. Your challenge is not to define the functional core processes, but instead to wire together those functional components by connecting their inputs and outputs in ever-more-automated fashion, starting by hand, then using simple belts (pipes) that eventually allow arbitrary load-balancing via "splitters", and eventually through to trains (the forking and load balancing happening via backpressure in the train system) and robots (where everything is managed essentially as a single state database of requests, and backpressure is provided by output limitations, usually per functional component).
Naively, I think that someday a decent chunk of programming might actually look like this, and parts even be represented visually (though in my opinion likely still defined formally as text). Only I think programmers will continue to write the functional components themselves, unlike in Factorio. They'll just live on different levels of the "codebase", and the "pipes" level will likely be a lot more abstracted than it is in Factorio.
As a software developer, I find this paradigm to map very well to serverless architectures, because you generally want to think a level higher than the per-machine basis. It does require a willingness to forgo handy and well-established tools like the filesystem and Unix pipes in favor of higher level abstractions around transfer and storage of data.
Factorio and Minecraft automation mods are a big inspiration! Check out InfiniFactory too :)
Bridging existing applications to the Serverless paradigm is far from simple. That's one of the biggest struggles I've experienced trying to build a Flow-based software platform.
Learn more every day though. Thank you for the interesting comment!
It looks very professionally done, but reading screens isn't as easy as it used to be for me, a video is better.
Good luck!
One of the most compelling arguments for a data-flow/flow-based programming is the mental model and the visualization aspect of it. This opens up opportunities for monitoring, visual, data-driven programming and it is white board friendly.
In Elements of Clojure[0], the author discusses the concept of "principled components and adaptive systems". And a flow based design reminds me of exactly that. The semantics of composition and communication are well-defined and universal, but internally the components can (should) be specific and concrete.
Similar can be said about Small Talk as well. A primary aspect of its design was the mental model, understanding and learning. The core idea was that learners (especially children) understand things in terms of their operational semantics.
> I don't mean to derail the conversation, but this really does remind me of the game Factorio, though sort of in reverse.
So no, I don't think this is a derailment, but likely one of the most important aspects of paradigms like this.
from 2015, on the same article: https://news.ycombinator.com/item?id=8867584
Related:
2019 (a bit) https://news.ycombinator.com/item?id=20215592
2019 (another bit) https://news.ycombinator.com/item?id=19203642
2019 https://news.ycombinator.com/item?id=18859019
2015 (not that good) https://news.ycombinator.com/item?id=10755250
2015 (a bit better) https://news.ycombinator.com/item?id=9718868
2015 https://news.ycombinator.com/item?id=8992281
The central concept is a "collection", an append-only set of schematized documents, which can be captured and materialized into other systems (e.x. pub/sub, S3 buckets, etc). "Derivations" are collections defined in terms of source collections, and stateful transformations/joins/aggregations which are applied to them.
A key twist is that collections are simultaneously a batch dataset (backed by cloud-storage) and also a real-time stream. They unify the current dichotomy of "historical" vs "streaming" data into a single addressed concept. Declare a new derivation, and it automatically back-fills over history right from S3, then seamlessly transitions to live data.
If this sounds interesting, check out our docs [0]. We're early, but love feedback!
(By the way, I'm a big fan of visual programming, so I'm just curious.)
Then there was this startup way back when: https://www.cloudamqp.com/blog/2017-05-31-work-queues-in-rab... https://librearts.org/2014/10/artificial-intelligence-design...
On the front-end side, Flowhub uses NoFlo for state management, a bit similarly to how you’d use Redux: https://github.com/noflo/noflo-ui
I have my little individual programs, grep, awk, sed, jq, etc, and then I can endlessly mix and match those different components to do what I want.
The limitation that I have seen with Unix pipes isn't often with the ability to process or manage the data it is that it only works if all the data is setup the proper way. As I have been in the industry longer it seems to me that most code that gets written isn't about actual computing but just centered around importing and transforming data that is expressed in different ways.
Is there something that makes reading in disparate data records easier so I can focus on the computing part of computer program and less time on parsing data?
You could have a conduit connector to join a bunch of wires/signals, I suppose. But, realistically, that is a mess always.
A node editor makes it somewhat less messy.
Curious to see if it (or concepts from it) gets picked up for programming education one day.
It is easier to design digital circuits when you have a whole catalog of 7400 and 4000 series gates, than it is using individual transistors. It is easier to wire a house when you're not making wires and switches with a hammer and a forge.
I welcome this new higher level of abstraction, and am willing to pay the cost in terms of CPU and Memory to get there, just as I'm willing to waste transistors or copper wire to have something done and working.
The most substantive addition in recent years is Akaigoro's comment about Actors, added Jan. 2020 (@guitarvydas, care to jump in?!), preceeded, I think, by Joe Witt's reference to NiFi in mid-2018. This general lack of activity I feel might lead readers to assume that FBP is an outdated concept, whereas this thread proves it's definitely alive and kicking! I, personally, am not allowed to update the article, due to the WP ban on self-promotion, so I would like to encourage people to add topics, controversies, anything, to the WP article... Freshen it up a bit, as it were! Thanks, and stay safe, everyone!
Apache NiFi comes quite a lot closer, with the main difference that they only have a single in-port, instead of separate named ones. Also it seems to be among the more heavy and complex implementations (for good and bad).
It would seem like I need to suspend my computation then "ask" the result from some component, get the result back, and then alter my computation based on that result.
Pure flow-forward would not seem to support this easily. Or can it? Or does it come down to that we always will need BOTH sync- and async- functions? (unless of course we limit the problem domain)
It's a bit like "pure functional programming". We can't really do that in practice because we need IO. The best we can do is to divide the program into two parts, purely functional, and imperative. Similarly I assume we could (and should try to) divide programs into "data-flow-part" and "synchronous-part".
The simplest way is with when/unless vertices which take a boolean input and some other inputs, and which either output the other inputs if the boolean input is true/false, or output nothing.
I understand that there can be loops in the dataflow and feedback can alter the processing logic of a component at some point in the future - since processing is async, the alteration can only happen in the future.
But when a component needs to make a decision, not alter its processing logic but simply make a choice, can it somehow 'ask a question" from some other component which is part of the data-flow, before it it decide which branch of the data-flow to send data to?
My question is simply is it possible and/or practical to do general programming with only (asynchronous) dataflow between the components/functions/units-of-computation?
Or do we need both synchronous and asynchronous components/functions/agents etc. ?
In pure tagged dataflow systems, asking a question doesn't happen: programs are directed graphs and like most other programming systems vertices have to wait for values. However, at a higher level, you could have vertices sending messages to one another and receiving answers back.
So computation can not start before I have all the inputs, right? But when I --a vertice-- am doing a computation, I would like to ask for more data from some other component, so that I can complete my calculation. So how could I get such data except as one of my input? But I must have all my inputs already, else I would not be executing.
This would then seem to be at the core of the difficulty with data-flow programming. Once a computation of a node has been activated, it can not ask for help from anybody, because it must have all its inputs present before it starts doing its calculation. Does this make sense ???
b = a
.filter1()
.filter2()
.filter3()
.filter4();
if (b.foo()) {
return a.filter5();
} else {
return a.filter6();
}
For visual programming, AFAICT you have to feed `a` downstream twice: once to filter5 and once filter6 in preparation for the conditional branch. That requires one of the following:* intersecting lines, which are way more visually complicated to read than the code above (and can quickly create visual ambiguities which as a class do not exist in a text-based language)
* one line which loops under the bottom of one of the downstream objects, which is visually distracting (especially if you have more than two branches)
edit: clarification
A diagram could not (easily?) describe the fact that we get the 'b' from one component and then we send the message "foo" to that component 'b". The data-path from this component to "b" could hardly show in the diagram because "b" is the response to a previous request which could be one of many possible alternative responses.
I'm not sure if I can express or even think about this very clearly but questions like this make me doubt the generality of dataflow as an alternative to current-day programming practice.
There are historical reasons why we're programming in text in mostly imperative languages: graphics terminals were either not very good or very expensive, and computers had single CPUs, until quite recently.
I'm aware that text is more compact, and understand the various user interface issues (e.g. crossing lines), but there are advantages in representing programs as graphs as opposed to text. One of these is the possibility of very strong type checking while a function is being defined, and I'm currently working on that. Another is that it should be possible to take advantage of homoiconicity: programs are graphs, and graphs are the most general data structures. Also, graphs (i.e. data) can double as programs.
Tutorials for my language can be found here: http://www.fmjlang.co.uk/fmj/tutorials/TOC.html These are out of date, so suffer from the problem you hinted at above, and they also don't explain my aims. I'll write updates to them when the new type system is properly working.
I do understand that. I thought about implementing something like that for Pd, but in reverse. Programs there tend to look like write-only spaghetti, so it would be helpful to have something like a "spotlight" feature that temporarily removes (or blurs, etc.) the spaghetti that isn't part of a currently selected chain of objects. That would make spaghetti at least navigable with a mouse and some time.
Anyhow, I've never found a dataflow or visual programming that really deals with the ergonomics of writing the code. At some point I just want to do a little math, or a quick conditional, or some such conceptually simple business for which the graph just gets in the way. At that point it's easiest to drop down to some kind of text-based DSL inside a single box and be done with it.
F,v,x,m,b,k
(F - b v - k x) / m
For Lisp, I have to define a function using defun and make it available to the system.I'll add some new tutorials and a FAQ once I've got the new type system working, which is taking me longer than I expected.
Here's a simple case structure for a decision based upon a boolean value: https://imgur.com/a/YyzXVhU
This is really no different than the type of dataflow you have in other functional languages like F# or Racket. Values flow into special syntax forms such as if or match that allow branching.
The problem is everyone wants to write little nodes and "just" wire them up. But all the real complexity is not in the nodes, but the nature of the wires and their composition.
We'll get there though, stay tuned.
https://en.wikipedia.org/wiki/Tacit_programming
All programs are actually pipelines of data flowing from a low entropy state to high entropy state with IO and state as the endpoints of the pipes-. Using the point free style or flow based programming makes this entropy and the pipelined nature of all programs more explicit.
Using OOP or regular programming the pipeline nature of data flowing to and from state and IO becomes less evident and more convoluted.
This allows me to build ever greater programs, that does very complicated things, and has thousands of lines of code and logic, but everything gets boiled down to small bite-sized pieces.
It’s a more data oriented approach. And it seems to follow an assembly line model.
I liken it to building code pyramids. Where I keep stacking and chaining one function to another, to build even bigger pyramids.
At the end, all the code is heavily unit tested, and I have a high degree of trust in the fidelity of the codebase.
The other thing I noticed when I started using this style is that I was able to ditch unit testing mocking frameworks. This made me realize, when the Function interface was introduce it was a backdoor way to get folks back to coding to an interface, which we should have been doing all along.