Data flow syntax: The missing piece for the success of functional programming?
bionics.it
bionics.it
While you can do that, Python doesn't make your life easier if you want to go functional, really. There's only map/reduce/filter, one-line lambdas (that get really messy really fast), no currying, etc. (however, comprehensions are fine)
I appreciate the effort of the OP in questioning the fact of dataflow implementations, and I suggest him to further explore the world of functional languages - it's huge and resourceful, while the quality of the content written is surprisingly high on average - as he may even find a working implementation for the problem he was trying to solve (happens way too often to me).
BTW, a quick google brought me to this [1] - I'm not sure if it's exactly what we are talking about, but seems related.
Thanks for the cells link, will have to check!
And a direct link to the SPL source code for the application: https://github.com/IBMStreams/streamsx.demo.logwatch/blob/ma...
http://doc.akka.io/docs/akka-stream-and-http-experimental/1....
There are split semantics where the tuples are simply fed into other flows. And then merge semantics where flows recombine into a single flow for further processing.
Even though they have specialized functions and constructs for branching out and merging, I'd argue that being limited to chaining functions, IMO it quickly grows unmaneagable when dealing with any more complex data flow networks. AFAIS, you'll get super-long chains of function calls, that are really hard to change, or inspect.
Compare that with the syntax proposed in the post, or some other FBP syntax like the one in GoFlow [1], where you define each connection between an outport and an inport on one line. IMO this gets exponentially more manageable, readable and maintainable.
If you want to deal with multiple streams of data, I think combinators do just fine. If you want to choose which function to pass into higher order functions at runtime, then indeed you can define functions and just reference them. If you want to split data up, you can use all of the nice function composition and pattern matching available in languages like scala and haskell.
To split a string, take half to upper and half to lower, you can use tuples, or zip and unzip, partition etc. Stuff like that. None of it seems messy but I'm probably not understanding this article correctly.
Nonetheless, the proposed solution is most in the component/object realm rather in the functional world.
As an evidence the use of nouns rather verbs as in :
splitter.input = hi_generator.his
lowercaser.input = splitter.left_part
uppercaser.input = splitter.right_part
This is not a judgement, just a remark. A functional solution would rather use transformation combinators to propagate a transformation on a specific value part.
For instance : input |> map split_two_halves
|> map_fst to_lower
|> map_snd to_upper
Or even better, use lenses as a generic way to unpack a specific part of a data structure and to repack an updated value: input |> map split_two_halves
|> map_lens fst to_lower
|> map_lens snd to_upper- http://docs.openstack.org/developer/taskflow/
- https://wiki.openstack.org/wiki/TaskFlow
We recently had a brown-bag (with a community member from HP and myself from yahoo) that is shareable and maybe useful to folks @ https://docs.google.com/presentation/d/1EZoY4FE2SDjfCqMCgBRr...
e.g.
applyStoreFilterStream.recv(function () {
// do some stuff
})
// when either removing a filter or adding a filter
.compose.serial(removeStoreFilterStream, applyTypeFilterStream)
.recv(removeAllElements)
.recv(pager.reset.bind(pager))
.recv(requestProductsStream.send);
This was using an in house streams library (before i'd found Rxjs).I did find that, because there is only one line at a time within which to express a dependency (from one stream to another) it was quite difficult to conceptualize the graph structure being created. I think perhaps having editors that can interpret this sort of code into a visible graph structure and conversely, compile a graph structure down into this form, would be interesting.
This article illustrates a problem I have with FP in general, coming from procedural / OOP codebases: For the generator example, how can I be sure that the initial generation of strings isn't getting executed twice? Will python analyse that I'm only ever reading from "left" or "right" within the gen_uc / gen_lc loops, thus optimizing away the the other side in the upstream generator? So I'm basically trusting the runtime to optimize for me? Do we have benchmark showing the overhead for this runtime analysis? Does it work consistently with JITed pythons like pypy?
So far I'm just always too skeptical that I could program myself into a performance hole that require complete rewrites to get out of, so I stick with code where I know what will happen on the machine. Python for me is just the toy language to try these concepts before diving into actually performant ones, but not finding any significant Haskell/F#/.. benchmarks that can perform as well as good C or Fortran code, hasn't helped in reducing my skepticism. IMO the thinking of "correctness now, performance later" is a fallacy in that the choices I'm making in the design phase can lead to complete rewrites in order to get something performant - in the worst case switching between languages.
This is in contrast with list comprehensions, which have a similar syntax, but which generate a full list immediately after being declared.
> yield (s[len(s)//2:], s[:len(s)//2])
I'm aware that the generator is being called lazily, but here we're actually calling each yield twice - once the left side is thrown away, once the right side downstream. Now, are you still sure that left and right is generated only once?
In this case, I think it'd make more sense to divide gen_splitted_his into two functions (or one which accepted a parameter), for generating one or the other.
Here's my conclusion: Not only is it not efficient, it doesn't even work. The generator cannot be rewound after it has been iterated over for lowercase - the uppercase version will never be executed. You essentially run into the problem described here: http://stackoverflow.com/questions/1271320/reseting-generato... .
It only works once you actually generate twice
> lc_his = gen_lc_his(gen_splitted_his(('hi for the %d:th time!' % i for i in xrange(10))))
> uc_his = gen_uc_his(gen_splitted_his(('hi for the %d:th time!' % i for i in xrange(10))))
It seems to me for this toy example, generators are simply not the right tool.
The solution being proposed at the end is actually not using generators at all.
Also, the abstraction allows for optimizations like delegating to Apache Spark or other processes/threads.
How is it not successful already?
http://www.tiobe.com/index.php/content/paperinfo/tpci/index....
I'm only pointing to Tiobe because it is popular, not because it is accurate, but you could use any other popular index and you would get similar results: a whole lot of imperative languages, but none of the officially Functional languages.