Apache NiFi
nifi.apache.org
nifi.apache.org
You'll note that a lot of the discussion here is, "What is it? Is it like X? Is it good for X?" Those are great questions to answer on the home page.
Apache NiFi is a dataflow system based on the concepts of flow-based programming. It supports powerful and scalable directed graphs of data routing, transformation, and system mediation logic. NiFi has a web-based user interface for design, control, feedback, and monitoring of dataflows. It is highly configurable along several dimensions of quality of service, such as loss-tolerant versus guaranteed delivery, low latency versus high throughput, and priority-based queuing. NiFi provides fine-grained data provenance for all data received, forked, joined cloned, modified, sent, and ultimately dropped upon reaching its configured end-state.
... and of course:
Really? I feel like it's more: "usable, powerful and reliable: choose one ... if you're lucky".
http://accelrys.com/products/collaborative-science/biovia-pi...
We already had Knime, but that is Java-based and I believe it only has a desktop interface ( http://tech.knime.org/getting-started )
https://cwiki.apache.org/confluence/display/NIFI/Persistent+...
That's what I did for a long time. HL7 contains information such as "Admissions, Discharges, and Transfers (ADT)" which is sent from the Hospital Information System (HIS) to other department's systems (radiology, pharmacy, medical records, and possibly dozens more) and vice versa (LAB results back to HIS, for example).
HL7 interface engines unpack the HL7 data into separate segments, fields, and subfields, identify the type/subtype of message, route to various destinations based on the type and any other fields, map the data and reformat as necessary for the destination, encode back into HL7, and send it. It also needs to queue the messages as needed, re-transmit or set aside the message, or just stop sending depending on what it gets back in the form of response codes. So you can see in the diagram some of those steps. The also handle input/output either in the form of TCP/IP ports, reading/writing files (for batch processing), or pretty much any other method you can send and receive data (serial ports, in some old cases).
I'm kind of curious how (and how easy it is) to define which fields in the input for a given event goes to the output (and do certain transformations on those fields if needed).
Something like this would make that much, much easier. Cloverleaf would route to this Apache product, which would then route to the appropriate service. The data on traffic, etc., would be especially useful (i.e., the service that sends faxes is slow, etc.)
Each processor in the flow receives the full flow file, so each step along the way gets the "state" in full. This helps in case something interrupts the flow - you can resume an hour or a day later.
Does the NiFi designer have the ability to do something like this, the ability to create abstractions in the front end?
At my last startup, we did user tests every week, and my cofounder was very careful never to let on which test items were made by us. We didn't have our name on the door or on the buzzer. Nothing with our logo was visible between the entryway and the user interview room. And when he started showing products, he'd always start with somebody else's.
It was great. We got some incredibly honest feedback (sometimes brutally so) on what we were building. It helped us kill a lot of bad directions early. Most people just don't want to tell you that your baby is ugly, but they'll happily dish if they think it's somebody else's.
For example, you could set up a fake market research org, and have them say they are "conducting a study on [your market] and are looking for users of products like [competitor 1], [competitor 2], and [your product]". Then for the user testing sessions you could rent a conference room from somebody like Regus, or even rent a user testing lab. Then during the tests, make sure to start with general questions and test (or show screens from) the other products before you get to testing yours.
I'm not sure if those specifics will work in your environment, but I hope you see what I mean.
However, the functionality of the engine itself is top notch. We've had some Mirth servers running for a year at this point without incident. It handles alot of things around the nuances of HL7 nonchalantly, which has usually been my hang-up with non-healthcare specific toolsets like Mule or Camel.
That being said, I'll definitely be checking this out.
http://www.zdnet.com/article/nsa-partners-with-apache-to-rel...
Thanks, NSA!
Furthermore, as a 501(c)(3), the ASF is limited in what it can do with do with donations and very, very rarely accepts targeted donations aimed at a specific project -- it's just not worth the hassle or risk. So while outside entities might contribute by sponsoring people to work on the project, no project at the ASF has someone "behind" it in the sense of direct funding.
This is all part of maintaining project independence.
Hope this helps to explain why it is not always easy to discover "who is behind" an Apache project. :)
More info: https://www.apache.org/foundation/how-it-works.html#hats
But it is usual to be able to deduce at least something about contributors from email address, github accounts, other contact info. But in this case there was absolutely nothing. It was almost as if they were experts in covering their tracks! I just thought it was quite funny.
http://www.forbes.com/sites/adrianbridgwater/2015/07/21/nsa-...
Hopefully it's better than what it was in Jan/Feb timeframe.
- We wanted to collect files from various locations and push that in hdfs. Nifi seemed like a good way to build a self service setup. Once we setup source and sink, if we read + processed + removed file from destination, Nifi copied it again. We did not have control of always removing source files as soon as it was copied on destination. Components we used were GetFile, PutFile and some options in Conflict Resolution settings for these components
- Inspect file name, run a script to generate new subdirectories for partitions and place the file in appropriate partitions. Attaching a script was easy. Changing the destination path on the fly was not.
Some Complex cases - there are other ways to do this but setting this up in Nifi was a breeze
- Set up file collection from ftp, sftp, file copy from 20+ locations. This was painless, few minutes per source
- Add REST interaction within data flow easily
- Read CSV files and convert to Avro/Sequence files
- Read files and route part of data to different processors
We also ran into some strange bugs where Nifi got stuck in some type of loop and kept copying data over and over again.
We were able to do all this testing in 2 weeks. Give it a shot, it might work out for your use case.
The guys developing it are incredibly responsive. I submitted a bug one morning via the mail list at around 9 am, and by 11am they had a patch slated for their next release.
I'm not complaining about not understanding, I just had one of those moments of "wow, the world of software is really big and there are vast, heavily funded, corners of it that I've never even heard of".
If you never had to process that kind of data (we're talking about Google, Facebook and Twitter kind of data) then you probably have no exposure to this kind of system.
One thing that is discussed is how NiFi, in contrast to "proper FBP", has only one inport, to which all incoming connections connect, so incoming information packets are merged, and subsequently need to be "routed".
[1] https://groups.google.com/forum/#!searchin/flow-based-progra...
Years later, I had to maintain a software application built by a bunch of idiots who thought it would be a great idea to use LabVIEW as their GUI layer (because apparently that was easier than learning/using a proper GUI toolkit). This monstrosity communicated with a back-end running on Solaris via LabVIEW's freaking horrendous TCP control. The whole thing made absolutely no sense. Though LabVIEW did provide a rather nifty visualization of spaghetti code.
I suspect this was a resume-padding exercise for the original authors. Rumor had it that one of the engineers responsible for the decision to use LabVIEW had been hired away by National Instruments. And we all cursed him.
So anyway, after that experience I'm pretty well prejudiced against graphical 'programming'. I took one look at this NiFi thing and said 'Ha! Nope.'
Does it have the ability to integrate with Hadoop ?
<disclaimer: I'm an employee of one such vendor>
Wikipedia says Dataflow programming languages share some features of functional languages, and were generally developed in order to bring some functional concepts to a language more suitable for numeric processing.
The thing closest to Dataflow that is most commonly used is the concept of DAG operations in Spark, but Dataflow usually makes time windowing a first level concept. Spark Streaming is moving towards this type thing.
It is true that there is overlap with ETL tools, but that undersells what Dataflow is.
Dataflow is a programming model for performing actions on data. A system that implements dataflow programming will probably have functions to load external data into the structure needed for the system, but it isn't primarily about moving data to another system.
For example, Google Dataflow[1] has functions for reading files etc, but there aren't really the huge number of things for cleaning and processing data that a real ETL system has. Instead, you load the data into the system, and then process it for a specific task.
[1] https://cloud.google.com/dataflow/what-is-google-cloud-dataf...
Still investigating.
Why visual programming tools don't (at least in documentation) address version control is an interesting question, given how important that is in programming generally, and the fact that those tools are often focused on enterprise audiences.
So there's a distinction there between the graph, and layout of the graph ie x,y coordinates of nodes, and routing of edges. What would version control track? This is a non trivial problem imo. Edit; specifically, storing a graph in some canonical form so that trivial changes would not create massive 'false positives' in the diff. For example changing an edge indicatin an entire sub graph was added or removed. The problem is similar to xml tree diffing.
This is why I think versioning in these tools is pretty much non existent, imo of course.
I haven't tried this, but this looks f'n awesome for what we do. We've been using LinkedIn's Azkaban or Confluent as two seperate paths waiting for cloudera's dataflow.
https://www.youtube.com/watch?v=jctMMHTdTQI seems a little slow but could get your started.
After a few minutes fiddling it seems you make "processors" which can do things like file input/output and data manipulation, chain them by dragging a connector; outputs can be things like SFTP. Closest thing I've used to this would be YahooPipes but this is self-hosted and so can access local files and such - seems powerful but I [as a non-coder] wonder if it wouldn't just be quicker to knock up a bit of python script for the sorts of things I imagine doing with this; that's probably a lot to do with limitations in my knowledge (of both).
Some good demos/examples one can mess with would be great.
I have been playing with both, and I prefer Spring XD's DSL. It looks a lot like Unix pipes, which I could easily grok (and it has the full suite of Spring monitoring and logging tools built in). That being said, they are both excellent projects.