The New Data Engineering Ecosystem: Trends and Rising Stars
insightdataengineering.com
insightdataengineering.com
Migrating and integrating data is incredibly expensive. When you implement a new system you either have to migrate the data from the older system (= expensive) OR having both systems working in parallel one for "older" data and one for newer (= really expensive)
Currently the established risk-averse corporations, the owners of most of the non- social media and internet data, seem to be still testing enterprise Hadoop distributions like Cloudera and Hortonworks for daily enterprise data operations, and maybie have some projects on R&D for harnessing new "types" ok data (like sensor data from a factory floor or high granularity transport data for supply chain).
Still, I hope the best of the new tech can get a place inside the modern corporations that work on important problems like Energy and Healthcare.
But, I may be completely missing the best practices though.
[1] https://wiki-bsse.ethz.ch/display/JHDF5/JHDF5+%28HDF5+for+Ja... [2] http://www.hdfgroup.org/products/java/hdf-object/##DOWNLOAD
Log / Stream Processing - Kafka \n Scalable Storage - HDFS \n Data Processing - Spark, Map reduce (in that order) \n Historical Analytics - Hive/ Spark SQL \n Real Time Processing - Spark Streaming, Storm \n NoSQL - Cassandra/ HBase \n NoSQL (In memory) - Redis \n Search - Elastic Search \n
Some more honorable mentions: kibana on elastic search - for analytics visualization \n druid - for analytics \n
Above are the basics - if you add them you will have 90% of the standard stack for big data.