For ‘Big Data’ Scientists, Hurdle to Insights Is ‘Janitor Work’
nytimes.com
nytimes.com
1. Data hygiene issues show up inconsistently, even within the same data source: I experienced this first-hand as a quant trader: I only had half a dozen, well-structured time series data coming from our trading apps and switches. You would think that I can get reasonably clean data every day, or failing that, being able to automate the data cleaning script. Nope! I experienced every data hygiene issue imaginable: from ntpd being broken, local time/UTC inconsistencies, switch firmware acting up, all the way to human errors in the ETL process. Building a reliable data pipeline has so many moving parts that I am fairly pessimistic about any one-stop solution.
2. Then there is a volume issue. If the data is small enough, humans can correct them fairly reliably, especially with some help from automation scripts/software. That said, doing this at scale is a very hard problem. One thing I learned working at a big data company is that many folks use MapReduce as a data cleanser at scale, and for that, MapReduce is a pretty awkward tool.
3. Anecdotal evidence: I have talked to employees at various big data platform/software companies, and though they have a wide ranging opinions on stream processing, Hadoop, Spark, etc., they all agree data cleaning is a huge pain/deal-slowing/-killing nemesis/unsolved problem that their employers semi-solve with increasing Sales Engineer headcounts. If this issue of solving data cleaning at scale was easy, I feel that someone would have come up with a very effective answer by now (and as an industry insider, I should have heard about it).
For example, the common denominator data format for many use cases (not all) is a relational database. Why isn't there something that can grab data from anywhere in any common format and import it into a relational database with only a couple commands or mouse clicks? Right now the situation is that you either have to pay a lot of money for such a tool which usually doesn't work very well or you piece together the various data conversion tools with scripts. There is of course a lot of other steps in the ETL process, but that alone would save a lot of time.
I agree with you here as well. I recently wrote a blog article about this, at least for log data > http://www.fluentd.org/blog/unified-logging-layer
IME, the issue is not tooling so much as dirty data. Even if a tool is a perfect solution, the data is (in the majority of companies) so dirty and inconsistently sparse that automated tooling breaks down.
Agreed. Even though all these companies in the 'big data' space claim they do everything under the sun because they run on hadoop or spark, data cleansing is still an area people are hesitant to dive into. Rightly so. What surprises me more is that ETL is also widely over-marketed by a lot of new big data companies. I don't think people have paid enough respect to how subtle the 'transform' stage can really be.
We have found that determining the best value of truth from incomplete and sometimes conflicting data sources is one of these messy problems with no easy answers. It takes a lot of fiddling with match rules and considerable domain knowledge to get good results. Accordingly, our product is designed to be very flexible and configurable.
I suspect improvements will come from making it easier for analysts to see a) the effects of proposed match rule changes and b) what decisions contributed to producing each output record under the current rules.
After twenty years of financial operations, including a decade in the back-office of a top hedge fund, I eventually accepted that I'm a bit backwards for my desire not to jump straight to a pivot table when encountering a new data set. Like a farmer reaching down to touch the soil, my first step is exploring the rows and columns with little tests here and there to find weaknesses within the information at hand rather than paving over them with instantly flawed reporting.
Anyway, for all the big data scientists out there too sexy to clean up their own data, I'm looking for work and don't mind pushing a broom. Check my profile for more info.
This needs to be done by more people, more often. I've worked with more than one (often SQL, but it doesn't matter¹) database whereby someone will claim that X is always true. It X is something that can be put into a DB query, then you should do that, and run the query. Every time. In my experience, if it isn't enforced by constraints in the database software, then it isn't true.
¹The nice thing about SQL is that you can but constraints on the data, and use the database to keep various invariants invariant.
Often, you'll find people making assumptions about the data based on the business logic, not the constraints in the database. Such thinking is flawed: you cannot reason about what you think the data is, you must reason from what the data actually is.
Worse, you'll find people making assumptions about what they think the business logic is.
I can't comment too much on it publicly but well, it looks like it was worth it. I definitely agree that there is a split between statisticians, and "data engineers" in both personality and interests, and you should have both on your "data science"/"big data"/whatever they call analysts these days team. It is however bloody hard to convince upstairs to pay a lot of money for the work! (as with anything "without deliverables", really)
Oh, but their data warehouse thing that nobody knew what it was called, inserted a random space every so often and removed random spaces every so often making the text completely unreliable. Compiled with all the normal misspellings, fat fingered acronyms, slang and other nonstandard ways of writing things it turned out to be a huge job.
"Bicycle stolen at 0930 from residence on the 300 block of E 10th St, East Village"
would get turned into "Bicycle stol en a t093 0 from r es iden ceon the 300 block of E 10 thst, Ea st Vil lage."
Good luck with that.
Maybe google "machine translation segmentation" or OCR for references to papers on the topic, but methods for doing this are really successful (especially on english).
Doing it better is simple: If you do ETL, version control your code. If you do hand edits, track the before and after.
Founder of an ETL startup here. This is exactly what we believe: the end-user of the data should be involved as early in the data pipeline as possible, including the wrangling. If you eliminate the engineer from the ETL process you remove a lot of painful back-and-forth and get more flexible pipelines.
But in many cases it's a technical debt issue. In the early startup stages when you're struggling for product-market fit, cleaning or auditing your data is not a good use of time. When a business like that starts to get traction, it finds itself with clients that now want more reliability, scaling issues, and a bunch of data that no-one really remembers where it all came from or how it was generated. A lot of the big-data startups are hitting that stage now.
Each file had about 100k objects, and there would be less 10 objects with errors on avg if they were mal formatted. It was easy (but annoying) to do at first but scaling up the clusters made it exponentially harder to do by hand in the terminal.
It was conceptually easy to solve once I abstracted the problem a bit: I had a rough idea of where I was getting the data from on the mining servers (urls, ocr'd pdfs , etc), and since there are a bunch of libraries out there for parsing json that give an idea of where the errors are occurring in the files by the byte, after that it was a combination of traversing file forwards and backwards (in memory, luckily files are only ~20mb each, but crashed every text editor I tried before trying to do search and replace) looking for any data within the file that could help me reconstruct the mal formatted object or remove the object if not enough information was available (which if I didn't want to do if my error occurrence wasn't ~10/100,000, I could use the other objects near it to reconstruct the object from inference).
From Dirt to Shovels, Fully Automatic Tool Generation from Ad Hoc Data: http://www.padsproj.org/papers/popl08.pdf
The website for the project is: http://www.padsproj.org/
Trying to outsource this stuff is appealing, I suppose, but that has turned out in my life to surprisingly difficult as well. I once worked with a supply planning division of a manufacturing organization to try to reduce inventory by anticipating when orders would come in, so I got access to the sales database. It appeared that there was a large spike in orders later in the quarter, and that we were carrying inventory unnecessarily. Actually, as we discovered during a pilot, sales people who got new orders from the same customer later in the quarter would just delete and re-enter an entire order rather than doing small updates, which they found to be a hassle (this was quite a while ago, when these systems weren't as easy to use).
These attempts to section off and outsource the "low value" work often just make things worse. It's just too unpredictable.
On a side note, it's narcissistic and cocky to use the word "janitor" for such an issue. It's not like data scientists should only be worthy of doing the illustrious part of the job and never the dirty tasks, right? I'd still go and take the broom and clean up the mess when no one else can.
There is an enormous need in several industries for this sort of thing; but most people don't really know they need it yet.
Also, I observe that the data science community does not talk about these issues nearly as much as they should either because 1) it doesn't make a very inspirational topic 2) they are shielded from this kind of issue by a data engineering team.
For instance, statistical programming languages like R, which generally operate on nice tabular data, should come with built in methods to enforce data validity. I should be able to tell R that I expect certain columns to only contain values from 0-1 and no NAs. This is an easy example because languages like R are somewhat dictatorial about how they want you input and process data, but one can imagine the same sort of methods built into the base libraries of general purpose languages.
On the other hand, maybe we need to be more rigorous on the whole about data validation. We accept that automated/continuous testing is an effective mechanism for preventing bugs. We need the same automated systems to check data files and flag them when there are issues.
It is often misused to mean anything that doesn't fit on an A4 sheet of paper, but for many companies the largest dataset that they have fits in the RAM of an ancient laptop.
If you need to analyze a couple million sales transactions or a dozen gigabytes of web visitor logs then the most appropriate methods won't include 'big data' in their descriptions.
Semper chi^2, bro.
However categorizing data into correct groups, removing what you consider is "non-essential", or simply rounding off decimal numbers all can have impact on the analysis down-stream.
In some ways data science is pretty much all about learning best practices for quantizing data.