Too many tools not enough carpenters
ckmadvisors.com
ckmadvisors.com
Data science is a pipe dream for many companies currently. Having up-to-date sales figures that can be sliced by dimensions as simple as product hierarchy, customer, and region is a realistic goal for many of our clients.
I am not trying to discount data science. We have a data science practice that does some cool things, but it's so much smaller than the BI practice, because there's currently so much more opportunity for doing basic ETL correctly and simple reporting and dashboarding.
There are multiple clients who were looking for "real-time access to data", which was actually synonymous with "automatic weekly refresh would make our lives so much easier. Right now we have to wait on $overworked_analyst to get this out monthly, and depending on workload that can lag by a week or two after the end of the month." These are not exact quotes, but entirely accurate paraphrases.
This is the reality for so many organizations that I think sometimes falls out of context in an HN-like environment.
Something that drives me to really ask what the requirements are is building something for my Dad. It was a tool to do deconvolution of MS spectra, and he said it needed to run "quickly". I got versions down to an hour, then 15 minutes and bottomed out at about 5 minutes for a decent result. After a while I talked to him about the timings and he said that "quickly" meant "under a day". A failure on my part to clarify what is a really fuzzy term. Linked to that, at the time they had a process which involved someone frequently manually looking up an item in a ~3-5k list.
I'm a data scientist, and I think that the amazing tools available now can really cloud the problems that are faced by many organisations. Sure, we can use word-sense vectors to create a deep neural net to do a thing, but 4% of your data has a country of "NONE". Or you've got dates that don't make sense (1000 years into the future), or a suspicious amount at 1/1/1970, 1/1/1900 and 1/1/1904. I've seen important things with "ZZ TEST DO NOT USE" as a field, there's truncated data, broken encodings and more.
This isn't to make fun of the people with these errors, getting and keeping good data is hard and often overlooked. But unless you've got that, you're probably not going to be able to get anything useful from the fancy algorithms. And even then, there's a huge amount to be gained from improving simple interactions at the point humans and computers interface.
A few responses:
I hope that list was sorted.
4% of Country = "NONE" sounds like they're running a tight ship compared to some places I've seen.
At least those date fields don't have "Unknown" as a value. Yes, I have seen the date field stored as text.
You are absolutely right. The ETL side, nailing down appropriate business logic, chasing down source errors and inconsistencies, altering processes to capture the data we need to give good answers; these are the challenges that take >80% of our time. If we do our job right in these pieces, the reporting can be done by an intern who learned how to use a pivot table yesterday.
Ish. If I remember rightly I think they were trying to compare multiple fields on those elements but it could be narrowed down.
> Yes, I have seen the date field stored as text.
Oh yes, that is always fun. Particularly when you see both 23/05/99 and 05/23/99 in the same column. Something I've ended up building bits and pieces of is work to try and find these kinds of inconsistencies. I'm slowly trying to automate a lot of the initial checks on a new dataset:
* Does it have a consistent number of columns in the CSV file? * Does it have fields with a surprising amount of question marks in? * Are the dates parseable with a single format? If you need to be precise, how many can be parsed by only one of the formats that's seen in the whole dataset? * How many things are blank? * How many things are blank-ish? NONE, FALSE, Empty, N/A, etc. * What does the encoding look like, are there any particularly weird characters? * What control characters can you see? (after hitting an enormous XML file which failed to parse half way though) * What does the type look like for each column? Currency (and then proportions), date, etc. * Are there number separators? Are they consistent? (1,000.00 vs 1.000,00)
Basically, what will trip me up later and leave me scratching my head before having to add yet another bit of code to ignore a field?
One of my side projects at the moment is to pull this stuff together from rag-tag bits of scripts and split up code to something I can just throw files at and get an initial report.
Also, something very important in this is how things overlap. 5% of each column being empty/broken might mean you have 7% of your data with almost no information or 90% of your data missing at least one thing. Depending on what you want to do, either might be OK or terrible.
> If we do our job right in these pieces, the reporting can be done by an intern who learned how to use a pivot table yesterday.
Yes, exactly! The best end point is where new questions and updates can be done either by someone else like this or by the experts who really know their field.
Ahh, thank you, I needed a bit of a data rant :)
I am skeptical about the possibility of throwing bodies at the problem because what you find is that everybody is excited about getting things done yesterday and not about getting things done correctly, consistently, or repeatably and without that mo' people, mo' data and mo' money just means mo' problems.
I'm not sure how my company compares to the typical retail customers BI companies like to target but hopefully this is useful.
Rather than real time access I think Automation would be be the most desirable outcome. The Overworked Analyst you refer to above is definitely something I relate to in our company. If we can free up analysts time from all the mundane data extraction tasks they can add value by actually 'analysing' - not 'extracting'. The same applies whether this is financial data (accountants) or production data(engineers) being analysed.
The problem is an industrial plant is not a clean homogeneous operation like a retail store. Equipment breaks down, new equipment gets installed parts of the production line can be decommissioned routed around etc. Our business rules are dynamic constantly changing. A lot of the off the shelf products BI guys pitch to us almost inevitably be out of date the moment it was commissioned. I've not really seen any solution that wouldn't require constant tinkering to keep it up to date and often that requires domain knowledge. Engineers often don't have the software skills and software people don't have the scientific/engineering background.
We've had some absolute horror stories with out-sourcing work to contractors without domain knowledge. Even simple stuff can cause complications. One amusing example I can think of is getting back a program full of bugs because the contractor confused the name of a database field "Fe2O3" (Iron Oxide, Fe two 'oh' three) with Fe203 (Fe - Two hundred and Three).
If you'll forgive me starting with a quote-snip, I think you very much underestimate the complexities of a retail operation. Let me just say that we've had multiple day discussions just about dates with one of our retail clients, specifically how they should be compared year-over-year. In manufacturing, I'm sure you can also understand some of the challenges that exist around inventory.
Moving beyond that to the meat of your post, I am familiar with manufacturing. That's one of the major verticals we cover (several of our biggest clients are names you would recognize), and I came up as an analyst in a manufacturing organization.
We are a true consultancy. We don't sell products. We're a Microsoft partner, so by the time we're involved the product decision is pretty much made. Our tools are SQL Server, SSAS, and a collection of tools in the presentation layer. None of these are off-the-shelf BI products, but tools to build a BI infrastructure. Our selling point is being able to transform the knowledge and subject matter expertise of people like you into robust models that are suitable for data exploration, ad-hoc reporting, and structured reporting. Our work is bespoke, if you will.
You're definitely right that BI is a moving target. If anyone tries to sell you on a build it and leave it solution, they're full of it. Like you said business rules change. A lot of the time (in a well-designed BI solution), these can be captured as data and incorporated into the architecture seamlessly. Sometimes business rule changes do require re-architecting portions of the solution. One of our more recent projects where I was tangentially involved was implementing work-in-progress reporting for one of our manufacturing clients. I can tell you that the equipment turnover and production line rerouting problems are all well-covered. They can spin up and down production lines, factories, reroute product between steps and that's all covered in the existing architecture.
Even in less dynamic environments, it's hard to call a BI solution done. If the processes and data don't change, that's just an opportunity for the sophistication of the questions and analysis to progress. I don't say this from a position of greed and job security (though the security is nice, for sure (: ), but from one of honest observation. Once you know what has happened and is happening, you start asking why and what will happen, and once you can answer those, you ask how to change what will happen. This latter portion is where the data science of the parent article comes in.
I think your company's 'bespoke' approach is the best method. Ideally I think you'd get the best value by embedding 'experts' in the organisation not sure how possible that is.
Even across projects, we generally tend to keep consultants within a major vertical, so we do develop some subject matter understanding (I would hesitate to call us experts in most subject matter), which certainly aids in ramping up on new projects quickly.
Beyond manufacturing, I personally have had a number of engagements with healthcare clients. You should never ask me for medical advice, but I'll most likely understand a problem statement regarding medical data faster than you will.
So, we don't necessarily have in house experts in our clients' businesses, but we do have people who understand the common problems in a vertical.
Retail's challenge most of the time is the lack of resources on those situations when the business is a low margin one or they just don't want to invest. Dealing with unreasonable and incompetent external customers is frustrating too. Receiving incompatible POS data from each customer, realigning sales organizations or cleaning products hierarchy, and even, getting the numbers from sales and finance to match is a pain. But the amount of processes is about one or two orders of magnitude smaller. In addition, the amount of data collected is huge (Plant Historians alone collect 200k to a 1M values every second), and your KPIs can be as diverse as the many internal and external regulatory agencies that oversee your processes.
Consider just a few of the systems in addition to PI Historian: ERP (you know this one), Manufacturing Execution System, Laboratory Information Management System, Laboratory Electronic Notebook, Change Control System, Non-Conformance/Corrective Action System, Process Control System, Building Management System, Preventive Maintenance System, Learning Management System, Raw Materials Information System, Regulatory Submission Information System....and those are only the big ones.
When the plant is in a highly regulated business like pharma, the amount of bureaucracy needed to change a single parameter in a process control system is disheartening. FDA specifically, is mandating pharmaceutical industry to perform process monitoring to prove that the processes remain stable. In these industries we do multivariate Statistical Process Control, not because of the big data trend, but because we have been able to save multi-million dollars batches by observing weak multivariate signals that would pass unnoticed in a typical, retail oriented OLAP cube.
Tools can help. People are essential, whether it be getting consultants, hiring in-house, or training to acquire expertise.
I often joke that my job is simply the application of arithmetic under a handful of logical rules, for varying values of "logical".
Edit: I don't type so well.
They tied the sales/finance systems to their manufacturing division via some ad hoc lotus notes apps!
Most of my clients are "afloat" because they have a viable product..."big data" is something far down the road to them...
Data science is an interesting market for three main reasons:
1.) The client owns the data.
2.) The skills needed to effectively manipulate the data are very tied up with domain knowledge of the data itself. This tends to throw people accustomed to the rest of the software industry for a loop: if you're a frontend engineer, your effectiveness is much more based on your knowledge of JS/iOS/Android than on your knowledge of the clients' problem. But when I watched the best data-scientists at Google work (and did some unstructured data-mining and machine-learning myself), I found that the majority of your time is actually spent looking at data - it's pulling examples, collecting golden sets, determining outliers, graphing data, etc. And it's often non-transferrable: people who were very effective with the News corpus were often totally ineffective with Social, people familiar with the Web corpus might be ineffective with News, etc.
3.) You often need an interdisciplinary grab bag of tools to get useful insights. The most effective data scientists at Google were the ones who knew some basic HTML and JS or Flash charting libraries, because they could really quickly graph their data and send the graphs around for comments. Deep knowledge of machine-learning is usually less effective than pragmatic knowledge of machine-learning combined with a basic knowledge of stats combined with an intuitive sense of the users who're generating all this data combined with basic presentation skills.
This makes it very difficult to develop effective all-in-one tools that are broadly applicable. Unfortunately, many VCs follow the hype machine and aren't interested in consulting businesses, which means that a lot of money has chased the last wave and not gone into efficiently solving problems. There's probably a tradeable business opportunity here, but I have little interest in going into consulting, so unfortunately I'm not the one that can take advantage of it...
If you were, how might you trade in on it?
Or just as applicable to cybersecurity, where hiring is a total disaster and not even well known/liked multi-billion dollar companies can't find enough good people.
Everyone agrees that you want good people, but you just can't find them because there are so few of them.
The ones who are really feeling the pain are hiring more junior people and trying to train them on the job, but you don't manage to successfully train them all, so you either need to fire them or find them something else to do, or you buy tools that let them leverage the skills they do have to provide something useful, and maybe they get better over time.
Reminds me of when I was an application security consultant and I thought Web Application Firewalls were really dumb since I knew several generic ways to bypass all the firewalls on the market, but 7 years later WAFs have gotten better, and consultants still haven't gotten any cheaper.
So while I'm sure there is immaturity in the big data tools space, there is probably an 80/20 solution there that lifts much of the load off your skilled data science professionals.
I see too much of this. Blog posts that attack a business's decisions, without demonstrating adequate insight into the tradeoffs involved in making that decision.
Skilled people are rare and valuable and hard to keep, even if you do everything right. Obviously the most effective choice in any situation is to use and develop human capital. But if and when that person leaves your company, your investment just went up in smoke.
Fungibility of skilled labor is something we should all learn to love. We want it to be easier to jump around and not harder. We want more options and not less.
That means you'd like to get results from that person earlier than he/she leaves. Developing a person takes a measurable time and efforts; can you estimate how badly you need those results from data science? Maybe those anticipated benefits aren't that big?
That's not to say you can always incentivize the person to stay longer. Of course if you need the results.
You don't invest in capital just to solve problems. You hire a specialist / contractor / consultant. You invest in capital if you expect the problem to recur often enough to make the investment worthwhile.
Derek Sivers wrote about how he had everyone in his company answer customer phone calls. There were phones everywhere, and there were incentives for employees to pick them up. It took a lot of work, but it paid off.
i hesitate to use the words "big data" or "data science". really, unless you're talking about terabytes+ and/or sophisticated stats, it should just be called " analysis".
The picture of success in manufacturing orgs is having domain experts (manufacturing/operations/quality folks) routinely performing tasks with data that are NOT sporadic one-offs (for PowerPoint slides). The actual tools don't matter too much but rather the fact that people are proficient at using them. It means, for instance, that you should see non-software-engineers able to work with databases and reporting tools directly (eg sql-server management/report studio), comfortable with basic scripting, and skilled in some programmable analysis tool like jmp, R, or even excel macros. This doesn't seem strange until you actually see it. This can exist ONLY because leadership has made it an actual priority and backed that up with serious resources. As expensive as enterprise BI tools are, they're still cheap compared to training and cultivating expertise in the staff.
Sadly, however, the norm is organizations that say they do data-intensive stuff like 6-sigma but in fact the core of their data practice is nothing more than an overworked database guy and a bunch bs-artists whose main tool is PowerPoint.
As someone who does skilled labor, I don't see why this is true.
> Don't let your enterprise make the expensive mistake of thinking that buying tons of proprietary tools will solve your data analytics challenges.
...
> The truth is that a top team of data scientists can achieve great results ... This is broadly the approach that our own data science teams take when working with clients.
You'll only hear about the advantages of, say, git, from people who use git. Does that mean they're telling you about git in order to get git some more market-share, so that development effort will be even more concentrated on it? Probably not; they probably just think git is a good idea, and therefore use it themselves.
I know plenty of teams are using Flume/Storm/etc, with unit tests, code reviews, etc. But that's like <0.01% of all ETL-related jobs on Earth. For everyone else, the biggest challenge with BI is doing basic, absolutely trivial stuff the right way. Why? (Forgive my generalizations, I'm too cynical to tone them down).
ETL is usually done by inexperienced people because it's unrewarding, uncreative work that doesn't require a lot of skill to do. With remedial SQL skills and some Excel magic, a high school dropout can do ETL tasks that are up to par with businessy (read: not Engineering) standards. Why are Engineering standards not followed for ETL? Because ETL happens whenever EndangeredMiddleManagerX says "We need this new report because SomeGuy wants to feel relevant/smart at his next meeting." If they can produce that report, even once, and the numbers are believable, everyone gets a pat on the back. Repeat ad nauseam. In that kind of work environment, you can't hire a decent engineer even if you try. Best case, you get someone with ~2 yrs of experience who will leave the moment they find something better. Everyone else knows to stay the hell away from that bs.
Big surprise when the same EndangeredMiddleManager types look at the past couple years of data and notice undocumented tables, columns with NULL+"N/a"+" " for missing values, databases that have disappeared and can't be rebuilt because the code was never checked in...Business schools are supposed to teach you that focusing on the short-term is dangerous. Guess that doesn't matter when your career plan is to keep switching companies/getting promoted every 1-3 years, while you leave untold trails of festering crap behind you.
Whenever I see this level of execuspeak I have only one advice: run. This is not the pitch of an efficient outside contractor. This is the pitch of a resource-sucking nightmare of a project that will survive for years generating nothing much beyond powerpoints.
I want to develop a UI that makes a complicated product simple, so users can actually use the features effectively and get their job done. I don't want to build a system that is just putting pretty colours on a convoluted mess of configuration settings and interacting corner cases. That's what some of my clients' competitors do, and their customers have to spend a fortune on consultancy just to make the box work, and my clients pay me because hopefully I'll do better.
Oh, wait, back up. My clients' competitors are making a fortune in consultancy because their customers can't make their boxes work. Good UI is an anti-feature if you can sucker customers into paying for consultancy as well as your product but you can't reliably convince customers to pay more for a product that is better in the first place even if it's cheaper overall, and both of those things may be true more often than we'd like to admit.
This does not leave me in a comfortable position either commercially or ethically.
> In this age of so-called Big Data, organizations are scrambling to implement new software and hardware to increase the amount of data they collect and store. However, in doing so they are unwittingly making it harder to find the needles of useful information in the rapidly growing mounds of hay. If you don't know how to differentiate signals from noise, adding more noise only makes things worse.
Among other benefits, training allowed them to step up quickly to solving real problems rather than doing time-wasting POCs.
Unfortunately the talk is not posted online that I can find but it was a great antidote to silver bullet technology fixes. Ironically the conference had a huge vendor pavilion with about 100 companies trying to sell silver bullets.
And its made more expensive by the having to have a very expansive security group attempting to fit external pieces into an internal puzzle.
What I think is the worst part of this is that the assumption is that software architecture and design then becomes a task that revolves around the existing and purchased tools, and very abstract pricing considerations can strongly influence these toolchains. It leads to very strange forces acting on your definition of a good technical hire. Almost none of them are beneficial for a robust and/or diverse technical organization.
At least it didn't prompt me to sign up for their news letter in the 30 seconds I spent on the site.
Data in the enterprise is a mess. A fancy tool won't fix that, only skilled people can.