Can this be considered an example of a tool that partially replaces the work done by a data scientist? At least it can save a lot of time.
Can this be considered an example of a tool that partially replaces the work done by a data scientist? At least it can save a lot of time.
Creating simple predictive models where your problem is already easily narrowed down to a "given x predict y" definition is pretty trivial. Having it automated is nice, but not exactly a hard thing to do.
Genuine question: how many people have jobs where those kinds of problems form any significant part of their workload?
I also often see a response to this sentiment along the lines of, "Yeah, but there's also data cleaning..." etc. My reaction to this is mixed. I mean, sure, there is also data cleaning involved, but is this really where people spend most of their time?
My team spends most of our time doing the following:
1. Formulating problems. Figuring out the various different ways that a real-world problem can be expressed mathematically and feasibly attacked computationally.
2. Engineering software to implement the solutions to these problems, sometimes using some of the (amazing) frameworks out there for ML or probabilistic programming, but often having to develop our own approaches from scratch.
3. Doing all the management, stakeholder relationship stuff, business cases, etc. that make your work relevant and possible.
4. Getting data. Always an issue.
I'm very genuine in my curiosity here: are we total snowflakes, and most data scientists spend their time cleaning data and building "given X predict y" models?
I know for me I've had things like a bunch of scanned images of tables as "data". Turning that into something useful took a lot of time.
Whether this is "getting data" or "cleaning data" depends on perspectives and definitions.
Data science will focus more attention on solution finding, data gathering, cleaning, ETL, and business.
CSV is about the barest amount of specification in a file format. It's common to run across files which are some weirdo encoding you can't easily detect, or are a mix of multiple encodings, or a mix of line endings, or which should be treated as case-insensitive (or only for some columns), or which have weird number formatting (or units, and not the same units in every row), or typos and spelling errors, or it came from an OCR'd PDF and there's "page 2" right in the middle of it, or they tried to combine multiple files together so there's multiple headers scattered throughout the file (or none at all), or the top has different columns from the bottom, or it uses quoting differently (obviously not per the RFC), or it's assumed that "nil"/"NULL"/""/"-"/"0" are the same, or ...
In short, data (which hasn't been cleaned by hand) sucks, and CSV doubly so. If you want to put your AI/ML smarts to work, write a program to take a shitty CSV file (or even better, a shitty PDF file!) and generate good clean data, plus a description of its schema. That would be an amazing tool.
So far, OpenRefine is the nicest tool for this that I've seen. Figure out how to make it fully automatic, and everybody with piles of raw data (governments) will beat a path to your door.
I have been thinking about making an open refine type tool for python. Every time I do data cleaning in python, it feels so repetitive.
Automating modeling is a bit easier than automating the other parts, though.