Data Is Never “Raw”
thenewatlantis.com
thenewatlantis.com
Instead, a more useful definition are that "Raw data" are observations of things. Those observations can occur in lots of different ways (surveys, logging calls, billing system records, just to name a few). It's important to always distinguish between the method of observation and the thing that's being observed, though.
Similarly, once you have "raw data", you can analyze it. That analysis too is yet another transformative step, which can introduce more biases and errors.
In general it just means the original source data as distinct from processed, selected, edited or derived data.
Et cetera. To some extent the ability to do science consists of saying, "Oh, this is good enough to go on," ignoring the bumps, and being mostly right about that.
Also..I'm fascinated by Delbruck's principle of limited sloppiness, where being lax with the protocols (e.g. letting your nose drip into your petri dish) is what leads to breakthroughs - be too careful and the unexpected doesn't happen.
1. There's no "raw" because even the data collection device and platform will do some form of processing (eg. quantization).
2. There's no such thing as "ground truth" only "ground reference" because truth is fuzzy.
The best raw data is "raw" as in food, but often it is "raw" as in sewage.
I was recently working with some open public data, and it was very much the latter. For example, there were a few important columns that were filled as free-form text by police officers, and contained all sorts of random entries that made the column basically useless. Sometimes they were unclear abbreviations or jargon, and sometimes they were what appeared to be keyboard-mashing, and possibly some bad OCR. It is obvious that the data collection was only designed for events to be reviewed individually, if at all, whereas in the tech world we know to always design our data to be analysed collectively.
I prefer to work with data that is "raw" as in food.
On a personal project, I added a processing stage before my "raw" stage. I decided to call it the "alive" stage.
In analytics, "raw" often means "the unaltered contents of the application database". This is hardly "unprocessed" or "natural", but to alter it ("clean" it) might lose information which turns out to be important later. An analytics person may express exasperation that the application database is so idiosyncratic. If it were up to them, the "raw" data would be cleaner, or more complete, or less noisy.
But the application database is the way it is because it would be infeasible to drastically change it. Certainly nothing can be done for the historical data that's already been collected. Perhaps in the future, data could be collected in a cleaner or less noisy way, the schemas normalized or redesigned, but any proposed changes must compete with the present inertia of the system, and with the need to maintain existing functionality. That is, any such changes must be feasible.
For physical experiments, "raw" data is that produced by sensors that were feasible to construct and operate given available technology and resources at the time. One might imagine that "rawer" data than that might be collected some day in the future. :)
Usually raw data is available but not included in it's entirety in reports. The dat included in reports will typically be data after processing. This has just always been my experience with the way studies and such tend to be carried out.
You're welcome to point me at an unstructured noise source to support your point, though.
I think we may be speaking past each other.
However, it's not like the electrical noise came to be by magic -- it's the result of many interactions that have a structure, and hence impart that structure on the "noise", we just lack key facts to be able to interpret that signal.
Data is most commonly used as a singular mass noun, and the headline makes more sense with it as a mass noun than as the plural of datum.