Sidehelm: a pipeline to validate, test, and pull CSV data
sidehelm.com
sidehelm.com
The real cost is that the fight is never over. Whenever a new customer comes in with a CSV so badly formatted that it is rejected by your system, further analysis reveals that there's actually some sanity behind the madness. That you could in fact add new rules to detect and correct the situation.
We've seen so many horrors. European dates ? Try a file with three different date formats over six date columns. Quotes that should be parsed in some columns but kept verbatim in others. Columns that contain raw binary data, without any escaping (but you can still infer the boundaries from context). UTF-16 with a byte skipped in the middle. Tab-separated files where the first line is space-separated instead. Cells that are 8MB of white space.
There is no end to the... creativity of CSV file authors.
Indeed, it never ceases to amaze. That's why it's so hard to imagine a "general" solution to this problem. You can think of something to handle to 90% most likely business use cases, maybe... but not the fat tail.
Eh, commas and new lines are common, let's use pipes. and count fields.
:'(
I cannot really think of many services off the top of my head that would be so willing to give significant chunks of data to another pipelining service.
It already includes some tools for working with and displaying CSV data.
Basically you don't want to rely on something that (1) requires a whole new service / monitoring later, (2) intrinsically exposes your data to untrusted sources and (3) can of course go "poof!" at any time -- unless you absolutely have to, or it's for something mostly ancillary to business and your data pipeline. (Or you're fairly small and just don't have a lot of people).
Especially when it's just CSV parsing, after all.
Could you please link to alternative solutions?
Does AWS Glue solve the same problems? https://aws.amazon.com/glue/
After a brief search, this seems promising:
I'm slowly putting together my own solution for this, but if suitable open-source software already exists I'd be happy to make use of or modify that.
What type of data/industry is it? It'd help us to know where to look for people that might need us.
In most cases programmers seem to have an incorrect set of assumptions around encodings, escaping etc.
https://docs.google.com/presentation/d/1vjm5YdmOH5LrubFhHf1v...