I wonder if there's anything better.
I wonder if there's anything better.
On the other hand, I used their python API a bit and for loading tables it's way too complicated, with the documentation going into great detail for faffing with metadata but not for actual loading (and nothing like a simple `read_datatable` function).
That said, because it's just a folder with CSVs you can just read them individually, although then there's nothing to take advantage of the metadata automaticall.
Point taken about the API. The Data Package [1] and Table Schema [2] libraries are generally designed as low-level libraries for building higher-level applications using the specifications. goodtables-py [3] is an example of a higher-level application built on top. But, point taken, we will look at it, and we'd welcome your feedback on the issue tracker [4].
[1]: https://github.com/frictionlessdata/datapackage-py/issues [2]: https://github.com/frictionlessdata/tableschema-py/issues [3]: https://github.com/frictionlessdata/goodtables-py/issues [4]: https://github.com/frictionlessdata/datapackage-py/issues
(I work on the Frictionless Data specifications and tooling at Open Knowledge International.)
CSV has many, many warts. However, it is the best thing we have right now for serialising data in a way that is easily read by humans (and consumer-grade software) and machines. Libraries like our Tabulator [1] which is used under-the-hood help provide an API to deal with many of the gotcha's when dealing with the format.
There's a brief explanation on why CSV was selected on https://specs.frictionlessdata.io/tabular-data-package/#why-...
Any text format can be broken by broken character encoding. I saw plenty of XML being used without any charset declarations. And JSON is in same position as CSV.