New Jersey Open Data Initiative
njleg.state.nj.us
njleg.state.nj.us
My city's school district makes images of the spreadsheets and then publishes them in a pdf. USELESS
"The policy and procedures shall include ... technical requirements with the goal of making electronic data sets available to the greatest number of users and for the greatest number of applications, including, whenever practicable, the use of machine readable, non-proprietary technical standards for web publishing"
"The Treasurer shall consider various means by which to develop a set of universal data formatting standards to effectuate the purposes of this act, including working with other State departments and agencies, and contracting, if deemed necessary, with nonprofit organizations, commercial vendors or third party groups for this purpose. If such standards are developed and adopted by the Treasurer, they shall be the format that each State department and agency will use to provide existing and future electronic data sets"
Wouldn't that cover PDF, even though it might be deeply inappropriate for it's content? PDF has several ISO standards, and is most definently machine readable.
PDF should also meet the goal as well:
> available to the greatest number of users and for the greatest number of applications
Huge number of PDF clients, such as PDF.js that powers Firefox's reader.
> If such standards are developed and adopted by the Treasurer, they shall be the format that each State department and agency will use to provide existing and future electronic data sets
Knowing how much government uses PDF... It is probably incoming.
But it makes little/no sense for spreadsheet data. Great for graphs and prose, however.
https://www.w3.org/2013/04/odw/odw13_submission_52.pdf
PDF with csv attachments is a good compromise but I rarely see that. Usually it is PDF or Excel and the PDF is a nightmare to munge.
EDIT: Tabula[1] is an open source tool to extract tables from PDFs into CSV/Excel format in case one really needs to process data trapped in PDF documents.
> Caveat: Tabula only works on text-based PDFs, not scanned documents. If you can click-and-drag to select text in your table in a PDF viewer (even if the output is disorganized trash), then your PDF is text-based and Tabula should work.
You should be able to sue them for misuse of public funds. Dedicating tax money to convert a useful format (CSV, spreadsheet, etc) into a useless or at least less useful one (ex: PDF) is a clear violation.
Not so yay: generally nobody’s going to be liable for incomplete or inaccurate data, with an exception of “gross negligence, willful and wanton misconduct, or intentional misconduct”.
Understandable: “No personally identifiable information would be posted online”
Unclear: “State departments and agencies would not be required to make data sets available upon demand”
[0] It seems that the availability and appropriateness is where things might stall. What if a department gives the Treasurer a dataset full of personally identifiable information? They thus fulfill their duty per this bill, the dataset is available, and yet public sees nothing, unless Treasury goes an extra mile and takes care of anonymizing the dataset.
The site: https://www.bathhacked.org/ The data store: https://data.bathhacked.org/
We run hackathons and also reach out to various organisations for data sets that although we may not be able to publish in the datastore, can result in very interesting tools.
The most recent data we got our hands on was the bike share hire programme. https://www.youtube.com/watch?v=4AfocuESnSg
Bath and North East Somerset council now have a policy to open as much data as possible to the public, directly as a result of BH.
"No personally identifiable information shall be posted online unless the identified individual has consented to the posting or the posting is necessary to fulfill the lawful purposes or duties of the department or agency."