With PUblic Datasets, the account making queries pays for the queries. NPS only pays for the storage (which is minimal).
With this API, NPS has to pay for every call to the API. That’s not cheap.
With PUblic Datasets, the account making queries pays for the queries. NPS only pays for the storage (which is minimal).
With this API, NPS has to pay for every call to the API. That’s not cheap.
> to access public data
Keyword is access. Hosting on AWS is an implementation detail that doesn't block the end consumer from accessing the data.
* AWS
* Comcast for their internet service
* Apple for their laptop
* A number of software providers for their development tools.
But asking the customer to pay google to query the data is crossing the line?
Site hosting is not a customer cost.
The rest you list are costs orthogonal to this service.
> But asking the customer to pay google to query the data is crossing the line?
Yes.
Why are you arguing for a US government agency to require its citizens to pay for access to data which they have already paid for by funding said agency?
That's the difference.
That's fundamentally a lot worse than the government paying hosting costs to one particular vendor for a commodity service.
People have quite literally died over the issue of public access to public data. It’s quite an important belay point to arrest the deterioration of the spirit of open networks.
What?
Taxes are, generally, money.
That was what was suggested upthread.
Requiring the user to have certain capacities to access data, where those capacities are provided by a number of competing vendors (and some by free, gratis and/or libre sources) is a very different thing.
So are you ok with some chinese APP company making 50 crappy NPS themed apps and having taxpayers pay for the backend?
Addressed in a comment in another subthread, which I know you are aware of since you responded to it, too: https://news.ycombinator.com/item?id=39086270
I will make that trade every day of the week if it means access continues to be through a standard protocol (HTTP) and not beholden to any particular vendor.
That said, I hate Medium with a passion and that things like the Netflix tech blog are hosted there.
I am perfectly fine with it being considered part of the basic, taxpayer-supported functions of government agencies to be providing the public with relevant data.
If there is a concrete abuse or wildly disproportionate cost problem in particular areas, that may need to be addressed with charges for particular kinds or patterns of access.
[0] Douglas Adams, Hitchhiker's Guide to the Galaxy
I'd be pretty happy with sqlite dumps too.
I don't really have an issue with the REST, though. I wouldn't be surprised if this was just a standard and cheap to set up Django+REST libraries stack. Yeah, the compute costs are higher than transferring static files, but I'd be shocked if this was taking enough QPS for the difference in cost to make a meaningful difference.
I get wanting the government to be responsible, but this veers a bit too far into Brutalist architecture as an organizational principle.
So someone can host a ripoff NPS app on the App Store and taxpayers now pay for content hosting?
You can access for free, but if you abuse or break the TOS your access is revoked. Done
Then use their API to populate a BigQuery public dataset and make available to all.
Otherwise, perhaps we, as outside observers, need to consider the possibility that those whom made the decisions to provide this service as such did so for reasons which we may not be aware.
nps-public-data.nps_public_data
parks
people
feespasses
What’s spammy about my post? I have asked people to focus on costs when they make general statements like “all govt data should have a REST api”.
I think we can dispense with the risky argument, because this API has existed for years without issue.
When you make the argument that "X is too expensive," the onus is on you to prove it's expensive in a relative sense, not simply in an absolute sense. Saving $100 matters if you're spending $1000; it probably doesn't matter if you're spending $10m. Feel free to convince us: make some estimates, crunch some numbers, and look at existing NPS IT spending and see if they seem ballpark reasonable. Otherwise you're just banging on about a left-field solution that almost no one wants because it's putatively cheaper (but by how much, you can't say).
Sure the concept of rest APIs is mature, the but each implementation is untested.
People here clearly don’t like a querier pays model and that’s fine. But should NPS still reinvent the wheel across the SDLC to serve this data? I think there’s a compelling argument in there.
REST API compute is very expensive when you include compute costs, transfer fees and admin costs to keep it up.
Not to mention the cost to implement a bespoke API and deal with security issues.
All to make CSV available!
With BigQuery they just copy the data in via CSV and Big Query handles the indexing & query engine.
Open source tools that will present a simple, read-only REST API over an SQL db with little to no custom code exist (so do proprietary tools, often from DB vendors and sometimes as part of SQL db products.) Same with NoSQL or sorta-SQL storage solutions.
The idea that they have to write a bunch of boilerplate code to do this is false. They might choose to do that, but its defintely not necessary.
> Authentication, throttling, threat prevention, encoding, etc etc.
Again, open source canned solutions that take a little bit of configuration exist for many of those, and some of them are likely shared services that they just point at whatever service needs them.
On that note, what would the processing entail? Processing the get request and packaging the entire dataset into a REST object right? Or is it a more complex API that lets you run queries against the dataset? For that matter wouldn’t downloading a CSV also have to be packaged into a REST object?