1. Indeed.com - analytics on how many job ads mentioned a specific brand or keyword.
2. Yelp - analytics on how many checkins a restaurant or chain got every month.
3. Apple Store - analytics on how many downloads an app got every month.
4. Amazon Reviews - an API to retrieve reviews for a product.
5. Google Trends - API to retrieve historical trends for a keyword
6. Google Search API - API to get search results for a keyword
7. Linkedin Company Page API - API to get the feed for any company page
8. Instagram API - API to get the feed for any instagram user or search results for any keyword
Here are the non-data API's I wished exist (that could be potential low-hanging fruit startup ideas):
....None
For example, here is the Amazon Reviews API: https://docs.aws.amazon.com/AWSECommerceService/latest/DG/EX...
https://suggestqueries.google.com/complete/search?client=fir...
This returns an array of 10 of matches for the term in `q`. This is how the AnyComplete command for Hammerspoon works: https://github.com/nathancahill/Anycomplete
Is that how fast the internet is in the US?!
I mentioned this on HN once before and got a "I thought it was just me!" response.
1. historical data
2. real-time data
3. trading api
credit card usage/history data (only for myself) that is across all credit card brand.
Oh the things I would build...
Actually, something that may be even more important is free and open access to APIs. I'm okay with registration procedures for larger volumes, but not for hobby use. It seems to me an unnecessary obstacle. If the hit rates are similar to what a user with a web browser would produce why is my script forced to register when the web user is not?
If I were planning to make a profit on that data, I might well have spawned 1000 boxes to speed things up (took 2 weeks of 24/7 scraping on 5)
But without registration, theres no way I can think of to associate multiple machines with a single person.
And presumably if you're rate-limiting on individual users of the api (and not the total usage across all users), then you're trying to maximize the number of users accessing the api simultaneously (ie you don't want 1 user maxing out resources, denying all other users).
But with things like cloud computing, its trivial for a single user to have an absurd number of machines (legally). AWS charges on compute-hour, not number of machines. Running 1 instance for 50 hours and 50 instances for 1 hour has the same cost.
But for the server, that difference is significant; you apply rate limiting because you don't want to be hammered for a short duration followed by nothing, at least not from one user.
Registration can also be bypassed somewhat trivially (temporary emails and aliases), but its a good deal more effort than bypassing non-registration rate limiting. And presumably most people who want to scrape a dataset large enough to bother bypassing the limiter are nautrally, by their job, aware of what I've described. Bypassing registration takes a good deal more work, and more specific tooling
So tl;dr
Rate limiting by ip only affects 1 machine, but not necessarily 1 user
Rate limiting by registration affects n machines, with 1 user
Before cloud vps, 1 machine basically correlated to 1 user. Now its trivial for n machines to correlate to 1 user.
I'm sick of searching 3 catalogs (hulu, netflix, amazon) and finding out none of them have what I'm looking for.
It's really useful for finding shows on services you use.
Probably never should have started using it in the first place.
Disclaimer: I work at Zinc.
Its kind of a developers wet dream if you like that kind of stuff. :)
I really wish I could provide integration with Numbers.app, but as a one-man shop I can't afford to get distracted with maintaining a reverse-engineered file format for one use case.