- Good stock market historical tick data (and streaming data). Opentick tried for a time, but mostly you need to contact expensive services. This makes it hard to mess around.
- A good local event API. Lots of companies have tried, few have good results.
- LinkedIn: They've been promising an open contact book API since 2007, but have kept it closed. If they had an open API (that lets you actually store data/invite), it would mean a lot of sites would be building on them.
- My dream: The big scientific journals would require everyone publishing a paper to upload all relevant datasets to a central repository which could be queried.
- Anyone you have an account with (Financial firms, banks, vendors) would have a standard commerce API. Sure, today you can export stuff in Quicken, etc formats, but Mint had to do a big deal with Yodlee to get the data in a uniform, queryable way.
And, mad props to Microsoft for opening the Bing API with pretty good terms. Google used to have a search API and it had horrible terms. Then they decommissioned it. Who would have guessed that MS would be more developer-friendly than G?
Anyone who's done Microsoft-based development? MS is incredibly developer-friendly compared to many platform vendors. It's a point of business strategy for them (if memory serves me, that's what Ballmer's infamous "developers, developers, developers, developers" thing was about in context).
Just don't try and mix in any technology from one of MS's competitors and you'll be fine.
Unfortunately their licensing is from an era that has come and gone. Building software that uses things like Windows Server, SQL Server, SharePoint or Office means to limit your scale to what they call "Micro ISV". You provide that package to your client and 95% of your revenues go straight to Microsoft.
You can build on their stuff but you won't scale and you won't grow, not because their technology doesn't scale, but because their licensing doesn't scale.
What about their licensing for Azure? Does that scale or is it more of the same?
However, look at SQL Azure. Microsoft thinks that their SQL offering is worth paying 66 times what you pay for google's database (or Azure BLOBs&Tables) and that's just for storage alone ($1 per GB, maxing out at 10GB right now). Add data transfer and you may get to 100 times.
Yes SQL Server has vastly more features than google's db, but does it make me 100 times more productive or profitable? I don't think so. If I need more SQL features I could run Postgres on Amazon or use Amazon's MySQL service for roughly 10% of what SQL Azure costs.
So if SQL Azure is the model of what's to come then I think it's indeed more of the same.
I work on a site that's trying to do that for the learning sciences.
http://pslcdatashop.web.cmu.edu/
Most of the data so far is from various studies in the Pittsburgh Science of Learning Center, because the head of the PSLC can tell the researchers to put their data in there, but we'd like to convince others to share, too.
I also happen to be working on a web services API to this data at work right now.
It always surprises me when people say this. Apparently it's not widely known that Google do make their search API available: http://code.google.com/apis/ajaxsearch/documentation/#fonje
(Before anyone says "that's just the AJAX API", please READ THE LINK, and scroll down to see the Java & PHP samples)
For example, if we had voting records, we could finally stop hearing all this "but you voted on -----" "no I didn't go back and check the record" "no look you voted on ----- which is basically the same" "shut up I served in the war" "yay war" stuff every election. We could just look it up and say "hey look you did". The Times "Congress API" is on its way to this, but last I checked, all it had was the attendance records.
Though, I think you're right. The documentation is quite poor.
storeNewBrainMemory();
and saveBrainToDisk();would be pretty cool...
I find that logic amazing, but considering how old the online banking software is and the high risk of changing it wholesale, I don't think they are in a position to fix it in the short term.
Specifically:
TuneCore, as it would make my startup idea a whole lot easier.
School registrations. I really wish there was some standardized API for universities. That'd make it possible to plug in the classes you want to take along with when you want to take them and get back a personalized schedule. As it is now (at least for my school), you pretty much have to write down on paper the classes you want to take along with when they're available and do it all yourself. I'd prefer something like this be standardized so one could make one website that would serve all universities and their students.
http://courses.illinois.edu/cis/2010/spring/schedule/index.h... http://courses.illinois.edu/cis/2010/spring/schedule/index.x...
I really wish that they will provide one in future.
But more than a specific API, it would be cool if websites simply provided an XML(/JSON/etc) version for every urls. Eg, http://news.ycombinator.com/item.xml?id=955077 would return the data in this page in XML format. This would be pretty simply to create (at least as read-only API), handle the situations where people resort to HTML scraping and effectively remove the need for API docs.
Definite business opportunity there, the two services I have found (in this case I'm looking for Texas data) both consider using an iframe "integration".
I'd love to see Google get ahold of the data and make it available through Base.
I think the leagues fight hard to disallow this although they have recently lost a couple high profile lawsuits.
http://arstechnica.com/tech-policy/news/2007/06/mlb-tries-to...
We offer stuff ranging from simple on-demand calls to our web service (you purchase credits, docs cost you X credits per access) to a Perl-based app that captures an XML feed and parses and inserts the data into a DB for you to use as you will at your end, usually MySQL but we support MS-SQL too.
We've got some pretty big clients: ESPN, USA Today are amongst them. We supplied Google with Olympic content during the Beijing Olympics.
http://l-mail.com http://postful.com http://earthclassmail.com http://postalmethods.com
I've been building a little side thing called snailpad (http://www.snailpad.com) and have a beta API running for some of the paying customers at this point.
1) Allowed me to find high quality, license free images easily.
2) Let me access LinkedIn data. I know LinkedIn has an API, but its not open.
No kidding. He's looking to get "comments for a given thread, score of posts and comments, and user information", which you can use an HTML parser (like Hpricot if you're using Rails) to retrieve any info you need.
EDIT: If you're going to downvote, please leave an explanation to teach me what I'm missing. I'd sincerely appreciate that. Thank you.
You're being annoying. Nobody wants to screenscrape, it's awkward and fiddly and fragile and verbose and requires reverse engineering and is wasteful of bandwidth.
It really doesn't matter if you can get the information by screenscraping. The thread topic is "what would you like an API for?", not "scorn people about what they want an API for."
Care to explain why I'm being downvoted?
Because I want less of this sort of thing on HN. I want it to fall to the bottom of the page, and to discourage it in future. I cannot articulate precisely what it is that I want less of, it's a big vague fuzzy blob of things, some of which I even agree with, and your few posts here and above here are within it, so down they go.
I think I will start without an API anyway. Most of the things I need to do should be possible with the RSS "API".
Not at all an API, but since we're on the subject.
I like reddit's format, where you can add .rss to almost any url and it returns the same result as an RSS feed
They're actually more detailed than the boards in the stations as they show current train locations.
I've played with it and seems good enough.
It's a whole lot better than nothing, but it's far from perfect.
Fortunately there are a few different startups working on this, and a rudimentary way to hack it with Google AJAX API but there's still nobody who can allow me to simply punch in name+city search and give me the address, geocode meta reliably for any city in the world and without cache restrictions.
Google has it all but needs to open this data up better.
For further googling, see terms ICR, OCR, character recognition.
As a user, I have had very good results with Finereader. Have not tried Tesseract. Parascript was good with online character recognition, but that market is small, and I have not looked at them for a while (disclaimer: I used to work in a previous incarnation of the company).
http://www.evernote.com/about/developer/api/evernote-api.htm...
"Not perfect" would be fine for many non-life-or-death signs.
This was the only way I could get the classes I needed to graduate on time
What would you build with it? Some people are using it in cool ways right now, e.g. hype machine uses it to power a twitter music chart.
There have been many CL clones(kijiji), but none gained enough eyeballs to make an interoperating API work. I suspect these attempts failed because they didn't pay attention to the community and instead acted liked closed corporations right out the starting gate. I thought the CL killer would appear on Facebook, but that hasn't happened because of totally different dynamics in Facebook's ecosystem.
Actually, I'm surprised that I haven't seen someone one here posting a hacked up API wrapper around Craigslist as a mini-startup.
And the good local events api someone mentioned would also be great.
http://musicbrainz.org/ is extremely good at this.
I've seen a few people respond with this "you can just html scrape that." Sure you can HTML-scrape for the information, but the topic of this "Ask HN" is about what APIs you would like to see. Maybe he already HTML-scrapes ESPN.com, but would prefer that there was an official API for it, no?
http://gd2.mlb.com/components/game/mlb/year_2009/
Here's an example day's scoreboard in JSON:
http://gd2.mlb.com/components/game/mlb/year_2009/month_07/da...
http://github.com/nickspacek/Net-Google-Tasks/
Works decently. Could easily be replicated in any language.
This is from the same framework that has winTheGame() and doMyHomework()
>>> data = urlopen('example.org/lovelydata/2009/') YetAnotherException: [...]
It needs some kind of web form cookie login non-standard authentication. Aaargh I can't spend any more time on this sub-project it was only supposed to take two minutes!
>>> import bigGuns Warning, Universe enters a fragile state. Tread carefully. >>> data = justMagicallyWork(urlopen('example.org/lovelydata/2009/')) Success. Cost: 12 Karma. >>> del bigGuns normality restored
Phew!
Maybe there's one now but 3 years ago there wasn't one and I had to build a scraper for Yahoo movies to build the product I wanted.
(From a Microsoft C# compiler/parser guy, no less).
make_this_many_dollars_magically_appear_in_my_bank_account(1000000000);