So you want to offer a public API …
zemanta.com
zemanta.com
Edit; hire 'average' programmers from oDesk etc to build 'stuff' against it 'for clients' (threat it as a serious project) to figure out how easy it REALLY is to make things with it. Average programmers are, well, the average, so they will be the most likely consumers of your API. If they have no clue at all, you didn't go far enough perfecting everything. For instance we hired a bunch of devs to do a Windows plugin for our monitoring system and we found out our documentation and examples were completely unusable for them. We'll have to rewrite/refactor all of that and have them re-try.
So if you really want people to use your system (API, library, IDE, programming language, compiler...) you should run usability tests, write easy-to follow documentation with examples and answers to common problems, make it super easy to download (if it's software), maybe make a web page where you can try it online (that helps a lot), have a nice, smooth user interface, make it behave similar to something users have probably already used, have a good name, a nice logo, a nice-looking webpage, take care of its search-engine positioning, offer services such as forums, maybe advertise a bit, sell shirts, give a free night at a spa to one of the users every month (hey why not)... wow that list came longer than expected. Well you get the point.
I have an alternative suggestion: "Your API (even more so than your product) is most likely never going to be used by anyone, so engineer to get it out the door as quickly as possible and test the waters."
Authorization, throttling, quotas, and monetization are all things you can worry about in the happy-but-unlikely case that your API takes off. In the mean time, just get it out the door. There are some things you can do to future-proof yourself with zero effort:
* Require every request to submit a contact email address as a parameter. This gives you a crude way to find out who is using your API and how much, and gives you the ability to contact them when you make breaking changes. Clients who don't submit contact information have nobody to blame but themselves when you make breaking changes.
* Put your API on a different domain from your main site, ie api.example.com instead of www.example.com. This gives you more flexibility to move it around without breaking things.
* Documentation is critical, but easy - provide examples. Just a series of "Want to do X? Formulate a link like this" and make a real link that actually works.
* Don't overthink the API. Yeah, use REST principles and don't be a total idiot. But accept that your first API is just an experiment. If your API is successful, you will end up rewriting it. The most important things are that 1) it works, 2) it's useful, and 3) users can figure out how to use it from the documentation. Just get something out there, even if it isn't pretty. If you need to make breaking changes, that's what the contact email is for.
We (https://www.voo.st/) added an API in response to a partner who asked for it. It took a day, half of which was formatting the documentation (https://www.voo.st/developers). Guess what? It's never been used. Chances are, your API will be the same way.
For a simple developer, output format is just another parameter and it shouldn't be hidden in HTTP stack.
At least in python, making requests becomes quite a few more lines of code if you want to add headers. There's usually a short-hand command if you just want to get full data from an url and a "create-an-object, then set these parameters, then this, then open a connection then read what's there".
Mandatory headers are just a way to make life of a regular developer painful.
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,/;q=0.8
This says "we would prefer one of these formats, but we'll take whatever you got".
Now I navigate off the page to a different site, and press the back button. Instead of the HTML page, I get the cached JSON response. Now if I change the ajax call to "example.com/objects/1.json" instead, that keeps the URLs separate and the browser won't cache the wrong response. Is there a better way to solve this?
[1]https://developer.mozilla.org/en/HTTP/Content_negotiation
Think of all the different "Accept" headers that clients can send you, each one of these will get an entry in your cache. And now add other headers that should also be in "Vary" like "Cookies", "Accept-Language", etc. and your cache will be virtually useless.
Also, I might be nitpicking here, but I think the difference between a URL and a URI is that the URL explicitly defines the media type.
If that's the case, then the advice would be to use URIs pointing to resources in combination with `Accept` headers.
No, no. An URL is an URI. And the URL doesn't define any media-type, those are hacks created frameworks; there's no such thing in the spec.
From the linked answer:
> * you break permalinks
> * The url changes will spread like a disease through your interface. What do you do with representations that have not changed but point to the representation that has? If you change the url, you break old clients. If you leave the url, your new clients may not work.
Putting the version somewhere else in the request doesn't fix this. If you drop support for an old API version it is going to break stuff. It is easier to spot the issue if you return a 404 rather than a 500 or 400 error, or even correct data that breaks the app consuming the API as it is expecting something else.
> * Versioning media types is a much more flexible solution.
The main issue I have with this is I have had the "pleasure" of working with pretty stupid people implementing clients on my API. They easily get POST and GET requests mixed up, HTTP and HTTPS, whether or not the request should be authenticated. I don't want to add something else to confuse them even more...
Where does the madness end? Do you treat everything as though it were GET and stick the real method (and every other request header) in a query parameter? I think it's ok to assume a basic level of understanding of one of the best documented specs in computing.
A public API is about 99% of what we (Crocodoc) offer. While designing the latest version of our API, I always erred on the side of staying true to REST principles. Fortunately there were others on our team that erred on the other side. The result is a happy medium that certainly veers from the canonical definition of REST but is arguably easier for Joe Developer to play with. Guess what...we still receive an incredible amount of support requests from developers getting stuck.
It's tempting (and fun) to build APIs "the right way" and use little tricks like putting versioning in the media type. However, that's pretty often at odds with product/user experience/business considerations. Why offer a product that only 10% or 20% of your customers are savvy enough to use?
So you are willing to do the work of those other devs for them, and cripple your own API in the process?
Due to licensing of content, it would be iffy to monetize directly (technically even the ads are in violation of that), so I don't even consider it.
Nice problem to have though.
EDIT: After reading http://thereisnorightway.blogspot.jp/2011/02/versioning-and-... I am not so sure anymore. They still have quite an awesome set of examples though.
I developed, and am responsible for, a moderately complex b2b web API that is been in use for about three years. A b2b API embodies a contract between organisations, and it's important to understand that this is different to b2c (e.g. Twitter) which is less restrictive.
The biggest issue for me has been dealing with change. User requirements change. The underlying system changes and accretes new features that need to be exposed in the API. How to deal with this without either compromising the integrity of the API's design or (at the other extreme) breaking systems that use the API, is not easy. I don't know of any easy answers, but design your URLs and data schemas with versioning in mind and get to know the developers who use your API.
Documentation is very important. If your API exposes complex data then make sure that you have good documentation on what every data element means.
Actually a lot of the APIs that use us are B2B because you often have different tiers of access rights and limits + want to have approval steps etc. for sign-ups. These are pain to code yourself.
The change part is a big deal - we normally deal with it in two ways: 1) launching parallel services for major versions so people consciously migrate their apps, 2) incremental updates - docs are versioned but all the same calls work on the same endpoint. How the version is flagged in the call is rather independent (URL, header, type).
We also use the swagger docs framework (http://swagger.wordnik.com) to create interactive docs to try keep questions on the API itself to a minimum (people can play with it themselves).
I must admit to being a big fan of RESTful APIs, but I'm skeptical of treating any technical approach as being universally applicable. Are there any likely scenarios where a RESTful API for a web service is not going to be a good idea?
When you have a lot of high-speed data, you will want to go with sockets and a live stream of small bits of data. Like Twitter's firehose for instance, or getting Wall Street data for your bot trader.
For example we implemented our API with both XML-RPC and REST + JSON interfaces (http://www.memset.com/apidocs/intro.html). All the methods are available and work the same way in both interfaces. There's some extra work to deal with some of the types (ie. dates in JSON), but it's not too difficult.
So we have one implementation/docs, but two interfaces to access to it.
Edit: One of the interesting things about their RESTful API is that different versions of the API are explicitly in there as resources.
Sorry if I didn't explain it very well, the docs are clearer though.
A good rule of thumb is: if the documentation mentions "methods" as anything other than the uniform ones (if you're using HTTP, that would be GET, POST, etc), then it's RPC, not REST.
Note: I'm not saying your API is bad; RPC is a completely valid model. I'm just saying it's not RESTful.
EDIT: to be fair, in the docs we never say it's a RESTFul API, it was MY mistake in my first comment.
Zemanta's API is one of them. We're not dealing with resources, we're dealing with analysis of text (natural language processing, information retrieval, etc...), so except of the basic url scheme, there's not much that you can take from the RESTful idea.
You could just POST the text, get back a bunch of links to the resources and then GET the ones you're interested in.
Assuming that certain "objects" are more expensive than others to generate, a resource-based view with lazily-generated representations could be more efficient than wasting resources without knowing if the client needs them or not.
What's expensive is the analysis of the original text, not creation of the objects. If you want to do what you are saying you would need to create and store objects server side in order not to have to re-do the analysis the second time around. There's no point in that.
Stateless:
The disadvantage is that it may decrease network performance by increasing
the repetitive data (per-interaction overhead) sent in a series of requests,
since that data cannot be left on the server in a shared context. In addition,
placing the application state on the client-side reduces the server's
control over consistent application behavior, since the application becomes
dependent on the correct implementation of semantics across multiple
client versions.
Cache: The trade-off, however, is that a cache can decrease reliability if stale
data within the cache differs significantly from the data that would
have been obtained had the request been sent directly to the server.
Uniform Interface: The trade-off, though, is that a uniform interface degrades efficiency,
since information is transferred in a standardized form rather than one which
is specific to an application's needs. The REST interface is designed
to be efficient for large-grain hypermedia data transfer, optimizing
for the common case of the Web, but resulting in an interface that is not
optimal for other forms of architectural interaction.
Layered System: The primary disadvantage of layered systems is that they add overhead
and latency to the processing of data, reducing user-perceived performance.
I fully recommend reading the thesis instead of relying on the vox populi that often times misrepresents what these concepts really mean and what advantages they bring to the architecture of the system: http://www.ics.uci.edu/~fielding/pubs/dissertation/rest_arch...