How Not to write a "REST" API
api.sharefile.com
api.sharefile.com
http://nordsc.com/ext/classification_of_http_based_apis.html
Sharefile is not even a HTTP API (since it doesn't use HTTP methods correctly).
For security purposes, authentication can be further increased by a POST of the "username" and "password" through the HTTP Headers as individual headers instead of the query string.
POST https://subdomain.sharefile.com/rest/getAuthID.aspx HTTP/1.1
Content-Type: application/x-www-form-urlencoded
password: yourpassword
username: email@address.com
What... I don't even...
>https://subdomain.sharefile.com/rest/getAuthID.aspx
Notice the https. The facepalm is just that there's no additional security in using POST vs GET.
And POST over HTTPS keeps the POSTed values hidden from those who might listen, whereas with GET over HTTP it would just be out there in the URL.
And HTTPS requests are encrypted. The whole request, including the "GET /someurl&password=s33krit HTTP/1.1" part. As I said, using POST doesn't add any additional security to this.
That's why professional security audits will ding you for putting anything sensitive in a URL.
And having read and written this particular audit line item about 29074894389734897 times in the last 15 years, let me assure you that logging is clearly an issue.
The bit about how "if someone can see your logs you're already boned" is also message-board-logic more than reality. Logs are shipped all over the place. On a network pentest, one of the things you do when you pop a box is hunt for logs; even after you get root on the machine, you don't automatically have all the passwords that get stuck in the log files. Same goes for all the random LogLogic-style consoles those logs get fed to.
If you have more questions about this stuff, my contact info is in my profile --- just click my username above this comment. I just wanted to chime in to say that 'carbocation was exactly right, but my comments are now repeating themselves, and so I'm done on this thread.
So will everything else, including the URI being requested, and thus the query string in it. Which is why it makes no difference using GET or POST.
And again, if they are compromised, then they are compromised. It doesn't matter if they have logging disabled, someone who would have access to the logs also has access to either the httpd account or the root account. Either way, they can already read your plaintext usernames and passwords directly when they are being submitted. Of course, they don't need your username and password anyways, as they already have full access to the system.
At least one person appreciates someone taking the time to correct this rather serious misunderstanding.
I disagree from the data I'm seeing in the access logs from my SSL-hosted site running nginx. In the logs I can see lines such as:
GET /path/script?variable=blahblah&another_variable=123
EDIT since I appear to have lost the ability to reply to comments: I disagree with SomeOtherGuy2 that The fact that it may get logged is a red-herring.
Ignoring how secure a server with a rogue user accessing it is, it's possible that there will be more than one server involved in this scenario, and central logging servers are common. Will the traffic sent to the logging server be encrypted? And what if the logging server is compromised? You're essentially storing passwords in plain text.
Why?
Also, the top line should read POST rest/getAuthID.aspx HTTP/1.1
Correction: Use of "X-" has been depreciated in a draft resolution.
Edit: As mentioned below, it isn't compliant in its use of the verbs (GET/POST etc).
Second, absolute URIs are perfectly valid, all HTTP/1.1 servers MUST accept them: http://www.w3.org/Protocols/rfc2616/rfc2616-sec5.html#sec5.1
REST was born out of a philosophy that the HTTP protocol already solved much of what you wanted to do. HTTP not only solved this, but solved this a long time ago and with 'great success' (aka The Interwebs ;-)
- Authentication mechanism
- Operations (CRUD) -> GET, PUT, POST, DELETE, ...
- Caching -> Use HTTP caching mechanisms...
- Resources -> URL's
- Formats -> mime-types + use a HTTP Accept: header
- ...
They reinvent the wheel on many of the bullet points above: custom operations, custom format handling, custom authentication mechanisms, ... That's why this really is one of the worst implementation of an 'RESTful' API
HTTP does not have "verbs". That is REST. Just because they aren't using a REST API, doesn't mean they are not HTTP compliant.
I thought that GET (along with a few other methods) were supposed to be "safe" and not result in any action on the server except possibly logging and stuff like that? Is this required to be HTTP compliant or is this more of a recommendation?
Shouldn't Fielding's original thesis be the bible?
1. Passing creds via GET will show up in server logs. 2. The way REST works, you don't want to put anything in a GET which will change anything on the server (which includes generating an auth token). POST is the standard here.
edit It's better that they support POST for the auth creds in a header, but they should still be using HTTP standard authentication methods and return a 401 if unauthorized.
Also, look for already complete solutions for what you want. The whole point of REST is that you don't need to create all these query string parameters or put stuff in headers - the HTTP specs already include most of this functionality. For example HTTP already includes several authentication mechanisms[1], as does TLS[2], which are more secure (and standard!) than a form-based login.
[1] http://en.wikipedia.org/wiki/Digest_access_authentication [2] http://en.wikipedia.org/wiki/Secure_Remote_Password_protocol
passwd: yourpassword
quoted-passwd: "yourpassword"I'm currently working with an unnamed credit API and all calls, regardless of status, return 200. Options can be strung together in single GET parameter. All calls resolve to a single URL and the method is chosen get a GET param (which variant of that method is yet another param). And a person's information can be passed in via any slew GET params, or as POST'd XML.
Its a mess... anyone up for a Stripe of the credit/authentication world?
The kicker is that they declined permission to open source my wrapper because "the API is proprietary."
- Use HTTP verbs: GET to retrieve one or more objects, POST to create a new object, PUT to update an existing object, DELETE to remove an object.
- Address objects by collection and by individual object: /users/#{user_id} is a specific user you can PUT, GET or DELETE. /users is where you POST to in order to create a new user.
- Use HTTP codes to return the result back to the client (ie HTTP 200 when you get an object, 201 when you successfully POST an object, 404 when you try to GET/UPDATE/DELETE an object that doesn't exist).
- I find it good form to return objects in JSON with the type of object at the top of the data structure, ie {:users => [ #array of users here ]} or {:user => { #single user }}
- Use OAuth or some sort of token system for authenticating the calls, don't use HTTP Auth.
You can get pretty anal about things but if you follow the above you'll have a cleaner API than 90% of the API's out there.
You forgot PUT to create a new object and POST to update an existing object.
> Address objects by collection and by individual object: /users/#{user_id} is a specific user
As far as REST is concerned, that doesn't really matter. The important thing is that the client doesn't generate the "/users/#{user_id}" URLs itself, but rather selects a URL from those it has been told about.
> I find it good form to return objects in JSON with the type of object at the top of the data structure, ie {:users => [ #array of users here ]} or {:user => { #single user }}
Apart from the fact that doing this allows for gracefully adding extra information, there's also an important security reason for doing this rather than having an array at the root: http://haacked.com/archive/2008/11/20/anatomy-of-a-subtle-js...
> Use OAuth or some sort of token system for authenticating the calls, don't use HTTP Auth
Why not do both? http://oauth.net/core/1.0/#auth_header
Sadly, I can imagine how they got to this point. They were tasked to create an API and they did it with the knowledge and tools that they had.
But. Their homegrown text serialization format is pretty wack. They couldn't just use CSV?
Luckily there are some constructive comments in this post which point out some valid concerns. This is a good chance to share opinions and learn a bit. Unfortunately there are a lot of developers out there who are tasked with projects and have no one in-house with experience. I was in that boat and it took developing and maintaining a lot of bad APIs to learn what made a good one.
I agree that maybe it's not necessary to carry over that cultural legacy to the internet, although so far it's worked pretty well.
If social pressure (public shaming) isn't the answer then what kind of pressure should be used? Public shaming is a pretty civilized way to enforce rules when you have no top down control that dictates those rules, I can't think of any alternatives that would be less harsh.
Public shaming can be a good idea, especially on people who are knowingly breaking the rules and holding back progress (See IE6).
However, I'm just glad people are creating APIs, especially because basically no one gets REST right anyways (Hint: if you can't click around your api in firefox with the JSONView extension, it's probably not RESTful).
Off the top of my head:
- API method urls are all the same .aspx, regardless of method used
- All calls are sent as GET (ignoring the whole point of HTTP methods in REST)
- No HTTP method codes as responses for automated parsing
- Custom authentication putting credentials in URL or in headers instead of relying on proven HTTP auth schemes we've been using for years (basic, form, etc)
- Return format is done with get arguments instead of HTTP content negotiation in headers (not so bad)
They've completely missed the boat here.
Using a token to sign requests is certainly a good idea, but you probably don't want to pass the token around using a cookie, even over HTTPS. Many browsers don't enforce good cookie security and will transmit cookies in the clear if an attacker can redirect the browser to a non-HTTPS URL on the same domain.
Thanks!
(but please correct me if it isn't a good example)
Isn't it more the fact that they spend all of their efforts documenting the URI structure and very little on the media types used - which is very unRESTful.
- API method urls are all the same .aspx, regardless
of method used
Extensions are meaningless. That is why we have Accept/Content-type headers. The HTTP spec even explicitly says to not use extensions to relay content information between client and server, from what I recall. - All calls are sent as GET (ignoring the whole point
of HTTP methods in REST)
I haven't read Fielding's dissertation in a while, but using the "Coles Notes" version from Wikipedia, I do not actually see using the verbs a requirement for REST. I do not recall REST even requiring HTTP. You can use any protocol you want, so long as it conforms to the principles.RESTful, on the other hand, does specify the use of HTTP verbs, but the site in question makes no mention of being RESTful.
By not using the proper verbs, the site does appear to violate the caching rules of REST though, I'll give you that.
- Custom authentication putting credentials in URL
or in headers (neither of which are encrypted over
https)
The entire https payload is encrypted, headers and all. REST says nothing about how authentication should be implemented. - Return format is done with get arguments instead
of HTTP content negotiation in headers (not so bad)
This falls under the same as using extensions. Though I will agree with you that it is a reasonable compromise in some cases, such as using a browser where you can't reasonably set your own headers.This company, along with others (e.g.: Flickr and their infamous "REST" API), are responsible for turning "REST" into yet another buzzword.
Shu - you don't know what you're doing (people here asking for help);
Ha - you closely follow the rules (people here complaining);
Ri - you understand the topic sufficiently to adapt and respond as necessary (people for whom this is really not a big deal).
[Actually, now that I read the Wikipedia page for that, it's not quite right - apologies to any martial arts people out there http://en.wikipedia.org/wiki/Shuhari]
RESTful APIs usually represent CRUD operations, each of these letter can be beautifully mapped to request types: - CREATE -> POST - READ -> GET - UPDATE -> PUT - DELETE -> DELETE
Second point: {error: false, value: actual_data} If we have an error variable, what is the HTTP error code then good for? Normal Web Servers use the HTTP error codes for a reason. Besides using standard webframework a json containing only "actual_data" means less code, less errors and so forth...
Third point: URLs should represent the hierarchy: GET /users/43/bookmarks/32442?... is much more beautiful and straight-forward to work with than /api_handler.exe?user_id=43&bookmark_id=32442&operation=get...
Regarding authentication: use a secret API key, that's simple and secure. Everybody does that, from small services to multi million user services like facebook.
As a hint: read the dissertation of Roy Fielding who "invented" REST. In my opinion REST means to exploit HTTP as far as possible instead of using any custom conventions.
http://greenbytes.de/tech/webdav/draft-dusseault-http-patch-...
If they were to change the auth method to use a secret key for signing requests and just call it an RPC API then I don't think there would be a problem.
- Doesn't use HTTP verbs (GET, PUT, POST, DELETE) to retrieve, modify, create and remove objects.
- Doesn't use a restful url path to action on (ie. GET /users/#{user_id} to get a specific user by id, POST /users to create a new user).
- Doesn't use HTTP codes to return results (ie. HTTP 200 for normal operations, 201 for created, 40x for error conditions).
That is specifically one thing which does not apply and is completely irrelevant. Good-looking URLs have nothing to do with REST.
The rest, yes.
He asked it in the context of API restfullness, as is pretty clear from the second phrase.
I was just talking about the restful part of the thing, your urls can look like this:
http://example.org/?wubwub=9cec6c6faef8f7ccefe1bbf409368f1e3d0fa626
and still be part of a completely and perfectly restful service.1. HTTP methods are improperly used. The documentation states "All API calls should be sent as a GET HTTP request". GET is used to retrieve resources, and should be cacheable. It should NEVER change the state of a resource. Yet we see in that it's used to delete, create, and modify resources! Examples:
GET https://subdomain.sharefile.com/rest/folder.aspx?op=create...
GET https://subdomain.sharefile.com/rest/file.aspx?op=delete
These should be as follows:
POST https://subdomain.sharefile.com/rest/folder.aspx
DELETE https://subdomain.sharefile.com/rest/file.aspx
HTTP verbs have specific semantics that allow for intelligent, scalable architecture.
2. Individual resources are not identified by an unchangeable URI. Let's say I want a folder:
GET https://subdomain.sharefile.com/rest/folder.aspx?op=get&...
Nope. This is more RESTful:
GET https://subdomain.sharefile.com/rest/folder/123
Two things to note: 1) no 'aspx' bullshit, that's implementation and shouldn't be visible to the user, and 2) the ID is in the URI itself as opposed to the params. URIs should seldom, if ever, change. Parameter names may change, URIs shouldn't.
3. Violation of content type retrieval. Let's say I want a folder as XML:
GET https://subdomain.sharefile.com/rest/folder.aspx?op=get&...
Wrong. This is the request I should be sending as a curl:
curl https://subdomain.sharefile.com/rest/folder/123 -H 'Accept: text/xml'
Note that I'm using the Accept header. This allows for flexibility in accessing my resource. With any luck, the server will give me an XML file, and its content type should be "text/vnd.sharefile+xml" (vendor specific content types are best).
4. No link relations in the body. I'm mostly assuming this because I haven't seen sample response bodies, but almost no one does this even though is a crucial aspect of a RESTful service. RESTful services describe the relationship between resources. Suppose I have an ordered collection of files. When retrieving a file, it should describe its relationship in this ordered collection. This is done either via the LINK header, or some sort of scheme in your response body (e.g. JSON Schema describes 'link' attributes).
A pretty good RESTful API is Github's (see http://developer.github.com/). If you'd like to learn more, I highly recommend "REST in Practice" by Webber, et al (http://www.amazon.ca/REST-Practice-Hypermedia-Systems-Archit...).
http://roy.gbiv.com/untangled/2008/rest-apis-must-be-hyperte...
A semi-related question - this is something I've wondered for a while - does the XMLHttpRequest single origin policy actually do anything for security? What kind of malicious resource might you try to fetch with an ajax call that can't just be wrapped in a jsonp callback?
- Linked to that, a significant part of the documentation is about URLs to hit instead of being about document types
* Misuse of HTTP verbs, a lack of use (by going through POST for everything) would be bad enough but they go through GET which should be 1. safe (should not alter the server's visible state) and 2. idempotent (should be callable n times with the same parameters returning the same result, as long as the server's state was not changed). This API deletes things through GET.
* Lack of use of HTTP headers, instead of using the Accept header the client has to use an informally defined query parameter
* Lack of descriptive content types (linked to previous point)
* Complete absence of use of HTTP status codes (standard or custom), the API returns a status of 200 in all cases
This is pure and simple (and terrible, since it's over GET) RPC tunneled through HTTP.
You guys are spoiled. In my day we had to guess what parameters to send.
In the past I've simply extended the uri to indicate additional methods like, POST /user/foobar/grant. I wonder whether putting those methods as additional HTTP verbs would be better.
Edit: extending HTTP verbs might not work too well with some firewalls/NAT that filter out non-standard verb requests.
Take a look at http://www.recessframework.org/page/towards-restful-php-5-ba... to learn how to get started. Also, the dude who wrote that has a framework that is built around REST. You don't really need it though.