What is the maximum length of a URL?
stackoverflow.com
stackoverflow.com
And, of course, insert here a long-winded rant about how broken your data model is if you need more than about 1000 chars. ;)
In Squid it's 4K by default: http://www.squid-cache.org/mail-archive/squid-users/200208/0...
I see no reference to a maximum URL size in Varnish, and a cursory glance through the source code is not revealing a hard-coded size. I'm not shocked. PHK is well-known for writing good code and good code generally doesn't have a lot of magic numbers.
I really can't stand contrafactual arguing from first principles. I gave you a legitimate non-aesthetic reason and you came back with a softly wilted notion conjured whole out of your imagination. Shut up already.
> I think if you have nearly 1KB of GET data, something is definitely wrong.
The only point anyone is trying to make is that this is incorrect; retrieving multiple records is a perfectly good use case for having a URL that is more than 1KB (though there may be pragmatic reasons to avoid it).
Is anyone seriously going to trip on your interface if you say something like "This API is essentially RESTful, but everything goes through POST and the CRUD operation is specified by the "op" POST parameter"?
The question I would have is whether it is ok to put your GET parameters in the body, which is generally frowned upon, but might be acceptable in this extreme circumstance.
1) Write small modules 2) Concatenate scripts
If you try to automate module concatenation then you'll get URLs like http://example.com/combo?foo.js&bar.js&baz.js, which combined with "write small modules" can mean a long list of small modules that easily reaches 2000 characters.
back in 90-something i had to research that for url and cookies.
msdn specs said something like "[url|cookies] should be at least [some size]"
mozilla's and w3c specs said: "[url|cookies] should be at most [some size]"
for everything. url lenght. cookie lenght. number of cookies per domain. number of cookies per subdomain... it was like that for ALL items. one said 'at least' the other 'at most', and for added fun, all the values where exactly the same on both.
I presume that the people writing HTTP servers and clients usually found it easier to allocate a fixed-length char array.
> Each label may contain up to 63 characters. The full domain name may not exceed a total length of 253 characters in its external dotted-label specification.
In this context "label" means a segment of the name, so in "www.example.com." "www", "example", and "com" are all labels.
The next sentence:
> In practice, some domain registries may have shorter limits.
alludes to your problem somewhat ;)
- Unique and unforgeable references for capability-based security (e.g. the Waterken Server)
- Serialized continuations, or unforgeable references to server-side continuations (as in continuation-based web servers)
The data URI scheme does not really apply here since it's never sent to any server. If a browser understands data URIs, it should logically also allow such long URLs.
Specifying what item(s) should be targeted in a pool of possibly millions, with endless possible combinations is bound to require some kind of precise pointer.
Here is an example URL I use: https://bacnethelp.com/vis/overview/KHs6cHJvamVjdC1pZCAiNTA1....
What I could do, however, is use some kind of shortening url scheme, kinda like google maps: http://goo.gl/maps/3uP8y.
I'm still uncertain about which way is better.
EDIT: Of course I mean a local shortening url, pointing to my own databases.
seems like it would be better stored on the server in redis or something (or, at least if leaving it in the URL, a more compact deduplicated format might be worthwhile)
({:project-id {:conditions "true); DELETE FROM projects WHERE (true"}})But nice reflex! ;-)
However the duplication overhead would only be really paying off with a large number of objects.
By the redis reference, I suppose you refer to a uniquely created key each time a user request a possible combination. Something like /short-url/abcd, where abcd would be a key matching {:project-id "505a125e44ae42e05a750c97"... ?
That's what I was thinking when talking about a shortening url scheme. It requires more work, but the final URL would indeed be more sexy.
Thank you for the input, I appreciate it!
I was probably stuck in a weird mindset.
Since you're using base 64 there, let's think for a minute. How many characters would you need to uniquely identify over a million objects? log_64(1,000,000) is about 3.3. With 4 characters, you could represent over 16 million objects. If you just store all of the objects that you need to reference along with an incrementing primary key, you wouldn't have to use more than 4 characters until you had more than 16 million objects in your database.
Have a billion objects? That's just five characters. Still not enough? With 7 characters, you could index more than 4 trillion.
But let's say that you can't actually keep a single database, with an incrementing primary key. You have multiple independent processes or people generating objects that need identifiers that will always be stable, you can't rely on manually picked names, and so on. So just use a secure hash: a SHA-2 or SHA-3 hash of the objects. If you use a 256 bit secure hash (44 characters in Base64, including the padding), and had 500 octillion items in your set, you would have about one in a quintillion chance of having an accidental collision. I'll give you a hint; you are never going to have that many items in your data set.
Now, you might object "what if SHA-2 is broken". Well, that may happen, though it's fairly unlikely. Most of the ways of breaking a secure hash involve making it a few orders of magnitude easier to compute a collision. But at 256 bits, you have a substantial safety margin; it would have to be pretty thoroughly broken before anyone would be able to find meaningful collisions. Heck, Git uses SHA-1 still, which uses 160 bit hashes, and is much closer to being broken.
Anyhow, the point of all of this is that a URL is supposed to be an identifier. It doesn't take that many characters to create an identifier that could uniquely identify each quark in the whole universe. You absolutely don't need long URLs to guarantee uniqueness; if your URLs are long, it's because you're including a lot of redundant information in the URL, or you are actually trying to store a description of the object in the URL, rather than an identifier.
If my goal was to get the smallest possible URL, you would be correct (well, you still are...) However, for the same reason people prefers a website named "ycombinator" over "zgrrc", even if the url would be shorter, I don't mind not being concise.
If I can check my url and without any database see that it's project X, device Y, object Z, it makes debugging easier.
But you are absolutely right: there are ways to make short, unique urls. I just don't want to use them.
People get confused about this all the time, but a URI is not necessarily a URL. A URL is always a URI. A URN is also a URI and might be a URL as well.[0]
Interesting side note is how, in 1997, with the original uuid: draft[1] I goofed and didn't recognize uuid: was really a URN. This was corrected in RFC 4122 in 2005[2].
[0] http://en.wikipedia.org/wiki/Uniform_resource_identifier [1] http://tools.ietf.org/html/draft-kindel-uuid-uri-00 [2] http://www.ietf.org/rfc/rfc4122.txt
See: http://tools.ietf.org/html/rfc2397
Entitled as 'The "data" URL scheme'.
Either way, it doesn't matter, because the OP has provided such limited details as to why this is required knowledge that the question is as irrelevant as the accepted answer.
I don't agree that a long URLs necessarily indicates a broken design. I've done it and configure my web server (Tomcat) to accept longer GET requests (so +1 to 'asiermarques' here, who pointed out that it depends on how the web server is configured too).
There are services out there, like Google Charts if I'm not mistaken (but I may be mistaking it with something else) which, by the way they work, forces you to create quite long URLs. It generates a graph on the fly and nothing is modified on the server, so a GET is used, not a POST (which, IMHO, makes sense).
From the thread I seem to understand that the Data URI scheme implies that company should start to upgrade or replace their broken load balancer that "mysteriously" truncate URLs and cause all manner of weird bugs ; )
. . . and all proxies that may be in-between.
I've debugged more than one client issue where people were doing something really weird over port 80.
REST never really considered requests that don't involve resources..