Java's URL.equals() Performs DNS Resolution
docs.oracle.com
docs.oracle.com
---
To quote a JDK bug ticket: https://bugs.java.com/bugdatabase/view_bug.do?bug_id=4434494
> We are well aware of the problem with URL.equals and URL.hashCode. The cause of the problem is due to the existing spec and implementation, where we will try to compare the two URLs by resolving the host IP addresses, instead of just doing a string comparison. Because hashCode has to maintain certain relationships with equals, namely if two objects are equal, they should have the same hashCode, the implementation of hashCode also tries to resolve the host in the URL into an IP address. As a result, we are facing problems with http virtual hosting, as described in the Description part, and performance hit due to DNS name resolutions.
> Unfortunately, changing the behavior now would break backward compatibility in a serious way, plus Java Security mechanism depends on it in some parts of the implementation. We can't change it now.
> However, to address URI parsing in general, we introduced a new class called URI in Merlin (jdk1.4). People are encouraged to use URI for parsing and URI comparison, and leave URL class for accessing the URI itself, getting at the protocol handler, interacting with the protocol etc. So, at present, we don't plan on changing the URL.equals/hashCode behavior and we will leave the bug open until Tiger, when we re-investigate our options.
- equals() or hashCode() called on java.net.URL object
- Map or Set may contain java.net.URL objects
https://www.jetbrains.com/help/idea/list-of-java-inspections...
It was already ancient advice then, but I kept finding myself having to teach it to a new group of people, usually after some incident had happened (because they won't listen until something has already broken)
Hell, even Rails maintainers rejected a PR I opened that tried to get rails link helpers to work with URI objects.
If I'm wrong: Why not invent a 'protection level' where this were a feasible thing... even if just to be able to remove old APIs and really force people to stop just adding @SuppressWarnings-deprecation all over the place. That's not a solution to anything.
It's there for backwards compatibility, which is not nonsense. It keeps the old code running.
A clear warning at top of URL, URL(String) would be good to make it recognizable in most IDEs when using it.
Also sounds like something a good static code analyzer should be able to detect.
Uhmm... yikes. Why are they resolving anything? A URL is a string, this should just be doing a string comparison. All of the parts of the URL object are strings or ints, so at worst they should just be comparing all of those individually, not resolving domains and comparing IP addresses. That makes no sense at all.
No, that doesn't make "perfect sense." Two different URLs are different, they cannot be equal if they're different, and a library called "URL" shouldn't use voodoo-magic to say that two different URLs are equal when they're not.
If it was e.g. Net.DNS.Equal(Uri, Uri), perhaps. But even then it is ambiguous as DNS has multiple record types, not just "A" records. So is it pulling all records (A, AAAA, MX, etc) and comparing them all? Or just arbitrary comparing one?
But as it stands, it is a URL library that is ignoring the URL part of the URL and using resolution to decide equality instead. It is nonsensical. And even if they did resolve identically they may not be treated identically by routing or endpoints.
$ dig amazon.com
...
;; ANSWER SECTION:
amazon.com. 60 IN A 205.251.242.103
amazon.com. 60 IN A 176.32.103.205
amazon.com. 60 IN A 176.32.98.166
Something resolving "amazon.com" is welcome to resolve it to any of those three addresses. Code that tries to "resolve" these IPs is at the very least coming perilously close to deciding that amazon.com != amazon.com, nondeterministically, if the underlying resolution call changes its behavior.The further obvious change of trying to compare the whole record just gets worse; what is the answer if domain 1 resolves to IPs (V, X, Y) and domain 2 resolves to (X, Y, Z)?
Oh, and let's not forget, DNS can be different depending on your geographical location, so in the US two domains may happen to resolve to the same IP but in Europe they may be different. What's the use of a URL equality operator that changes behavior based on where the user is? Whatever the use of such a thing may be (I mean, yeah, I get the internet contrarian impulse, yeah, I can construct some bizarre situation in which it is useful), it is certainly less useful than a simpler operator.
Basically, the entire idea is just fundamentally flawed and shouldn't be used. I make no claim to have exhaustively enumerated all the ways in which this is a bad idea, merely added to the pile, and demonstrated sufficient evidence for my claim that it shouldn't be used.
Edit: They know, see yzmtf2008: https://news.ycombinator.com/item?id=21766138 and consider my post here a lightly educational post on further reasons why it's a bad idea. If you only take one thing away, remember, a function that takes a hostname and yields an IP doesn't have very many useful properties beyond just that it yields an IP or failure. You can count on very little else about that IP. It should be treated as opaque and not compared, stored (other than logging), etc.
And yes, I know that overriding .equals() does not change the behavior of == in Java. But I would still consider .equals() to be the "default" equality operator for objects.
It's extremely common now, but it wasn't possible yet when they made this design decision.
(Edit: And really, virtual hosts are an extremely weird HTTP feature. What other internet protocol cares about what domain name you used to establish a connection?)
kerberos? (btw. this is older than 1995) or smtp (little bit different than the http version, tough)
I don’t agree with what this is doing... but you can’t say it makes no sense at all.
foo.bar/baz would be said to be equivalent to bar.wack/baz if they shared the same IP address.
If there is some sense to be made from that, it eludes me.
In this case, Java has thrown away the critical distinction between the two.
* A URL is a parameter to a function returning a resource.
* A URL is a string, conforming to a certain format.
Neither of these specifies that the resource it points to is always the same resource, nor that it's always possible to resolve it. That's by design; hostnames change. The Web is not permanent and that's why we have eg. HTTP 404.
Network topologies also change.
However, the default equals() comparison between two objects is supposed to compare those two objects, not the current topology of the Internet.
There is no way this behaviour ever made sense, nor any way it ever could make sense. It's moronic, through and through, for any language or library, to implement default equality in this way.
If you want to implement a comparer which does stuff like this and accepts a hostname resolver as a dependency, great. But there is simply no excuse for this kind of stupidity in a default dependency-less implementation.
It also means that if one URL changes, then they could be equal sometimes and not others.
It sends a different Host HTTP header; servers send different responses based on that header ~100% of the time.
They even seem to know it makes no sense.
... and what happens when you try to compare a URL which contains a nonexistent hostname, like http://asdfghjkl.example.com/ ? Does that compare as equal to all other URLs with unresolvable hostnames?
if (internalURLs.contains(submittedURL)) {
check. Then change your DNS records to point to some other server, once your domain is in their database and they assume it to have already been validated as internal.This seems absolutely ripe for abuse, as now there's a nice string you can search for in GitHub to see where that's used as a security feature. I'm imagining things like "if URL("http://example.com/some/path").equals(URL(checkedUrl)) { return AllowEditRights }", and checkedUrl = "http://wiki.example.com/some/path" or similar.
Java is as good as Windows in the sense, that its standard library and set of APIs is very stable and supports a lot of legacy software. I doubt there's a real need for URL class in the new code by now, given that URI class was introduced in JDK 1.4 almost 18 years ago. There's plenty of dependencies though, so URL will probably stay in the core library forever, but URI represents a superset for URLs, has reasonable implementation of equals/hashCode and is sufficient for majority of the uses.
I think https://imgs.xkcd.com/comics/ten_thousand_2x.png applies here, just in a more restricted programmer-y sense..
URL class is used to establish actual connection to a resource.
But, just like Vector remains in use in Swing, URL remains in use in the core java library (looking at you, classloader.getresource), so it's easy for me to have made that mistake.
then again i dont use it often. mostly come in contact with it when using ClassLoader API.,
For example, because the "http" scheme makes use of an authority
component, has a default port of "80", and defines an empty path to be
equivalent to "/", the following four URIs are equivalent:
http://example.com
http://example.com/
http://example.com:/
http://example.com:80/ scala> new URL("http://localhost/foo/bar/baz").equals(new URL("http://localhost/foo/../foo/bar/baz"))
res0: Boolean = false
It's just funny to me that they went to the effort to do a full network request to see if the hostnames resolve to the same IP, but didn't bother to normalize paths. doppio.eng.sun.com
it would be referred to simply as "doppio" from within the engineering ("eng.sun.com") domain, or possibly as "doppio.eng" from other domains inside of Sun. It was fairly rare to use FQDNs inside of Sun to refer to other hosts inside of Sun. Thus, the following URLs all referred to the same resource: http://doppio/foo.html
http://doppio.eng/foo.html
http://doppio.eng.sun.com/foo.html
It's a plausible point of view that URL.equals() should report true for any two of the above URLs. (That doesn't mean that I think it was a good idea, though.)So, no, this is not remotely Ok at all.
Regardless, given all Java's warts, this doesn't even make my top 10.
My company Datastreamer has been around for a decade. We provide crawl data (usually a massive amount of crawl data, north of 300GB per day) to our customers.
... so we have a LOT of real-world experience pushing data to customers in production over long time periods.
Here's what we've learned.
Networking libraries around HTTP are and have been fundamentally broken for a long time and they're broken in pathological ways that you don't realize until years later in production.
DNS caching is a good one. A lot of systems do infinite DNS caching. Java, until at least Java 8, does infinite DNS caching.
Some do infinite HTTP timeout. Timeouts are awesome. You should use a timeout. Without a timeout if the network breaks your code just locks up.
Some libs provide no API to change TCP buffer sizes (which you have to do at the kernel level).
So about 5 years ago we took a harsh stance. NO CUSTOM CLIENTS.
We have a streaming firehose client that we implemented from the ground up to do everything properly. The API is literally that we just stream JSON files to disk.
It's a docker container now so not too hard to deploy.
Your job is to just to listen to the disk and wait for new files to be written. We do a move from a tmpdir to the final dir so the entire file is written and you don't have to worry about partial reads.
About 80% of our customers love it. The other 20% of customers seem to initially hate it and we have to explain to their CTO or senior architect that, no, you DO NOT want to implement this from the ground up.
What happens is that it works immediately, but then 18 months in it will break pathologically and everyone running it has moved on or it's in some datacenter that no one has access too.
This causes us to break our SLAs and means we have upset customers.
This decision by far was one of the best decisions I've ever made and has really helped our growth and stability over time.
It's really really really nice to keep customers for 5-10 years. They're happy and you get steady checks and predictable growth.
When I see an infinite timeout, I suggest the developers change it to 30 years. They inevitably respond "That's crazy!" and prove my point.
equals() and hashCode() are probably one of the weakest points of Java. While it seems like an obvious candidate for a contract, the issue has always been that one person ends up defining equality for everyone, when often different usecases will warrant different definitions of equality. Are objects equal if they have the same identity? If they have the same data? If they resolve to the same thing? It's easy enough to leave them unimplemented but the issue then is a lack of standard library support for providing custom hash and equality functions for Maps and Sets.
It's less an issue so I get why it's been backlogged, but it's one that I'd love to see worked out.
I'm fine with not breaking old jars, but they could have made newer javac's by default die with a notice "use -legacy to build this". Same thing with Date, Vector, all the broken thread semantics, you name it.
These classes should have been relegated to a handful of ancient jars, they shouldn't keep popping up in modern Java because newbies have to learn a whole host of classes they aren't supposed to use.
Ultimately I think it’s a somewhat natural part of language evolution that they accumulate cruft over time, because the costs of breaking backwards compatibility are too great. And after a few decades or so, new languages come out that are (currently) much cleaner, and they take over until the cycle repeats. i.e. “cleaner Java” will not be a new version of Java, but a new language solving similar problems (Go?).
Design Patterns (Elements of Reusable Object-Oriented Software), Erich Gamma, Richard Helm, Ralph Johnson, John Vlissides, 1995]
Same year Java was invented. Hard to imagine Java inspired it.
It hasn't been deprecated, it's still in use in the standard library[1][2][3] and they don't offer additional classes that accept URIs.
> Easy to be a harsh judge now but back when java was first developed...
They've made no effort over 11 major versions to diminish its use beyond a brief note that encoding is easier with URIs. There's not even a warning on the equals or hashCode method, let alone the class documentation, it just quietly mentions that it's resolving a name. There's no way people are finding out that you should avoid URL except through lore.
[1]: https://docs.oracle.com/en/java/javase/12/docs/api/java.base...
[2]: https://docs.oracle.com/en/java/javase/12/docs/api/java.base...
[3]: https://docs.oracle.com/en/java/javase/12/docs/api/java.base...
https://github.com/biomics/icef/blob/d69f9be9b1f773598b47de5...
Query I used: https://github.com/search?q=%22Map%3CURL%2C%3E%22+language%3...
https://github.com/apache/dubbo/issues/5462
https://github.com/DoctorD1501/JAVMovieScraper/issues/310
https://github.com/MovingBlocks/Terasology
focusing on repos that actually seem to be used. Someone should write a tool to scan this
* DNS will fail even if it is implemented correctly on your clients' network
* It probably isn't implemented correctly on your clients' network
* Your software is probably doing more DNS queries than you think
This seems like a particularly unfortunate example but things like this are not uncommon. Doing any kind of RPC, even just to a server on your local machine? Half the time your library is doing DNS queries under the hood for no good reason. Or performing reverse DNS queries just to display a hostname in a log file.
Its easy to accidentally trigger a lookup. If it happens in a loop, a query that should take .2 seconds now takes 30 minutes.
“ This method assumes that path is a directory if it ends with a slash. If path does not end with a slash, the method examines the file system to determine if path is a file or a directory.”
People often overlook this and then wonder why their app stutters randomly when the UI thread gets blocked on this NSURL ctor :^)
Just avoid it entirely.
> The recommended way to manage the encoding and decoding of URLs is to use URI, and to convert between these two classes using toURI() and URI.toURL().
and at that point you might as well just use URI. The resource resolution functionality is basically the only thing URL offers over URI.
I don't think it's unreasonable to ask that an API be deprecated when core functionality, like comparing the thing, is broken. The stdlib docs should provide best practices to new developers and clearly mark old, broken APIs as deprecated.
Same with Date[1]; you have to make a defensive copy every time you accept or return it or you risk your internal state being changed. You're not going to pick this up from the docs unless you have a. learned about the problems with mutability, and b. carefully scrutinize the docs and notice, oh, hey there are all these setter methods.
What's nuts is the APIs in Date that they did deprecate aren't even broken, they're just not very general! It's the Date object itself that is hilariously dangerous to use.
[1]: https://docs.oracle.com/javase/8/docs/api/java/util/Date.htm...
That being said, URL.equals() is a terribly opaque and non-obvious method signature for performing DNS-based comparisons of hostnames. It lacks any indication that calling it involves network IO.
at java.net.Inet6AddressImpl.lookupAllHostAddr(Native Method)
at java.net.InetAddress$2.lookupAllHostAddr(InetAddress.java:928)
at java.net.InetAddress.getAddressesFromNameService(InetAddress.java:1323)
at java.net.InetAddress.getLocalHost(InetAddress.java:1500)
- locked <0x00000000800a1578> (a java.lang.Object)
at sun.font.FcFontConfiguration.getFcInfoFile(FcFontConfiguration.java:352)
at sun.font.FcFontConfiguration.readFcInfo(FcFontConfiguration.java:425)
at sun.font.FcFontConfiguration.init(FcFontConfiguration.java:94)
- locked <0x00000000d5af3c58> (a sun.font.FcFontConfiguration)
at sun.font.FcFontConfiguration.<init>(FcFontConfiguration.java:76)
at sun.awt.X11FontManager.createFontConfiguration(X11FontManager.java:768)
at sun.font.SunFontManager$2.run(SunFontManager.java:431)
at java.security.AccessController.doPrivileged(Native Method)
at sun.font.SunFontManager.<init>(SunFontManager.java:376)
at sun.awt.FcFontManager.<init>(FcFontManager.java:35)
at sun.awt.X11FontManager.<init>(X11FontManager.java:57)
at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)
at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)
at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)
at java.lang.reflect.Constructor.newInstance(Constructor.java:423)
at java.lang.Class.newInstance(Class.java:442)
at sun.font.FontManagerFactory$1.run(FontManagerFactory.java:83)
at java.security.AccessController.doPrivileged(Native Method)
at sun.font.FontManagerFactory.getInstance(FontManagerFactory.java:74)
- locked <0x00000000d5abbb00> (a java.lang.Class for sun.font.FontManagerFactory)
at sun.font.SunFontManager.getInstance(SunFontManager.java:250)
at sun.font.FontDesignMetrics.getMetrics(FontDesignMetrics.java:264)
at sun.swing.SwingUtilities2.getFontMetrics(SwingUtilities2.java:1113)
at javax.swing.JComponent.getFontMetrics(JComponent.java:1626)
at javax.swing.plaf.basic.BasicLabelUI.getPreferredSize(BasicLabelUI.java:227)
at javax.swing.JComponent.getPreferredSize(JComponent.java:1662)
Use URI
The semantics don't even seem desirable in 2019, let alone the performance characteristics.
Some quick googling suggests people are using java.net.URI instead to bypass this poor design.
> Two URL objects are equal if they have the same protocol, reference equivalent hosts, have the same port number on the host, and the same file and fragment of the file.
> Two hosts are considered equivalent if both host names can be resolved into the same IP addresses; else if either host name can't be resolved, the host names must be equal without regard to case; or both host names equal to null.
> Since hosts comparison requires name resolution, this operation is a blocking operation.
So it's not an implementation bug, it's a requirements bug. Now why on Earth they thought this was a reasonable requirement for an equals() method, that's a fair question.
That seems... less than useful.
> Note: The defined behavior for equals is known to be inconsistent with virtual hosting in HTTP.
So, yes.