Buying a single character domain – and 3 character FQDN – for £15
shkspr.mobi
shkspr.mobi
I’ll pick on the FQDN part first, because it’s the (only) part that is unequivocally wrong. FQDN is a very specific technical term in the domain name system. A fully qualified domain name includes a trailing dot, so Ⅷ.fi. would be four characters, even if Ⅷ and fi were the actual labels. But they’re not: DNS is strictly ASCII-only, so this normalisation is happening at a higher level (as the OP notes in another response here, tools are applying IDNA2008, per RFC5895). The FQDN is viii.fi., which is eight characters long.
Next I deny the claim that it’s a single-character domain. Perhaps I’m getting petty here, but even if people do colloquially speak of example.com as a six-letter domain, counting only the label at the level you register the domain, so that I would grudgingly allow Ⅷ.fi to be considered a single character domain (per proletariat vernacular), the domain name that was “bought” was not that, but viii.fi, which is a four-letter domain. Hair splitting is fun.
But my pettiness knows no bounds. Domain names aren’t bought, they’re registered for such-and-such an amount per annum. And I bet it wasn’t exactly £15.00 that was paid.
:-)
———
I went thinking about other TLDs that would work, and ℡ (TEL → .tel) and № (No → .no) occurred to me off the top of my head. Haven’t seen a .tel domain name in yonks. I never did quite see the point of .tel.
I also have a list of tld which can be shortened by this process. https://shkspr.mobi/blog/2018/11/domain-hacks-with-unusual-u...
And, yes, £15.30. But let's not quibble :-)
If we're being pedantic, "example" has seven letters!
:-)
:-) indeed.
The term is ambiguous, and many uses don't include the trailing dot. See IETF RFC 8499 section 2.
> DNS is strictly ASCII-only
DNS is not strictly ascii. At the protocol level labels are octet strings. You may be thinking of the LDH convention.
• Hmm, I wasn’t aware of that admission of ambiguity. Allowing that fuzziness in terminology puzzles me, because if you don’t have the trailing dot, then what you have fundamentally isn’t fully qualified. Sure, some resolvers may treat it as though it were, but some won’t, and are perfectly correct not to. (I haven’t the foggiest idea what the balance of implementation behaviours might be. I know enough DNS to be dangerous, but I don’t live and breathe the stuff.)
• I should have said DNS hostnames are ASCII-only, which I believe to be true. (And yeah, this still depends on conventional rather than rigorously defined terminology.)
As far as client-side tools... no one ever uses the trailing dot and that can lead to some interesting situations if you have a search path set and use a resolver that resolves wildcard. Use of the search path at all is fairly unusual for "client-ish" setups though. You could also view the common non-use of the trailing dot as one of the causes of the ICANN recommendation against top-level domains resolving, as the search-vs-hostname-vs-fqdn ambiguity would be less common (but still present) if people commonly used the trailing dot.
:)
Technically correct is best correct.
>Hair splitting is fun.
Surely it just can't get any better than this. We're done here.
>But my pettiness knows no bounds.
I'm dead. Wrap me up.
I suppose this is inevitable when you tasked with representing literally every symbol in existence. You couldn't pay me enough to touch this problem with a ten foot pole (this and text rendering).
Writing seems simple — children do it routinely — but like a biological system it evolved over millennia in a ton of different directions. It’s coupled with emotional, practical, and even, yes, moral issues that operate on both deeply personal and social issues. This is hard to capture in software.
Unicode made a couple of hard decisions right up front. I hate them but they were smart and Unicode would not have survived had they not made them. One was round trip with legacy character sets, which meant encoding a lot of redundant characters (English and German “A” have he same code point, but Greek “A” and Russian “A” do not, nor does an “A” that appears in a Japanese code table. Second was abandoning attempts at Han unification, which had its own linguistic, emotional and political issues.
People are complicated and so are their languages so wrestling the whole thing into a tractable system has been worth the effort.
Huh? Han unification happened.
Imagine if Unicode has to start dealing with the temporal change that for example the Olson TZ database [2] has to!
[1] https://shkspr.mobi/blog/2019/06/quirks-and-limitations-of-e...
One example of how this is a huge clusterfuck is that until recently, Windows Notepad opened and saved everything with the Win-1252 encoding scheme (labeled as ASCII in the app). The web, and the other popular OSes, on the other hand, are standardized around UTF-8. So if you download a txt file from the web or OS without a BOM, and you open it in Notepad, you can get characters that looked right in your browser, but not in Notepad.
There are smart algorithms out there that can detect character encoding pretty well, but none of them are perfect (as far as I know).
The Win-1252 default and the fact that most computer users have no idea about character encoding have caused all sorts of headaches for me with the reporting software I work on.
The idea that computers should support cultures, and not the other way around, is pretty recent.
Gruß, stkdump
Um, no. The words were originally written that way. Ä, ö, ü and ß actually developed from ligatures for ae, oe, ue and ss, long before computers were a thing.
I modified your first sentence to make it more generic and applicable to many other things in software.
Barring that, we could have all used UTF-8, but windows really screwed that up, and none of the arguments for 16-bit alignment vs 8-bit alignment for processing really hold water.
Sadly, ws. is not serving right now it seems - I had no idea root ccTLDs could vend an A record but there you go; I guess technically that means the root servers themselves could vend A records, which would let you have the ultimate website at http://./
I’m still waiting for someone to launch a .ux domain so I can grab macint.ux
A more probable course of events would be “tux” being registered as a new gTLD.
https://newgtlds.icann.org/en/applicants/agb/guidebook-full-...
As it looks like that TLD is not currently issued:
There are, of course, exceptions---the policy was put in place after the creation of ccTLDs and does not apply to them, so a few ccTLDs get to break the rule.
You mention this can be used to avoid filters, so I guess this is specifically to trick "string".length?
That's still shorter than V I I I though.
The domain resolved for me in Chrome on macOS though!
I own 0e.vc, which is on GitHub as a general purpose xss domain if you need it. Iirc it does eval(window.name), or location.hash. Whatever works for you. It’s also on the public suffix list which makes it almost like a top level domain for security purposes. So I can have subdomains that can’t ever share cookies :-))
A bit of a tangent, but: how and why?
https://ahreflink.com/domains/two-letter
Seems there are plenty available.
https://tools.ietf.org/id/draft-chapin-rfc2606bis-00.html#ne...
[root@host ~]$ curl -v http://.
* About to connect() to . port 80 (#0)
* Trying 127.0.0.1...
* Connected to . (127.0.0.1) port 80 (#0)
> GET / HTTP/1.1
> User-Agent: curl/7.29.0
> Host: .
> Accept: /
>
< HTTP/1.1 400 Bad Request
< Date: Sat, 15 Aug 2020 15:38:54 GMT
< Server: Apache
< Content-Length: 347
< Connection: close
< Content-Type: text/html; charset=iso-8859-1
<
<!DOCTYPE HTML PUBLIC "-//IETF//DTD HTML 2.0//EN">
<html><head>
<title>400 Bad Request</title>
</head><body>
<h1>Bad Request</h1>
<p>Your browser sent a request that this server could not understand.<br />
</p>
<p>Additionally, a 400 Bad Request
error was encountered while trying to use an ErrorDocument to handle the request.</p>
</body></html>
* Closing connection 0
# cat named.conf.local
zone "." {
type master;
file "/etc/bind/db.root";
};
# cat db.root
$TTL 60
@ IN SOA localhost. root.localhost. (
1 ; Serial
604800 ; Refresh
86400 ; Retry
2419200 ; Expire
604800 ) ; Negative Cache TTL
;
@ IN NS .
@ IN A 209.216.230.240
# dig a .
...
;; QUESTION SECTION:
;. IN A
;; ANSWER SECTION:
. 60 IN A 209.216.230.240
# wget http://./
--2020-08-17 09:38:26-- http://./
Resolving . (.)... 209.216.230.240
Connecting to . (.)|209.216.230.240|:80... connected.
HTTP request sent, awaiting response... 400 Bad Request
2020-08-17 09:38:26 ERROR 400: Bad Request.Doing a dig, there are no A records for . but they do provide NS records obviously, for all the other TLDs to resolve.
Devilish detail is probably in the RFC and likely implementation specific also, as to whether they accept it is a valid request or fail before trying.
dig A .
; <<>> DiG 9.11.5-P4-5.1-Debian <<>> A . ;; global options: +cmd ;; Got answer: ;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 26280 ;; flags: qr rd ra ad; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 0
;; QUESTION SECTION: ;. IN A
;; Query time: 3 msec ;; SERVER: 192.168.1.254#53(192.168.1.254) ;; WHEN: Sat Aug 15 18:19:41 BST 2020 ;; MSG SIZE rcvd: 17
if you want to query a dns server for a top level domain, you have to query the dns that handles it, which in our case is one of the 13 root dns servers.
read more here: https://en.m.wikipedia.org/wiki/DNS_root_zone
I suppose other TLDs could have different rules. But in general you'd want to have a canonicalization step so you don't have two domains that are just different ways of composing the same thing.
On the bright side, even the three character version is unlinkable on facebook -- it just redirects me to http://invalid.invalid/. I'll take that as a win. I still managed to get a pretty cool domain name for around $40, and it was definitely fun to mess around with this idea.
EDIT:
Interestingly, HN has automatically punycoded the URLs. That should be ⑭.㎳ and ⒕㎳
xn--jxb.fi would appear to be unregistered.
(U+2167) (U+002E) (U+FB01) is not a valid domain name.
viii.fi (all ASCII) is registered, and the registrar is gandi. But so what?
I don't get it.
When your browser see the Ⅷ character, it performs the IDNA2008 process listed in RFC5895 to normalise it. The same thing happens in OS tools like dig and ping.
Hope that makes a bit more sense.
This seems to be only useful if you manage to find a website with an XSS flaw and one that also limits the input to 20 characters? Are these situations really common enough to warrant this attack? It all seems rather arbitrary to me.
They shouldn't - but they do.
My brother's old .fi domain was taken by a company with the same name (because they have priority) and they didn't even do anything with it...
I think there might be one or two other domains that have a host as well. But beating two letters is difficult.
The punycode version is xn—x6h.ws
Recycling Symbol for Type-4 Plastics
I don’t use it as much anymore because of the previously mentioned Unicode filtering some sites use, like HN.
Does that count?
That may be a standard thing to do with unicode in domain names, run it through the standard unicode denormalization first? Understanding what browsers are "supposed" to do with unicode in domain names (and URLs generally) is very confusing for me.
I would be curious to learn more about what standards govern how browsers handle unicode in domain names, the history of it, how compliant browsers are, etc. I also don't entirely understand the goal here -- the original `Ⅷ.fi` isn't actually only two bytes in any encoding... what is the value of having something that shows up as two "glyphs" even though it's more bytes and denormalizes to something else with a yet different number of bytes?
Internationalized Domain Names (IDN) FAQ (https://unicode.org/faq/idn.html)
Unicode® Technical Standard #46: UNICODE IDNA COMPATIBILITY PROCESSING (https://www.unicode.org/reports/tr46/)
Internationalized Domain Names for Applications (IDNA): Background, Explanation, and Rationale (http://tools.ietf.org/html/rfc5894)
Internationalized Domain Names for Applications (IDNA): Definitions and Document Framework (http://tools.ietf.org/html/rfc5890)
Internationalized Domain Names in Applications (IDNA) Protocol (http://tools.ietf.org/html/rfc5891)
The Unicode Code Points and Internationalized Domain Names for Applications (IDNA) (http://tools.ietf.org/html/rfc5892)
Right-to-Left Scripts for Internationalized Domain Names for Applications (IDNA) (http://tools.ietf.org/html/rfc5893)
Since browser already turns it into ascii format before resolving, how would it work in XSS for server-side max length limitation as he mentioned in his other article, "Minimum Viable XSS" [1]?