Apple's whitelist of the 250k auto-completable domains in iOS
cdn.smoot.apple.com
cdn.smoot.apple.com
The list exists initially to support a feature, a utilitarian reason. Then, being in the list has SEO value or some such. The list becomes a legitimacy or authority test. It affects traffic, sales.
Maybe other stuff piggybacks this list. Anyone who needs a list of important domain names uses it. Ordering search results, caching, testing, crawling... More dynamics unrelated to Apple's original purpose.
Meanwhile, Apple continues to treat the list as a relatively unimportant thing... Just a clunky little hack that enables autocomplete. At some point, willfully ignorant of a whole industry of professional services out there promising to get you on list.
Does anyone know why this list is hosted and what it is used for?
Is it even used?
Edit: some 7 year old posts suggest it was used by Spotlight.
I also am a little amazed it loaded and scrolls smooth on Firefox mobile...
In summary: I don't have any indication it's in current use. I saw this URL in intercepted traffic from an iOS 13 device. I suspect it was used at some point but not anymore. I bet iOS 15 still uses something similar but doesn't use this particular URL (perhaps one with more obscurity/protection).
p.s. I also found this list of spelling corrections/autocompletes: https://cdn.smoot.apple.com/static/static_corrections_dict/2...
Some examples:
{"q":"hackernews","c":{"hackerneww":0,"hackernees":0,"hackernee":0}}
{"q":"ycombinator","c":{"ycomi":0,"ycombinatir":0}}
{"q":"apple store","c":{"apple stire":0,"apple stoe":0,"apple stoee":0,"aple st":0,"aplle":0,"aple":0}}It's a binary file, but if you look at the hex dump there's a list of vaguely "bad" words near the end including "whorehouse," "Zipperheads," "unabomber". I have no idea what format it is, though.
How could they realistically do that?
They could use a similar tool plus human review to maintain the list.
Moderation is a hard problem because it isn't just a matter of someone filtering between the polite posts and the less polite posts, it's a matter of filtering between the polite posts and the content that will sear your soul, no joke.
But that's not what this is. This is just, is the website still there and look correct? With the correct software setup it's roughly a person-month by my estimate to gets eyes on every site in the list.
(Though most people usually don't set write this sort of software very well, making someone laboriously click this, scroll around some, click some more, click a tiny radio button, click the tiny submit button, wait for the next thing to load, etc. It'll be longer & more work with this style. Someday I hope to have the chance to write some sort of classification program and implement the UI I've wanted for a while, which amounts to "right -> ham, left -> spam", and everything as pre-rendered as I can get it before it gets to the human. I'm sure some people out there have done something like this, but it makes me honestly sad how few I've seen.)
That would be a good point if this were the halting problem. It's not. It's a list of domains that you're suggesting to users.
For starters, a VERY basic solution might be to look up the domain name ownership information and see if that has changed. If so, flag for review.
Secondly, you can store the public SSL certificate and make sure that's still the same. If it changes, flag for review.
Thirdly, screencap the site, save it, periodically re-cap and compare how similar the images are. If it changes, flag for review.
> I can't begin to imagine how they would do what you're suggesting.
Did you try?
> They need to detect when a website changes in kind, but ignore day-to-day changes or normal UI revamps.
The solution doesn't need to be perfect, it needs to be good enough.
The only realistic option I can think of is some combination of:
• Make autocomplete operate on a blacklist instead of whitelist, with a more limited goal of only removing e.g. known porn sites.
• Make the list of potential matches machine-generated, without human intervention. (Aside, are we sure the current list isn't just the 250K most-visited sites on the internet, or something like that?)
Either of these would remove culpability since it's no longer a curated list. And yet, would that make it more safe in a meaningful way?
If Apple wants to protect users then autocompleting porn site domains seems like the place to start, not avoid.
The absence of porn domains raises the question of Apple’s intent.
Because if I’m sharing my screen on a business call, and I start typing something into my browsers’s address bar, I don’t want it to autocomplete something nsfw which just happened to share the same first letter.
Furthermore, we now know that in practice "machine-generated" seems to be even worse, because too many people are fooled by the "the machine did it" 'excuse'. (Like you seem to be doing here ?)
For instance : https://thedataist.com/book-review-automating-inequality/
2: Did the owner information change since last time?
2n: No action required. Maybe select some sites randomly to have a human compare, but it's probably fine.
2y: Have a human check the site. Did it simply get purchased by another (similar) organization, or is it no longer relevant to its original purpose?
3: Do the needful.
It might be a list that was made by each company request? Based on some form or authentication or as such? And the companies who didn’t approach Apple aren’t there. Just guessing. But I can imagine Apple doing such hacky things, though most of those remain behind the orchard curtain.
Now I know why...
Also it sometimes is too slow to pick up that I don’t want to go where it autocompleted, I hit go, then immediately hit back and typed more characters in to get the actual site I wanted. I could go and edit my history, but that process is a little clunky on my phone.
1. "tempo.ai", which doesn't actually resolve
2. "web.ai", an http-only site that looks kind of like a cross between a geocities website and a parked spam page, with the content allegedly last revised almost 20 years ago (2003)
I wonder how old this list is.
Currently on websites starting with 'joe'.
Because Internet entropy is fun. More fun than a recommendation engine.
These addresses are Onions v2 and now disabled in Tor.
For others, here's the other ones found:
- "correction_dict_url": "https://cdn.smoot.apple.com/static/static_corrections_dict/2..."
- "crowdsourcing_blacklist_url": "https://cdn.smoot.apple.com/static/crowdsourcing_blacklist_u..."
- "crowdsourcing_whitelist_url": "https://cdn.smoot.apple.com/static/crowdsourcing_whitelist_u..."
- "spotlight_model_resources": "https://cdn.smoot.apple.com/static/spotlight_model_resources..."
- "spotlight_stopword.map": "https://cdn.smoot.apple.com/static/spotlight_suggestions_sto..."
- "spotlight_phrase_dictionary.map": "https://cdn.smoot.apple.com/static/spotlight_suggestions_phr..."
- "silhouette_topic_mapping": "https://cdn.smoot.apple.com/static/silhouette_topic_mapping/..."
- "silhouette_whitelisted_topics": "https://cdn.smoot.apple.com/static/silhouette_whitelisted_to..."
- "silhouette_config": "https://cdn.smoot.apple.com/static/silhouette_config/5/silho..."
- "dictionary_resources_url": "https://cdn.smoot.apple.com/static/dictionary_resources_url"
Also I just realized that the source of all these doesn't require auth either and you can just view it here: https://api.smoot.apple.com/bag (but requires a spoofed user-agent if you use cURL)
$ curl -s https://cdn.smoot.apple.com/static/autofill_tld_whitelist_url | jq '.tlds|length'
249999I got 250k by running this:
$ curl https://cdn.smoot.apple.com/static/autofill_tld_whitelist_url | jq | wc -l
250004
And then I naively assumed with the extra JSON brackets the total would be 250k but I wasn't thorough enough!This attempts to scan WHOIS records for random domains in the list and try to find the most recent "registered at" date for a domain.
It's very unreliable. The whois-parser Ruby gem thought that giants.it was created on 2020-11-03 09:00:01, but I checked the record and it looks like it was actually created on 2009-09-28 11:00:43.
It looks like nature-in-art.org.uk was first registered on 21-Jul-2019
I don't know if this would provide any useful information. It would take at least a year for a brand new domain to make it into this list. And I don't know if the "created at" date can be trusted, maybe it gets reset whenever a domain is transferred, or lapses and then gets renewed.
In conclusion, I don't think this idea works at all.
I've found wiibrew.org and hackmii.com in there which are both Wii homebrew sites that became popular around 2008/2009 and probably declined in popularity starting in ~2012/2013.
Then there's also wiiu-developers.nintendo.com and wiiudaily.com which probably didn't exist before late 2012 or early 2013 when the WiiU was released.
I see lots of websites for small German towns, which should only have a couple of visitors per week.
Of the blogs and personal websites I read - some popular, some less so - only paulgraham.com is there.
I suspect the list is biased to official and very old domains.
I assume it must have come from user searches at some point?
tl;dr; I bet it'll never get updated. (but some similar list at a more hidden/protected location likely will be)
p.s. this list of common spelling corrections is interesting too: https://cdn.smoot.apple.com/static/static_corrections_dict/2...
The list has "orange.co.il" (now NXDOMAIN) but not "partner.co.il". Partner (one of 4 main cell networks in Israel) used to be called Orange but they terminated that agreement with Orange FR years ago (which I'm guessing is why they can't redirect "orange.co.il" either).
In a way it's a slight step towards allowing or disallowing visiting certain websites. Is the owner of the phone allowed some control here?
e: This is reinforced since we updated our domain/URLs from a ccTLD to a gTLD not long ago, also claiming the Maps entry, and the list reflects the gTLD domain!
I always long for Firefox's url bar, it easily beats everything else since Firefox 2 came out or so. Nowadays you have to un-configure two layers of "let's put a search bar in your URL bar" but then it's still just amazing. Especially compared to google, which obviously doesn't have an incentive to have you skip searching on google.com... and apparently apple doesn't want to you complete just any old urls, but only apple approved ones?! This seems like such a useless kind of (soft, but still) gate-keeping...
Firefox on Android e.g. used to (or still has – I'm not sure what the state is nowadays after the rewrite) have a similar (although somewhat smaller) list of popular domains which it would use for providing autocomplete results if it couldn't find anything in your history, but local history always had priority.
Joking aside - Firefox makes some strange choices and the I wish mobile and desktop could (optionally) share more of my browsing history. Maybe they do but the autocomplete doesn't behave the same across devices.
No. Contrary to popular belief, browsers on iOS aren’t all Safari skins. They all use WebKit, but there’s a vast amount of functionality in a web browser that isn’t handled by the rendering engine. This is one example.
Non-Safari browsers on iOS are free to use whatever address bar implementation they like. If Firefox on iOS has a crap address bar, that is 100% down to Firefox.
I know some people hate Apple for both good and bad reasons, but it's pretty wild to me people think Apple does this. I mean right out the gate that would be pretty illegal, no?
Every browser you use is forced by App Store policy to run off of the system WebView, which is Safari...other features being part of the browser doesn't change being forced into using that engine.
No. It’s WebKit. Not Safari. That was the whole point of this subthread.
Because Apple is selling the hardware device, they can decide what software they ship on it. This is not specific to phones, it’s true of all hardware. There’s no feature to load arbitrary browsers onto a PlayStation, for example.
Not anymore. https://en.wikipedia.org/wiki/OtherOS
For example, search for “nytimes.com”. Many of the subdomains haven’t been used in years - and are to blogs that were retired 6 or so years ago.
[1] https://cs.chromium.org/chromium/src/net/http/transport_secu...
Here's some other ones:
- "correction_dict_url": "https://cdn.smoot.apple.com/static/static_corrections_dict/2..."
- "crowdsourcing_blacklist_url": "https://cdn.smoot.apple.com/static/crowdsourcing_blacklist_u..."
- "crowdsourcing_whitelist_url": "https://cdn.smoot.apple.com/static/crowdsourcing_whitelist_u..."
- "spotlight_model_resources": "https://cdn.smoot.apple.com/static/spotlight_model_resources..."
- "spotlight_stopword.map": "https://cdn.smoot.apple.com/static/spotlight_suggestions_sto..."
- "spotlight_phrase_dictionary.map": "https://cdn.smoot.apple.com/static/spotlight_suggestions_phr...
- "silhouette_topic_mapping": "https://cdn.smoot.apple.com/static/silhouette_topic_mapping/..."
- "silhouette_whitelisted_topics": "https://cdn.smoot.apple.com/static/silhouette_whitelisted_to..."
- "silhouette_config": "https://cdn.smoot.apple.com/static/silhouette_config/5/silho..."
- "dictionary_resources_url": "https://cdn.smoot.apple.com/static/dictionary_resources_url?..."
Presumably to auto complete URLs, like "go" to "google.com", in Safari's address bar.
I (think) i have turned off all search-suggestions in settings.
There are several furry art communities listed, but Fur Affinity is not. Fur Affinity is the largest and most popular of these by a fair margin.
Unrelated to the furry websites the above, there are some nsfw websites listed but other, more popular ones are strangely missing. (The word "porn" does not appear in the list.)
redditenhancementsuite.com is also listed... which I did not expect.
I suspect this list is run through a keyword filter and a blacklist before anything is added.
Keep your friends close and your enemies closer.
> breitbart.com, stormfront.org
Lulz.
The threat model isn’t too different from other things that can happen if a malicious user is on the same network as you. The scale would be different though.
https://cdn.smoot.apple.com/static/autofill_tld_whitelist_ur...