Cached Chrome Top Million Websites
github.com
github.com
grep -oP '\.[a-z]+(?=,)' current.csv | sort | uniq -c | sort -n
...
15840 .pl
17914 .it
20182 .de
21690 .in
27812 .ru
29194 .jp
30359 .org
35741 .br
36675 .net
406052 .com
.com domains by popularity: grep -oP '[a-z0-9-]+\.com(?=,)' current.csv | sort | uniq -c | sort -n
...
365 tistory.com
370 fc2.com
408 skipthegames.com
489 online.com
515 wordpress.com
707 uptodown.com
880 schoology.com
2570 fandom.com
2651 instructure.com
3244 blogspot.comPeople seem to think it is somehow measuring visits to those origins. But it's measuring how many unique subdomains are listed for those domains
Also interesting I haven’t heard about half of them. Some are nsfw, apparently.
create table current as select * from '202211.csv';
select * from current;
┌────────────────────────────────────┬─────────┐
│ origin │ rank │
│ varchar │ int32 │
├────────────────────────────────────┼─────────┤
│ https://hochi.news │ 1000 │
│ https://www.xnxx.xxx │ 1000 │
│ https://www.wordreference.com │ 1000 │
│ https://finance.naver.com │ 1000 │
│ https://www.macys.com │ 1000 │
│ https://www.xv-videos1.com │ 1000 │
│ https://fr.xhamster.com │ 1000 │
│ https://poki.com │ 1000 │
│ https://salonboard.com │ 1000 │
│ https://clgt.one │ 1000 │
select tld, count(*)
from (select reverse(substr(reverse(origin),1, position('.' in reverse(origin))-1)) tld
from current)
group by tld
order by count(*) desc;
┌───────────┬──────────────┐
│ tld │ count_star() │
│ varchar │ int64 │
├───────────┼──────────────┤
│ com │ 406052 │
│ net │ 36675 │
│ br │ 35741 │
│ org │ 30359 │
│ jp │ 29194 │
│ ru │ 27812 │
│ in │ 21690 │
│ de │ 20182 │
│ it │ 17914 │
│ pl │ 15840 │
│ · │ · │
│ · │ · │
│ · │ · │
│ za:5002 │ 1 │
│ lk:8090 │ 1 │
│ org:1445 │ 1 │
│ co:14443 │ 1 │
│ ar:3016 │ 1 │
│ net:8001 │ 1 │
│ care:9624 │ 1 │
│ au:8443 │ 1 │
│ com:333 │ 1 │
│ edu:9016 │ 1 │
├───────────┴──────────────┤
│ 2076 rows (20 shown) │
└──────────────────────────┘
[0] https://duckdb.org/docs/installation/I’m always amazed that they have a data science team. It’s not something many would expect from the porn industry. I certainly didn’t expect it.
The data science teams likely provide a considerable ROI in that industry.
Seems like the "SafeSearch" filter is based on a list of "adult domains" instead of the indexed content at the URL.
Also to consider: China uses in-app browsing a lot, with interactive experiences very similar to websites built right in the bilibili/ali/wechat apps.
Contrary to popular belief, Google only pulled Search business out of China. The rest of services is still hosted on Google.cn inside China. To download Chrome:
$ curl -svk 'https://www.google.cn/intl/zh-CN/chrome/'
* Trying 180.163.150.34...
* TCP_NODELAY set
* Connected to www.google.cn (180.163.150.34) port 443 (#0)
However the "Make searches and browsing better (Sends URLs of pages you visit to Google)" data won't be collected, because the connection would be blocked.
But that's also just chromium isn't it, much like a PWA? Unless they made something of their own.
But in a market of near 1B internet user, not having a single site in top 1K suggest something is wrong with the stats. I wonder what are we missing from those numbers.
It contains data about ~7 million top websites, and for every website, it also contains: - the full content of the main page; - the verbose output of curl, containing various timing info; the HTTP headers, protocol info...
Using this dataset, you can build a service similar to https://builtwith.com/ for your research.
Data: https://clickhouse-public-datasets.s3.amazonaws.com/minicraw... (129 GB compressed, ~1 TB uncompressed).
Description: https://github.com/ClickHouse/ClickHouse/issues/18842
You can easily try it with clickhouse-local without downloading:
$ curl https://clickhouse.com/ | sh
$ ./clickhouse local
ClickHouse local version 22.13.1.294 (official build).
milovidov-desktop :) DESCRIBE url('https://clickhouse-public-datasets.s3.amazonaws.com/minicrawl/data.native.zst')
DESCRIBE TABLE url('https://clickhouse-public-datasets.s3.amazonaws.com/minicrawl/data.native.zst')
Query id: 6746232f-7f5f-4c5a-ac68-d749d949a2dc
┌─name────┬─type───┬─default_type─┬─default_expression─┬─comment─┬─codec_expression─┬─ttl_expression─┐
│ rank │ UInt32 │ │ │ │ │ │
│ domain │ String │ │ │ │ │ │
│ log │ String │ │ │ │ │ │
│ content │ String │ │ │ │ │ │
└─────────┴────────┴──────────────┴────────────────────┴─────────┴──────────────────┴────────────────┘
4 rows in set. Elapsed: 1.390 sec.
milovidov-desktop :) SELECT rank, domain, log, substringUTF8(content, 1, 100) FROM url('https://clickhouse-public-datasets.s3.amazonaws.com/minicrawl/data.native.zst') LIMIT 1 FORMAT Vertical
SELECT
rank,
domain,
log,
substringUTF8(content, 1, 100)
FROM url('https://clickhouse-public-datasets.s3.amazonaws.com/minicrawl/data.native.zst')
LIMIT 1
FORMAT Vertical
Query id: 8dba6976-0bf6-4ce8-a0f1-aa579c828175
Row 1:
──────
rank: 1907977
domain: 0--0.uk
log: * Trying 213.32.47.30:80...
* Connected to 0--0.uk (213.32.47.30) port 80 (#0)
> GET / HTTP/1.1
> Host: 0--0.uk
> Accept: */*
> User-Agent: Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:84.0) Gecko/20100101 Firefox/84.0
>
* Mark bundle as not supporting multiuse
< HTTP/1.1 302 Moved Temporarily
< Server: nginx
< Date: Sun, 29 May 2022 06:27:14 GMT
< Content-Type: text/html
< Content-Length: 154
< Connection: keep-alive
< Location: https://0--0.uk/
<
* Ignoring the response-body
{ [154 bytes data]
* Connection #0 to host 0--0.uk left intact
* Issue another request to this URL: 'https://0--0.uk/'
* Trying 213.32.47.30:443...
* Connected to 0--0.uk (213.32.47.30) port 443 (#1)
* ALPN, offering h2
* ALPN, offering http/1.1
* CAfile: /etc/ssl/certs/ca-certificates.crt
* CApath: /etc/ssl/certs
* TLSv1.0 (OUT), TLS header, Certificate Status (22):
} [5 bytes data]
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
} [512 bytes data]
* TLSv1.2 (IN), TLS header, Certificate Status (22):
{ [5 bytes data]
* TLSv1.3 (IN), TLS handshake, Server hello (2):
{ [108 bytes data]
* TLSv1.2 (IN), TLS header, Certificate Status (22):
{ [5 bytes data]
* TLSv1.2 (IN), TLS handshake, Certificate (11):
{ [4150 bytes data]
* TLSv1.2 (IN), TLS header, Certificate Status (22):
{ [5 bytes data]
* TLSv1.2 (IN), TLS handshake, Server key exchange (12):
{ [333 bytes data]
* TLSv1.2 (IN), TLS header, Certificate Status (22):
{ [5 bytes data]
* TLSv1.2 (IN), TLS handshake, Server finished (14):
{ [4 bytes data]
* TLSv1.2 (OUT), TLS header, Certificate Status (22):
} [5 bytes data]
* TLSv1.2 (OUT), TLS handshake, Client key exchange (16):
} [70 bytes data]
* TLSv1.2 (OUT), TLS header, Finished (20):
} [5 bytes data]
* TLSv1.2 (OUT), TLS change cipher, Change cipher spec (1):
} [1 bytes data]
* TLSv1.2 (OUT), TLS header, Certificate Status (22):
} [5 bytes data]
* TLSv1.2 (OUT), TLS handshake, Finished (20):
} [16 bytes data]
* TLSv1.2 (IN), TLS header, Finished (20):
{ [5 bytes data]
* TLSv1.2 (IN), TLS header, Certificate Status (22):
{ [5 bytes data]
* TLSv1.2 (IN), TLS handshake, Finished (20):
{ [16 bytes data]
* SSL connection using TLSv1.2 / ECDHE-RSA-AES128-GCM-SHA256
* ALPN, server accepted to use http/1.1
* Server certificate:
* subject: CN=mail.htservices.co.uk
* start date: May 15 18:36:37 2022 GMT
* expire date: Aug 13 18:36:36 2022 GMT
* subjectAltName: host "0--0.uk" matched cert's "0--0.uk"
* issuer: C=US; O=Let's Encrypt; CN=R3
* SSL certificate verify ok.
* TLSv1.2 (OUT), TLS header, Supplemental data (23):
} [5 bytes data]
> GET / HTTP/1.1
> Host: 0--0.uk
> Accept: */*
> User-Agent: Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:84.0) Gecko/20100101 Firefox/84.0
>
* TLSv1.2 (IN), TLS header, Supplemental data (23):
{ [5 bytes data]
* Mark bundle as not supporting multiuse
< HTTP/1.1 200 OK
< Server: nginx
< Date: Sun, 29 May 2022 06:27:15 GMT
< Content-Type: text/html;charset=utf-8
< Transfer-Encoding: chunked
< Connection: keep-alive
< X-Frame-Options: SAMEORIGIN
< Expires: -1
< Cache-Control: no-store, no-cache, must-revalidate, max-age=0
< Pragma: no-cache
< Content-Language: en-US
< Set-Cookie: ZM_TEST=true;Secure
< Set-Cookie: ZM_LOGIN_CSRF=b2dda010-d795-4759-a9c3-80349f3b46ed;Secure;HttpOnly
< Vary: User-Agent
< X-UA-Compatible: IE=edge
< Vary: Accept-Encoding, User-Agent
<
{ [13068 bytes data]
* Connection #1 to host 0--0.uk left intact
substringUTF8(content, 1, 100): <!DOCTYPE html>
<!-- set this class so CSS definitions that now use REM size, would work relative to
1 row in set. Elapsed: 0.539 sec. Processed 4.60 thousand rows, 273.86 MB (8.54 thousand rows/s., 508.28 MB/s.)Is it using HTTP range header tricks, like DuckDB does for querying Parquet files? https://duckdb.org/docs/extensions/httpfs.html
If so, what's the data.native.zst file format? Is it similar to Parquet?
It works for Parquet as well:
SELECT * FROM url('https://clickhouse-public-datasets.s3.amazonaws.com/hits.parquet') LIMIT 1
And for CSV or TSV: SELECT * FROM url('https://clickhouse-public-datasets.s3.amazonaws.com/github_events/tsv/github_events_v3.tsv.xz') LIMIT 1
And for ndJSON: SELECT repo_name, created_at, event_type FROM s3('https://clickhouse-public-datasets.s3.amazonaws.com/github_events/partitioned_json/github_events_*.gz', JSONLines, 'repo_name String, actor_login String, created_at String, event_type String') WHERE actor_login = 'simonw' LIMIT 10
Note: the query above is kind of slow.
Here is the query from preloaded data - your activity in GitHub issues:https://play.clickhouse.com/play?user=play#U0VMRUNUIGNyZWF0Z...
https://clickhouse.com/docs/en/getting-started/example-datas... says "Dataset contains all events on GitHub from 2011 to Dec 6 2020" - but I'm seeing results in there from a couple of hours ago.
Do you know if that's continually updated and, if so, is that documented anywhere?
The source code is here: https://github.com/ClickHouse/github-explorer
This shell scripts updates it: https://github.com/ClickHouse/github-explorer/blob/main/upda...
Disclaimer: I'm not a Clickhouse user, but I have a bit of experience with Parquet.
It looks like the native format is (very briefly) described here: https://clickhouse.com/docs/en/interfaces/formats/#native
It looks similar at a high level to Parquet: binary, columnar and has metadata that permits requesting a subset of the data.
Looking at:
> Processed 4.60 thousand rows, 273.86 MB
I'd guess it's chunking the rows into groups of ~4,000.
The OP must have a nice connection if that completed in 0.5 seconds! (Or perhaps the 273.86MB is the uncompressed size after zstd compression, or perhaps there were other parts of the session that caused that chunk to get cached, and it was elided from what was pasted in to HN.)
EDIT: I was curious, so I ran the tool and watched bandwidth on iftop. It uses about ~50MB each time I run the query. From this, I conclude: it does not cache things, the 273.86MB is the uncompressed size, and OP has a much better internet connection than me. :)
54679
So over 5% of the top 1m sites still don't use HTTPS.
https://play.clickhouse.com/play?user=play#U0VMRUNUIGZsb29yK...
SELECT
floor(log10(rank)) AS r,
count() AS total,
sum(log LIKE '%TLS%') AS tls,
round(tls / total, 2) AS ratio,
anyIf(domain, log NOT LIKE '%TLS%')
FROM minicrawl
WHERE log LIKE '%Content-Length:%'
GROUP BY r
ORDER BY r
┌─r─┬───total─┬─────tls─┬─ratio─┬─anyIf(domain, notLike(log, '%TLS%'))─┐
│ 0 │ 6 │ 6 │ 1 │ │
│ 1 │ 61 │ 58 │ 0.95 │ baidu.com │
│ 2 │ 599 │ 562 │ 0.94 │ google.cn │
│ 3 │ 5591 │ 5057 │ 0.9 │ volganet.ru │
│ 4 │ 51279 │ 44291 │ 0.86 │ furbo.co │
│ 5 │ 476181 │ 361910 │ 0.76 │ funygold.com │
│ 6 │ 3797023 │ 2927052 │ 0.77 │ funyo.vip │
└───┴─────────┴─────────┴───────┴──────────────────────────────────────┘
7 rows in set. Elapsed: 0.844 sec. Processed 7.59 million rows, 43.74 GB (8.99 million rows/s., 51.83 GB/s.)Download it as:
curl https://clickhouse.com/ | sh
Connect to the demo service: clickhouse-client --host play.clickhouse.com --user play --secure grep -o -E "://.*?," current.csv | sort | uniq -c | grep -v "1 ://" | wc -l
8310(I haven't checked whether the documentation is complete/accurate, of course.)
The paper tries to justify its ethics with Google's privacy policy, which is laughable. There are so many papers about how meaningless privacy policies are. If Apple or Mozilla did anything remotely like this, Hacker News would riot.
Edit: I don't want to be a conspiracy theorist, but this post suddenly got a bunch of downvotes at the same time as defensive comments from a current Googler and recent ex-Googler. Then one of my responses below to a Chrome developer got flagged for no obvious reason. Hmm.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
Someone defending this privacy debacle on Hacker News is a Google employee on the Chrome team and was a business cofounder with the Stanford collaborator. That person not only failed to identify how very close they are to the topic, but also phrased their comment in a way that falsely represented distance from the topic. It seems to me essential for understanding their misleading comment to be aware of the factual context.
I thought I had phrased this factual correction in a way that was neutral and not a personal attack. My assumption was that the commenter may have violated Hacker News guidelines by being so misleading. What did I do wrong?
As for the downvotes, I see that I should have emailed you rather than adding a note in the comment. Nonetheless, could you see what's going on?
The commenter publicly identifies themselves in their HN profile and you're using that to attack them. It's completely backwards to say they've misrepresented anything. The essential thing is to assume good faith and not go on weird innuendo-laden witch-hunts.
I'm not disagreeing with you about the underlying issue—there's an argument to be made that the kind of "publishing" that Google/Chrome does here is is really a way of obscuring it from the majority of users, and so on. HN commenters are certainly welcome to make that kind of argument. But we need you to err on the side of not posting in the flamewar style. If I see a commenter posting in the flamewar style and then also bringing in someone's personal details as ammunition, it's no longer a tough-thing-to-balance, it's just out of line.
"Comments should get more thoughtful and substantive, not less, as a topic gets more divisive."
This includes only listing publicly discoverable pages, only including data from users who have turned on "Make searches and browsing better (Sends URLs of pages you visit to Google)", and only including pages that are visited by a minimum number of users.
If this is news to Hackers News, there is no way that regular Chrome users are aware of it. Saying something in a privacy policy or on a developer website just can't be enough for analyzing a person's URL data.
> This includes only listing publicly discoverable pages, only including data from users who have turned on "Make searches and browsing better (Sends URLs of pages you visit to Google)", and only including pages that are visited by a minimum number of users.
Since when does aggregating this type of data make it fair game? This is analyzing a person's URL data from their own devices. There has always been a big bright red line for browsers touching a user's browsing history. Google crossed that line.
Also, I just checked on a fresh Chrome install. The "Make searches and browsing better" option is enabled by default and buried in Chrome settings. How is that acceptable consent for analyzing a person's URL data?
One big problem there is that we don't know what percentage of users for whom "turned on" is a euphemism for "didn't notice."
I’m not using Chrome on all my devices.
2. Chrome prompts you to opt-out of metrics collection on install.
None of the reasons you've listed for this being ethically dubious are true.
Because the default should be "opted out by default, let the user opt-in if they so wish"
That's not merely a good idea but also
https://news.ycombinator.com/newsguidelines.html
Please don't post insinuations about astroturfing, shilling, bots, brigading, foreign agents and the like. It degrades discussion and is usually mistaken. If you're worried about abuse, email hn@ycombinator.com and we'll look at the data.
There's also just not writing in the high-dudgeon flamewar style which helps with the downvotes.
I've noticed similar behavior in HN voting. Down vote spikes but few if any comments in-line with the voting. Not sure if it's bots, human-based click farms, or too just don't understand that disagreement is not grounds for down voting.
Perhaps a bit of all three?
It is perfectly fine on HN and always has been.
The Guidelines are clear about why we're here and expectations. The emphasis is on discussion, learning and objectivity. Yes, disagreement is mentioned (i.e., allowed) but even that needs to be constructive, yes?
A down vote - with no discussion - well, frankly in the context of the Guidelines, is:
1) Not in the spirit of the guidelines; 2) Perhaps redundant to 1, but lazy; 3) At best, small-minded and childish;
If people want to pout about reading something they don't like, this isn't the place for them.
Yeah, I see who you are. And I'm ok w/ pushing back. That's what make HN what it is ;)
https://news.ycombinator.com/item?id=22910444
https://news.ycombinator.com/item?id=16131314 and there are many many others
Yeah, I see who you are.
I'm literally a random scold on the internet, I just happen to be right about this.
I'm not going to explain why.
How does that feel? What value does it add? (Sweet FA, eh.)
You're right, you might be right. But that does make it right. I get zero satisfaction from context-less down votes. I don't do them. I ignore them when I get them (i.e., they have zero influence on my HN behavior). If I'm changing my mind over some lazy a-holes' click, I'm losing. Big time.
I can't imagine why anyone feels any differently. The reality is, there are pointless noise. There's not enough context to drive anything actionable for anybody.
But while I have your attention: how about a feature request: Karma points that consider the discussion below a top-parent comment.
My perception is that, collectively, HN hates and criticizes Google much more than Apple and Mozilla. I mean, much more. This last sentence accusation sounded bizarre to me.
Not the entirety of HN. As I have more than once delicately pointed out[1], Mozilla is Google's bitch.
That is mostly because Apple almost never does something like this, and Mozilla literally never does.
Because Google is a web advertisement company that dominates many large spheres: search, browsers (including standards committees), email, mobile (Android is 77% market share) etc. All are things that we've come to view as crucial to modern life.
And time and again they've shown that they only view that dominance as a funnel for ad revenue, data collection, and whatever benefits them at this particular moment.