99 karma · joined March 31, 2023
Some findings from the dataset:
Geography (top countries):
United States: 253,589
India: 34,127
Canada: 20,263
United Kingdom: 18,701
Pakistan: 10,124
(392 total countries)
Industries:
E-commerce: 164,010
Adult & Gambling: large category overall
Gambling (L2): 84,353
News & Blogs: 49,424
SaaS products: 39,105
Niches (L3):
Online Casinos: 81,608
Clothing: 33,553
Niche SaaS tools: 31,881
Home Decor: 31,273
News & Journalism: 29,750
Platforms (detected on ~295k sites):
WordPress: 116,250
Shopify: 84,407
WooCommerce: 42,615
Squarespace: 25,328
Wix: 23,598
Webflow: 3,049
TLDs: .com: 435,622
.store: 38,223
.org: 26,474
.online: 23,422
.site: 22,919
.ai: 6,167
Happy to answer questions about methodology, accuracy, crawling, classification, or detection heuristics.
I get most of it, but I think especially around the holiday some stuff is getting through... Some black friday deals were actually hitting like news does...
Some stories are very clearly manufactured
Thanks for that bug feedback - ill get fix.
I'm not pulling from social media yet.
The embeddings themselves will (pry) cluster ok in different languages (but I have not tested this yet)
I think checking source in story is next step...
No translation yet.
I think the biggest problem is im relying on published date from the news source itself too much and its wrong sometimes... not super often, but if 1 out of 100 sources get its wrong then it can steal credit for being source article when its not.
Yea, I need to do some work on improving first to publish... currently I'm relying pretty heavily on the published date provided in the story itself, but sometimes that is wrong and makes it look like a later publisher was first to publish.