Cool! But how did you get the initial dataset of 643,000+ Shopify stores (data as per your “About” page) in the first place, to then scrape the products from their /products.json feeds? Or did you just try a huge list of domain names at random?
Does "Built With" provide that data? How accurate do you think it might be?
# / <//_\
This is being interpreted as a malformed HTML closing tag, which (according to the HTML5 parsing algorithm published by WHATWG) gets treated as a comment. The file doesn't contain any > past this point. This leaves the uncommented contents from lines 1–6: # ,:
# ,' |
# / :
# --' /
# \/ />/
# /
Or, with whitespace collapsed: # ,: # ,' | # / : # --' / # \/ />/ # /
Which should be exactly what you observe.Ref: https://html.spec.whatwg.org/multipage/parsing.html https://developer.mozilla.org/en-US/docs/Web/CSS/white-space...
The robots.txt does not exclusively list what not to scrape.
It provides information on which parts are allowed and wich are not (disallowed).
It also provides sitemaps for crawlers as a starting point with more information (eg. which sites are available and how often are they updated, etc.)
Shopify sites also have shop-name.com/products.json which has URLs that point to cdn.shopify.com