https://www.shopify.com/robots.txt lists a lot of sitemap files, which tend to be a good starting point.
# / <//_\
This is being interpreted as a malformed HTML closing tag, which (according to the HTML5 parsing algorithm published by WHATWG) gets treated as a comment. The file doesn't contain any > past this point. This leaves the uncommented contents from lines 1–6: # ,:
# ,' |
# / :
# --' /
# \/ />/
# /
Or, with whitespace collapsed: # ,: # ,' | # / : # --' / # \/ />/ # /
Which should be exactly what you observe.Ref: https://html.spec.whatwg.org/multipage/parsing.html https://developer.mozilla.org/en-US/docs/Web/CSS/white-space...
The robots.txt does not exclusively list what not to scrape.
It provides information on which parts are allowed and wich are not (disallowed).
It also provides sitemaps for crawlers as a starting point with more information (eg. which sites are available and how often are they updated, etc.)