I plan on writing a long post on my experience writing a web crawler in Elixir, and I believe it is a harder problem than many—including me one year ago—think it is. Web servers out there are just broken, pages have invalid HTML that somehow browsers are able to interpret, etc.
Here's some things Bernard does that many other don't:
* I respect robots.txt for each link I visit, and keep a single connection per {host,port} not to hammer popular domains. Maybe not ideal speed-wise, but I try to be a good Internet citizen.
* Modern MITM-like services such as Cloudflare, Akamai, etc. need to be accounted for, as they might return 403s even if the link is working perfectly. I am planning on registering as a "Cloudflare good bot" very soon, right now I simply ignore and log 403s for Cloudflare-server domains. The goal ultimately is to have as many false negatives as possible.
* In the future, Bernard will be able to parse CSS and JS files, to find and test links to @font-face URLs, imported JS scripts, etc.
* If there is enough demand, as it is quite involved infrastructure-wise, I might consider adding support for client-side only websites, through headless Chrome (on a separate, more expensive tier).
* Not shipped in this MVP, the original goal of Bernard is to remember all the links its seen, so it is able to notify you if one 404s because you changed its URL, moved it around, and forgot to set up a redirect. This happens far too often in my experience, breaking bookmarks and causing SEO issues.
And last, but not least, none of the alternative I have tried, freeware or FOSS, felt good, correct or well-designed enough to use, from my point of view. I do believe there is always a space to do better than the status quo.