"Given how much of the web is rendered client-side these days you need to start out with a headless option."
What does "rendered client-side" mean.
Assuming that "rendered client-side" means interpretation and execution of Javascript is necessary to read a site's textual content, then how much of the web is rendered client-side.
If the focus is on textual content, e.g.,, someone is primarily "scraping" text as opposed to images and video, I would guess that only a minority of the web is "rendered client-side". How would we prove otherwise.
This guess I am making would not be an uneducated one. I have been accessing the web without using Javascript for over 30 years. Today, I still use a text-only browser to render HTML as text/hypertext. This allows me to read the site's textual content, quickly and easily. I initiate most HTTP requests with TCP clients, not the browser. All requests, whether from TCP client, browser, or otherwise, are made through a localhost forward proxy. If most websites were truly dependent on Javascript, it stands to reason I would not be able to read much of the web. In other words, another web user who reads the web with a Javascript-enabled browser should be able to read websites that I could not read. This has not been the case. In fact, I often see commenters on HN complaining that they cannot read a site that I am having no trouble reading. The culprit is often Javascript.
The truth is that I rarely encounter a site that cannot be read with the text-only browser. For example, I can read the content of almost every site submitted to HN. A very small minority of sites I find are, more or less, empty shells with links to some Javascripts but no textual content for the visitor to read. These "landing pages" expect a Javascript-enabled browser that automatically follows links in the page (e.g., to remote Javascript files), and that retrieves, interprets and executes Javascript automatically and indiscriminately.[FN1] In what some might see as a Rube Goldberg design pattern, the scripts then make HTTP requests to the "real" site. In such cases it generally only takes me a few minutes to find the "real" site, often what some refer to as a "JSON endpoint".[FN2] However this process has not lead me to rely on a "headless" browser to read websites.
Honestly, if a majority of sites adopted the "JSON endpoint" approach to serving textual content it would make reading websites even easier for me. I could just retrieve JSON and reformat it to a uniform brand of simple HTML that I prefer, as I already do for some sites. I could make the format of all websites 100% identical. IME, a web of uniformly-formatted content is much easier and faster to digest. I would imagine it would easier for machines to digest as well. The text-only browser I use currently makes the format of all sites look almost the same, since it only uses a single font and so many websites use similar designs. Because it does not automatically follow links or execute Javascript, it also tends to make the "load" time of all sites very similar. For me, this uniformity speeds up the ability to digest web content as compared to using a graphical browser for the same purpose.
FN1. Today we see "modern" browsers incorporating an ever-changing array of "features" and options to try to mitigate the risks of this behaviour.
FN2. Generally, IME, these "endpoints" serve the textual content with minimal markup or sometimes no mark up at all. Thus, the end user is free to format the text into whatever design suits their personal tastes. As a website visitor, this is relatively more efficient IMO than trying to read an infinite number of possible "web designs" which is the approach we currently see on today's www. It is more predictable. With the later approach, visiting a new website with a Javascript-enabled, graphical browser is always a "surprise". It might be easy to read or it might not. Visiting "endpoints" generally does not suffer from this problem.