There's been some pushing for browsers to automatically decide what language you want to view a webpage in, and that is even more insane than the User-Agent header is. The correct solution for content that exists in multiple languages is to make each language available at its own URL. Just like I already said. Why in the world would you want to lose the option to select the language you want?
User-Agent spoofing was added to NCSA Mosaic in 1996. The public www was three years old and text-only clients were still in widespread use
https://raw.githubusercontent.com/alandipert/ncsa-mosaic/mas...
NCSA Mosaic is the early graphical browser that begat Netscape Navigator that begat Firefox. Later came Internet Explorer, Safari, Chrome and so on
What was the point of spoofing in 1996
Maybe it was just for fun, judging by the examples in the "mosaic-spoof-agents" file
As a matter of practice, by default I do not send a user-agent header. I only send the minumum headers required. For me, that's almost always 1-2 for GET and 4 for POST
For the vast majority of websites I have accessed,^1 this header minimisation has zero effect on the success of the HTTP request
1. For example, I have used a database of sites submitted to HN as way to test if header minimisation affects HTTP request success
Generally I do not use a web browser to make HTTP requests. I perform text processing on the HTML, JSON or whatever is returned, using custom utilities. I store information in SQLite. I read this information as plain text, preferably 7-bit ASCII. I dislike UTF-8
Onlilne debates about "browser fingerprinting" always seem to focus on trying to "blend in", e.g., via "spoofing"
As such, because browser continue to get more bloated with "features", online commenters argue in favor of sending more and more data points to servers that can be used to create a fingerprint instead of reducing the amount of data sent
Because, according to their reasoning (or lack thereof), sending less data would "stand out"
True, but it's generally easier to "spoof" a client that sends less data than one that sends more. And the number of sites that require a user-agent header is still smaller than the number that don't
For example, https://www.apple.com does not require a UA header
For example, http://www.slackware.com does not require a UA header. However, it does require Accept and Accept-Encoding headers
This site can also be accessed over HTTPS
They're pretty brutal in blocking non-browsers, I tend to have to add a lot of useless headers.
Blocking "non-browsers" and blocking IP addresses are two different things, so to speak
The first is based on dumb heuristics and (incorrect) assumptions about behavior based on what software someone is (mistakenly) presumed to be using
The second is based on past behaviour
I send the minimum HTTP headers
I only request what I want; I do not send requests for ads, tracking or telemetry
I accept text formats, e.g., HTML, JSON, etc.
No need for images, CSS, Javascript, etc.
IMO, there is a fundamental difference between (a) not using the data collection, surveillance and advertising-friendly software, i.e., popular graphical web browsers controlled by so-called "tech" companies that, surprise, said so-called "tech" companies and their business partners want people to use and (b) engaging in bad network behaviour, what some used to call poor "netiquette"
The heuristics used to allegedly identify bad behaviour, for example, sending the "wrong" value in a User-Agent header, are beyond stupid
But the false positive totals, "collateral damage", are apparently not large enough to matter
For example, something like "90% of the current internet is behind Cloudflare, Google, Akamai, etc."
The www is only subset of the internet and "90% of the internet" does not use a CDN
In any event, even if "internet" means "www" there is no evidence to support this statement
https://www.netcraft.com/blog/may-2025-web-server-survey
Further, being "behind Cloudflare, Google, Akamai, etc." does not necessarily mean a User-Agent header is required
For example,
https://www.cloudflare.com does not require a UA header
https://example.com is "behind CF" and does not require a UA header
https://www.google.com does not require a UA header
https://developers.akamai.com does not require a UA header
And so on
Individual www users, unless doing mass scanning or the like, probably do not access 100% of the www's sites (according to Netcraft, most are inactive)
More likely, individuals access only a fraction of the active sites
It's unlikely that individuals access exactly the same list of websites^1
For me personally, the vast majority of sites that I access do not require a UA header
The sample set I use as evidence to support comments I make about this on HN is a list of all sites submitted to HN over a recent period.^2 Currently this is 12,641 sites
1. To have a meaningful discussion about UA header requirements, a sample set of sites must be agreed upon
2. Perhaps HN readers access these sites and possibly access some of the same sites. Of course it's also possible that there are HN commenters who do not actually access HN submissions and instead only read comment threads and submit replies
https://news.ycombinator.com does not require a UA header
In the bad old days there were so many differences between html, css and js behaviors that if you wanted your site to be nice you had to change it for the browser. The way css padding worked wasn't even the same. Feature detection was rarely viable for any of this.
No user agent would probably have only entrenched IE6 dominance even more by blocking you from deliberately making a site that works at all on other browsers (including IE7 for that matter)
JS devs were kinda able to patch around the nonsense because they were able to feature-detect - part of the reason this stuck around was because no legitimate user or dev cared (or should care). But the header was mostly (useless) noise, and the people spoofing were dealing with the couple bad apples of the time.
Of course, defining features is easier said than done, and a standards body is a challenging environment to define these in...
I get why people are fingerprinting bots and others are working around it, but neither are "legitimate" applications - if your content is public, it's public, end of story. And working around these controls to sell botnet access to sites is equally illegitimate - nobody has a right to resell content they do not own...
I wish this was the case. Unfortunately, companies that work on non-Chromium browsers need to employ dedicated web compatibility teams to either a) help website users fix non-standard (i.e. Chrome only) HTML/CSS/JS, or b) replicate Chromium-like behaviour for specific (very popular) websites so that they work "correctly".
There's also the websites that deliberately block certain browsers which is what tools like "chrome-mask"[1] are built to solve.
[1] https://addons.mozilla.org/en-GB/firefox/addon/chrome-mask/
Not really, there's just only one browser (Chrome). Firefox has declined so low it's a rounding error, and all other browsers are Chrome forks.
In practice when you make your new set of hacks the string you can always evaluate whatever cruft in the useragent today, but next browser shows up.
It does make the user agents insane but I don't know if there's any obviously better system for the problem, even with hindsight
And detecting their platform doesn’t stop you from presenting supplemental links for all platforms. In fact I’d say that’s completely normal.