Googlebot’s JavaScript random() function is deterministic
tomanthony.co.uk
tomanthony.co.uk
Ads often have something like this attached to the end:
load("https://adserver/track.gif?{stuff}&_=" + Math.random())
My ad server collected multiple tracking pixels -- for various events and after five pixels I could fingerprint the browser (Firefox, Chrome, MSIE, etc) and identify someone fiddling with the user agent string, or using a proxy server to mask this information.Going to extreme measures to be untraceable is like wearing a Ghillie suit to the airport.
Ultimately, the sizing of a font can effect element widths, and element widths can be queried.
• Did proper random before anybody else
• Active countermeasures against cookie-based retargeting
• Popular enough market that merely having an iPhone in a geographic area doesn't single you out
There's a guy who downloads every page with curl. I see him on web logs. I think he must have some script that parses the amp pages out and does something with it, but because he's the only person in that geographic region who browses the web with curl, he's very easy to spot from a tracking perspective. On the other hand, because he's using curl, I don't think anyone wants to bother trying to show him an ad.
On the other hand, I think I've read a FAQ where he says he views web pages by having a script fetch the pages with wget and them emailing them to him, but I'm pretty sure that's simply because wget is free and does the job (being able to recursively download the necessary resources), not because it's GPL licensed.
I generally do not connect to web sites from my own machine, aside from a few sites I have some special relationship with. I usually fetch web pages from other sites by sending mail to a program (see https://git.savannah.gnu.org/git/womb/hacks.git) that fetches them, much like wget, and then mails them back to me. Then I look at them using a web browser, unless it is easy to see the text in the HTML page directly. I usually try lynx first, then a graphical browser if the page needs it (using konqueror, which won't fetch from other sites in such a situation).
I occasionally also browse unrelated sites using IceCat via Tor. Except for rare cases, I do not identify myself to them. I think that is enough to prevent my browsing from being connected with me. IceCat blocks tracking tags and most fingerprinting methods."
If you just want to prevent yourself from being identified as an individual, that's a different problem. Tor browser does a pretty good job of solving that.
Cons: interactive infographics and courseware don't fit neatly in either.
Or just use an extension that replaces Math.random() with something more random, but it's possible that could cause weird performance problems on certain pages and it would be hard to debug.
Also, probably useful for determining two pages are the same, which may be needed to help prevent the crawler from crawling a million paths into a SPA that don't actually exist, for example.
PRNG != Random.
You've just ruined most stream ciphers. That's not true at all.
>Predictable – Googlebot can trust a page will render the same on each visit
It's probably important for Google's crawler to identify whether a page changed or not, if some elements in a page are randomly generated they may want to limit the impact.
I mean, after all they seem to use a real, changing value for their Date, so if they wanted they could just seed their RNG with that.
But I believe Googlebot always faithfully sends it's user-agent. Is there a reason Google would care about 'fixing' this to make Googlebot harder to detect via random() predictability, when you can always just detect it via user-agent anyway? I'm not sure, curious if others have thoughts!
If the "disguised" googlebot is the same as the actual one, chances are it is since it would want to be as close as possible to not flag false positives, and use the same seed for consistency then you might be able to use that to avoid detection on the fact that you are serving google something different than normal users.
Newspaper used/do that to be able to have their full article content indexed while serving a paywall to everyone else.
[0] https://intoli.com/blog/not-possible-to-block-chrome-headles...
[1] http://antoinevastel.github.io/bot%20detection/2018/01/17/de...
https://www.elephate.com/blog/chrome-41-key-to-website-rende...
The warning you're seeing is
a) probably shown based on UA, and Googlebot has a very different UA to the Chromium browser
b) warning users who will need to browse/actively use the site. Googlebot simply parses content, so has no use for quite a lot of the active functionality on the site. As such, it will generally not need to support all of the features used by that active functionality, just barely enough to be able to read content.
YouTube copyright violations?
The chance of getting a response from Google support?
I'd also assume that the Twitter firehose could be a great source of randomness.
Bottom line: the PRNG is, as all PRNGs are, deterministic. The user-facing math.random API using the underlying PRNG may or may not be. Those are distinct things.
From the spec's description of math.random with regard to random or pseudo-random: "...using an implementation‑ dependent algorithm or strategy."
http://www.tomanthony.co.uk/fun/googlebot_puzzle.html
The idea being you could use a function to always send Googlebot down one execution path and users down another, and make it look like you are doing an AB test on a small set of traffic. You could then do something nefarious, such as add spammy content to the page for Googlebot.
However, in reality it is not likely to be a viable tactic for any decent quality site, and unlikely to have much of an impact on lower quality sites that may be willing to risk it.
It is an interesting idea though, and we research this sort of thing such that we can better identify such behaviours in competitors of our clients.
On another note, this might be a neat way to avoid or deliberately screw up Googles "knowledge scraping" thing that they're doing which is pissing people off mightily (where when you search for a question, Google gives you the direct answer that they've scraped from another site.)
It is actually truly random if you have a fair dice.
There's no way to prove, of which I am aware, that a string of 24 fours is not random.
Of course, to get 4 out of a PRNG you probably have to ask it for a random number in a range - if you’re requesting numbers in the range 0-1000 then you would expect fewer long sequences of fours than if you request in the range 0-5. And you can get arbitrarily many 4s by requesting numbers in the range 4-4...
You can't prove that a string of 24 fours is not random, and the fact that some process returned 24 fours once is not proof (but is some evidence) that it's not random - however, a process that always returns 4 is not random.
A process that returns a single fixed number (which was chosen by a fair dice roll) once is a random process. Using that same process twice or more is not.
The way my professor explained: a drunkard can always find his way back to his building, but not to his apartment.
That might be true for a drunk Kitty Pryde, but is it true for a drunk ordinary human who is constrained to only move between floors via the stairs or elevator?
--js-flags="--random_seed=12345"
to Chromium.Unless you are using some form of entropy, e.g: dedicated hardware, that will be the case.