“Fetch as Googlebot” tool helps to debug hacked sites
mattcutts.com
mattcutts.com
Still nothing is perfect and this is a good read about a issue alot of people are not aware of.
For most sites that store their templates as regular files and their contents in a database, this is plenty good enough. For sites that store their content as regular files too, it only takes a few extra minutes to separate the good stuff from the bad stuff.
This is super easy to implement. Every web host should be doing it.
Thanks for thinking of that though.
As a reporting tool it could be interesting though.
The nature of the compromise is something we'd be interested in, and hopefully something that they'd bring to our attention. I'd like to be able to have our log monitoring software watch for attempts at common exploits and automatically block them, but it doesn't do that yet. Which is one reason why it's still not ready for launch yet. :-)
The PHP files where the content and the system are one and the same (hand written pages not using a packaged CMS) aren't part of "the vast majority of hacks" category. Compared to exploiting a WordPress vulnerability in 50+ million installs, someone trying to mess with the black box that is someone's custom written page happens insignificantly rarely. Your retort doesn't hold water.
Speaking from experience, this is simply not true. There are automated scanners in the wild which will attempt to detect and exploit common vulnerabilities in simple PHP templating systems and CMSes. One frequently exploited vulnerability is in applications which use URLs of the form:
index.php?page=foobar
With supporting code along the lines of: $page = $_GET["page"]; /* if register_globals isn't set */
include("pages/$page.html");
Until relatively recently, when PHP started rejecting filenames with embedded null bytes, code like this was vulnerable to input such as: index.php?page=../../../../../../proc/self/environ%00
Applications like this are relatively easy to detect in an automated fashion, and were for a time being exploited on a very large scale.The files must be created by a different account. For certain setups this can be problematic, but it's a good idea for most.
How do I, as someone outside of the music industry, gain access to that communication channel?
$ curl -v -A Googlebot example.org
$ curl -v -e www.google.com example.orgThere are certainly differences between just setting the user agent and running Fetch as Googlebot. (The incoming IP address being an obvious one.)
https://www.google.com/search?num=100&hl=en&safe=off...
(In this case, I specifically put in "For Sale" to highlight the spammy drug ads, but they come up even without this).
Puzzle: I scan through looking for pages with "XYZ For Sale" in the title and then check out Google's cached version of the page. Sometimes, I see the spam in the cache, but often enough I don't.
So: how is it that the search result is different than google's own cache for that page?
Most of the time they (hackers, sorry pronoun overload) just naively check the referrer. Going to Google and searching for the site and clicking it is often sufficient