Robots.txt for the NYT has a specific exclusion for an 1996 news article
twitter.com
twitter.com
Edit: newspaper -> newspaper editor
Old, to the point, websites hand-written in notepad.exe are rarely in the top 5-10, even when they have precisely the answer you’re looking for.
Not sure if it's Google discriminating against old sites, or if the new sites just have such incredible levels of SEO-fu that a static, just-contains-what-you-want site has no hope of getting selected for results
Google does this to make every tech worker not working for Google, and thus Google's competition, less productive.
</outlandish conspiracy theory>
Month after month, year after year, search results get worse, and harder to sift through. At this point, google is mostly a way to find wiki articles despite typos (because Wikipedia’s search is bad and can’t handle typos).
I don't know off-hand if it's possible with the search (e.g. it might use a last-selected version cookie) but we should probably be using version-specific 'bangs'. `!pg12` or whatever, in my case.
Information rots over time, decaying from true to false. Some facts are eternal, many aren't.
Worse if you are looking for Dave Smith. You'll get David Smith, because Dave=David, even though you know Dave never identifies as that.
Then add it fixing spelling errors. Which often are not. Rare term? Must be a spelling error!
Part of this, is because most people are non-precise, and also because Google wants voice input to work well. So there, their, they're are the same thing, but can also synonym to things like them, and they.
Point is, Google is trashy for any search not involving cats, or explosions. Quotes barely work, so the only way to mitigate this a bit, as google removed the + search modifier a decade ago, is always, use verbatim, under search tools.
And hope Google hasn't broken it that day. Which they do often.
The newspaper doesn't have to do anything.
Something that Wikipedia reminded me, rtf is obligatory only in the EU and search engines are not obliged to comply with it on its international websites.
Since then it seems he really has gotten it together, and even had a project show up on the first page here. So now I suppose the robots.txt entry doesn't matter much, but it's still there anyway.
To reliably exclude a URL from indexing, you have to serve a “no index” instruction with that URL, either in a meta tag or an HTTP header. And for this instruction to be read, the robot has to visit that page! So disallowing the URL in robots.txt can actually be counterproductive to de-indexing it.
Google also offers a tool specifically for removing URLs from their index in Search Console.
that's pretty messed up.
Source: https://www.nycourts.gov/courthelp/criminal/sealedGoodResult...
edit: a word
https://www.wired.com/robots.txt
It's concerning someone working for the news doesn't understand why.
This is their power of ruining lives forever.
This is a new thing, it should be taught, it's not hard to understand.
I don't know why journalist think they deserve respect when these things are not fundamentally in their ethos.
There should be a better process than robot.txt and some news sites are doing better. Europe has brought in laws. But if journalist want to be thought of as more than writing blog spam, they need a better answer to this.
What is the alternative? I do not know anything about the particular case and make no claims about it but suppose the charges were dropped for political reasons or because the victim was intimidated? Shall we in such scenarios silence them entirely in order to protect a criminal?
Also his name plus his alleged crime gets an AP News article as the top hit.
[1] https://pastebin.com/raw/RE2tpyR3
8<-----------8<-----------8<-----------8<-----------8<-----------
#!/bin/bash
snapshots="20120713050942 20121013154343 20121010165822 20120921054221 20130413152313 20130113162428"
# orig source http://state.gov/robots.txt but also on pastebin in case they delete it:
wget --output-document=robots.txt http://pastebin.com/raw.php?i=RE2tpyR3
for x in `echo $snapshots`
do
for i in `cat ./robots.txt|cut -d ' ' -f2 | tr -d '\15\32'`
do
if [ -e `basename $i` ]; then
echo "$i already fetched"
else
wget https://web.archive.org/web/$x/http://www.state.gov/documents/$i;
fi
done
doneActions based on right to be forgotten don't affect what you're allowed to know or what you tell individuals, they affect what you're allowed to publish to the entire world. Much like publishing a photo of someone; sometimes you need permission.
Would you ask a rapist or their victim for permission to discuss the crime?
RTBF isn't ownership of information about self it is a privelege for villains to censor their victims and the general public.
This is repeated every time that the RTBF is mentioned yet none has ever given an example of RTBF being abused. Can you give an example?
Does raise the q should states publish so much information they collect Swedish tax records, Names and Addresses of accused persons (lots of countries) publishing Mugshots like the US does.
IIRC it was a private person who wanted to remove an article about his house being foreclosed due to his debts. His request to remove the article was not granted since the article was lawful, but the court agreed that google should stop pointing to that article since it was like 15-20 years old and the man had paid of his dept and his house was no longer foreclosed.
What do you find villainous about that?
> Does raise the q should states publish so much information they collect Swedish tax records, Names and Addresses of accused persons (lots of countries) publishing Mugshots like the US does.
The right to be forgotten has nothing to do with publishing records of something. It is about search engines not pointing to something published in the past that can have an adverse effect on somebody today while the published information is irrelevant to anybody today ...
Publishing Names and Addresses of accused persons and their Mugshots is a great example. Publishing that information is relevant since the person is being accused at the time of publishing. A years later the person is found innocent. It is still true that the individual was accused but after being found innocent the information of them being accused isn't particularly relevant to anybody typing their name in a search engine but can have an adverse affect on the individual if he is looking for a job or whatever.
For that very reason a convicted criminal in my country (in the EU) can get a document saying that they are not a convicted criminal after they paid their dues to the society. The point is that people make mistakes and there is no good reason (in the very wast majority of cases) that those mistakes should follow them their whole lives. The point is rehabilitation not eternal punishment. That is a reason why the right to be forgotten exists.
That’s an interesting phrasing. Even you (who seems to support the policy) describe the convicted criminal as a convicted criminal, yet they’re able to get a piece of paper that says they aren’t one.
Does everyone have to produce such a paper routinely in life, or is this a case of “hi, I’m @sokoloff and, even though you didn’t ask, I’d like you to read this paper which says I’m totally not a convicted criminal”?
“A, are you a convicted criminal?” “No.”
“B, are you a convicted criminal?” “No.”
“C, are you a convicted criminal?” “This piece of paper says I’m not.”
They can get that paper after they payed their dues to society (served their time).
> Does everyone have to produce such a paper routinely in life
Not routinely but I don't think anybody goes through their life without ever needing that paper. It is required for a vast variety of government programs and some jobs also require it. Working in IT I had to get it a few times when applying for jobs in private sector.
"Isn't particularly relevant" according to whom? Relevance in the eye of the beholder.
The daily mail claims articles on Josef Fritzl a man who kept his daughter in a dungeon for 24 years and Tory MP Jonathan Djanogly who hired individuals to spy on fellow party members were effected but years after the complaint about them being delisted they are now findable on google again.
Google doesn't have a great history of doing the right thing with automated moderation. It's entirely likely that a script could easily deny a request to censor trivialities while covering the crimes of a pedophile without a human being in the loop.
More recently someone convinced google to delist the url
https://www.techdirt.com/blog/?tag=right+to+be+forgotten
Their list of articles on abuse of the right to be forgotten for maximum irony.
Then there is Thomas Goolnik who is using RTBF to hide his efforts to hide his efforts. That is to say he is using RTBF to delist articles about him misusing the RTBF It seems the start of the chain of forgetting is Goolnik defrauding people out of a million dollars. Something that might be worth knowing if you were participating in any endeavor he was involved in.
Had the victim not been a unique UNION of {phone with battery + full video of scene with audio + harvard grad + well dressed} he would likely have ended up arrested, probably tried, and quite possibly convicted/pleaded, possibly even end up on a sexual predator listing.
Think of how many times this has happened. Now think of how star-crossed lucky the victim here was that he had irrefutable evidence to clear the false accusation.
I hadn't heard that detail before. I can see how someone might legitimately interpret that as a threat.
Imagine being in her shoes: you're peacefully walking your dog in the park when a stranger confronts you, says "I'm going to do what I want, but you're not going to like it", and approaches your dog. Haven't we all seen enough videos to be concerned about what might happen next? From what I can tell, he seems like a good guy, and I don't think anything bad would have happened, but she couldn't know that at the time.
Which all reinforces the larger point: when the Internet reacts to a story, important details are buried and overlooked. Politicians and activists ignore the particulars of the case and shape it to fit their own agenda. The media sensationalizes the story for clicks. An outraged mob dismisses nuance as x-ism or y-phobia. Employers acquiesce to the demands of the mob simply to avoid becoming a target.
The right to be forgotten is important, because Internet justice is rarely just.
Most Americans support right to have some personal info removed from online searches https://www.pewresearch.org/fact-tank/2020/01/27/most-americ...
they prefer to be able to remove information about themselves.
they strongly prefer you don't have the ability to remove information about someone else.
RTBF does not use robots.txt (and then require that it be respected) on the site hosting the thing to be forgotten, it's an exclusion on the search engine side.
If you can’t get NYT to remove it, maybe you “know someone” who can.
The correct thing would be to serve pages with the appropriate HTTP header to disable indexing. Of course search engines are still not obliged to follow the header, just as they are not obliged to follow the robots.txt file, but you are not leaking more information that you need.
Really, robots.txt file is only useful to reduce the load on the server by crawlers, it shouldn't be used as a protective measure!
https://www.wnycstudios.org/podcasts/radiolab/articles/radio...
robots.txt version from 2017: https://www.robots-viewer.com/robots/checksum/8029662cfb040c...