20% of requests for Wikimedia Commons are for one image of a flower
phabricator.wikimedia.org
phabricator.wikimedia.org
Finally, I replaced the image on there with a 'Netscape Now' button. Within 15 minutes the matter was resolved.
Did they continue to link to your software after that? (I'm curious - what was your software?)
http://web.archive.org/web/20000510010712/http://www.camarad...
Hard to believe now, a typical blog post will already pick up 30K visitors without too much trouble.
https://jacquesmattheij.com/story-behind-wwcom-camaradescom/
And apologies for the non-working images.
I didn't see that consequence coming when camarades.com shut down. I really should dig up those images and repair the blog but the todo list isn't really getting any shorter on this end.
> Starting in 1996, Alexa Internet has been donating their crawl data to the Internet Archive. Flowing in every day, these data are added to the Wayback Machine after an embargo period.
I invited my parents to come see what I had done, and somehow typed the website wrong and ended up on a spanish-language porn site. I could not hit the back button fast enough. Possibly one of the most embarrassing memories of my childhood.
I have no idea what my parents thought I was up to.
One of the most popular cams for years was an old person that was extremely ill and that rarely moved but he had pretty big fanclub and he thought it was quite funny that he was more famous on what eventually became his deathbed than he had ever been while he was still active. After he died his family asked to remove all the images and close the account which of course we did. Makes you wonder if all those people wishing him well over the years kept him going a bit longer. What is interesting is that if you did this today I'm pretty sure the jerks would drown out the nice people by a considerable margin, of course there were jerks back then as well, but on the whole the internet seemed to be a much nicer place to hang out than it is today.
https://jacquesmattheij.com/sorting-two-metric-tons-of-lego/
When I was a kid I asked my mom to print me out Grand Theft Auto cheats from Gamewinners.com while she was in work.
Somehow I got the address wrong and she wanted to know why I wanted to print out pages and pages from a site dedicated to men cheating on their wives. Got there in the end though and I still have some of those GTA cheats memorised.
https://www.youtube.com/watch?v=BzAdXyPYKQo&ab_channel=yate5...
https://silicon-valley.fandom.com/wiki/Russ_Hanneman
I'm just glad I didn't turn out like Erlich Bachman! (OR DID I???!)
https://www.reddit.com/r/SiliconValleyHBO/comments/4jmlv9/wh...
Edit: grammar.
...because attacks from China are horrified at the thought of disrupting Falun Gong?
It won't stop a hacker who is probably bypassing parts of that anyway, but the more casual requests such as those caused by deep linking will generally stop getting through.
So replace it with the pakistani flag to solve the problem (or start WW3)
I was tempted to replace with goatse but I think I just changed it to a screenshot of his website saying that it was illegal to deep link.
It soon got changed.
Instead of just outright replacing the image, I set up rules in Apache to check the referer, and if it was our site, serve the correct image. Anyone else, it served up something...questionable.
Problem solved.
Ebay didn't care. So obviously the only option was to create a script that would randomly change the images to something unpleasant (think early 2000s rotten internet content).
Good times.
See, for example, their statistics at https://grafana.wikimedia.org/d/000000102/production-logging...
In large part because 99% (+/-) of their traffic is read only. While Facebook and Google have to do heavy workloads for every click and action taken on their services, Wikimedia can cache basically everything. Allowing them to operate on a tiny fraction of the number of machines (and infrastructure) that the rest of the players do.
The contempt on here is crazy sometimes.
Nobody in this thread is saying that. Parent to you said:
> they could just stop being absurd instead [of building more DCs]
implying FB could build fewer DCs by scaling down some of their per-page complexity/"absurdity". Basically saying their needs are artificial or borne of requirements that aren't.
> conglomerate entities are fundamentally opposed to my right to privacy
That's a common view, but it's not on topic to this thread. This thread is mostly about the tech itself and how WikiMedia scales versus how the bigger techs scale. It has an interesting diversion into some of the reasons why their scaling needs are different.
You could instead continue the thread stating that they could save a lot of money and complexity while also tearing down some of their reputation for being slow and privacy-hostile by removing some of the very features these DCs support (perhaps) without ruining the net bottom line.
This continues the thread and allows the conversation to continue to what the ROI actually is on the sort of complexity that benefits the company but not the user.
Let’s suppose the Facebook cluster spends the equivalent of 1 full second of 1 full CPU core per request. That’s a lot of processing power and for most small scale architectures likely adding wildly unacceptable latency per page load. Further, as small scale traffic is very spiky even low traffic sites would be expensive to host making it a ludicrous amount of processing power.
However, Google has enough traffic to smooth things out, it’s splitting that across multiple of computers and much of it is after the request so latency isn’t an issue, and it isn’t paying retail so processing power is little more than just hardware costs and electricity. Estimate the rough order of magnitude their paying for 1 second of 1 core per request and it’s cheap enough to be a rounding error.
So it wouldn't necessarily be contradictory if most of their core functionality could be replicated very simply, yet the actual product is immensely complicated. I forget where I first read this point, but probably on HN.
Edit: I don’t know what I’m talking about. Happy Monday!
Also, there's way more than just the web tier out there.
Everyone, please listen to Rachel and never ever me.
It hasn’t been this way since around 2013 but again I am fuzzy on how. I think that’s when most such data was switched to TAO, which has local read what you wrote consistency. As long as users landed in the same cluster (and thus TAO cluster) what they wrote was visible to them, even if the DB write hadn’t yet replicated to their region.
FlightTracker postdates my time at FB (ended 2018ish) so I’m not sure how that is used. These systems evolved a lot over time as requirements changed.
I don’t remember anything about writes being batched in memcached and merged in on page load.
Do you have a source for this?
I'd be willing to bet that the ONLY reason why 15% of their daily queries "haven't been seen before" is because they add un-needed complexity like fingerprinting. You're making it seem like they've never seen a query for "cute animals" before when obviously they have. They choose to do a lot of extra leg work because of who you are.
So your claim that 15% of their queries have "never been seen before" is probably inaccurate. I'd be willing to bet that "15% of their queries are unique because of the user, location, or other external factor separate from the query itself."
They've seen your query before. They've just never seen you make this query from this device on this side of town before.
15% of the queries themselves are unique. https://blog.google/products/search/our-latest-quality-impro...
https://www.google.com/search/howsearchworks/responses/
I work for Google (and used to work on Search).
Internet users who came online later, from GenZ to many boomers, will often just write conversational sentences and questions.
I know many folks IRL who work at big tech who have no interest in posting here because the community comes off as very unwelcoming. That’s a shame, because they have insight that would be great to hear. Regardless of anyone’s opinion of their employer.
Apologies in advance if your intent was purely about the topic. I just thought I read something in your tone that might hinder discourse rather than encourage it. I wanted to point it out, in case it was unintentional.
On the other hand, if I'm not responding, it's not because I find HN too abrasive — it's because I am afraid of leaking non-public information. That's why whenever I talk about Google, I try to cite a Google blog post or other authoritative source, or talk about my own personal experience; hence, "I rarely search for the same query twice."
I didn't dismiss his argument; I said that he was correct right after he posted: https://news.ycombinator.com/item?id=26073488
"That's not the point" can be interpreted as respectful, but it also can be interpreted as argumentative. I chose to assume good intentions, but I offered a different phrasing that would have a higher chance of not being misinterpreted: i.e. using "yes and" instead of "no but": https://www.theheretic.org/2017/yes-and-vs-no-but/
Having said that, I do think this is clearly a sensitive issues, not a purely technical one. I can appreciate the nuance of working for Google and doing excellent work while seeing the company criticized left and right for its business model. I think given the community, while there is opposition to how Google may at certain points conduct itself as a corporation, there is no lack of respect for any individual working there. I certainly view my comment and the discussion of privacy as having 50% to do with Google’s strategy and 50% to do with the technical aspects of whether you can build a search engine that holds user privacy as a core priority rather than trying to launch an ad hominem on you or anyone. And I saw your other comment that agreed with me and the GP comment so I think my first sentence aside, we are on the same page :)
I doubt you store the history of all searches ever? People don't need a google account to query the engine, others disable history, etc.
Are you saying you still have all searches ever made ever? Because you would need this to say a query hasn't been made before wouldn't you?
"There are trillions of searches on Google every year. In fact, 15 percent of searches we see every day are new"
Does it mean literally the text string typed into the box by the user is new?
Or does it mean the text string combined with a bunch of other inferred parameters we don’t know about is new?
I guess that could be the case. Many could be related to things that are on the news. Like, 'the cw powerpuff girls' for the new show that was announced. No one was searching for that until the announcement, probably
Alternatively, if we assume that google has already recorded 20 Trillion unique search queries (~ 1 Trillion new ones per year for 20 years), the odds that a query composed of 3 correctly spelled english words that are not names has been seen before is 1.6%. Even if we restrict queries to those using the most common 1000 words, there's a 50/50 chance of a query composed of 4 words being unique.
Of course people do not just type random words into the search bar and some terms will be searched many thousands or even millions of times, but still if anything the fact that 85% of searches aren't unique seems surprising.
They briefly mention the statistic in the last paragraph.
Also, speaking of people behaviour, it would not make sense to search everyday for "cute animals", but the volume of searches done for new things people discover as they get older would make more sense. I mean just look at search trends for things like "hydroxychloroquine" for example (and that's not to mention people who get it wrong, i.e. other factors for differing search queries too)
Also, other languages can change the queries depending on how you phrase the sentence too. Add to that the people using other ways to search instead of just visiting google.com and I think you can get pretty close to 10%.
If fingerprinting is the reason, 15% would be a figure too low I surmise. Would that be the case I think that would make probably 20-25% of searches rather than 15%.
It could very well be that they do classify fingerprinted search differently only in some countries and not others? That would/might explain the 15% figure.
I might be wrong and under-estimated fingerprinting techniques for Google. If they have really good fingerprinting techniques, that would reduce the estimate I have in mind to a better number (close to 15, maybe?)
Nobody has ever searched for hydroxychloroquine before today. Today is the day the word is hypothetically invented. Today 2 million people will search for hydroxychloroquine. But only one of them was the first to do it.
What I know about pop-culture and viral internet culture is telling me that 15% of 1 trillion searches being unique is shady math.
So I am not fully convinced that the 15% claim is completely transparent.
The folks who use keyword-based searches are largely those who got on the Internet before ~2007. Tech-savvy, relatively well-off, usually Millenial or Gen-X, plugged into trends. This happens to be the demographic dominant at Hacker News. But there's a much larger demographic who just types in whatever they're thinking of, in natural language, and expects to get answers.
Come to think of it, this is also the demographic that doesn't use tabbed browsing, and uses whichever browser ships with their OEM, and often doesn't realize that there's a separate program called a "browser" running when they click on the "Internet", and issues a Google Search for [google] (#3 query in 2010) when they want to get to Google even though they're on Google already but don't realize it, and doesn't know what a URL is. When a big-tech company makes a brain-dead usability decision you don't like, first consider how that usability choice might appear to your grandmother and it might not seem so brain-dead.
I'm not sure, on my productive days maybe >50% of my Google searches are not very cachable. (for example, I just googled "htop namespace", "htop novel bytes", "htop pss", "htop nightly build ubuntu 14.04")
What can be interesting, I think, is that you have a completely open infrastructure that has to solve problems on a global traffic scale.
If people are interested in knowing more, I suggest you also take a peek at the wikimedia techblog, specifically to the SRE category https://techblog.wikimedia.org/category/site-reliability-eng... and the performance one https://techblog.wikimedia.org/category/performance/
They need it to be a certain way in order to operate. The limitations and advantages of how software gets made. Why it gets made. The way the software works. How and why product decisions were made over the last 2 decades. What resources they have/had available. It's all a totally different game. Not surprising that different soil and a different climate grow different plants.
One of Google's early coup d'etats, when they were a strategic step ahead of the boomers, was bankrolling gmail, youtube and such. Gmail offered free giant inboxes. They got all the customers. This cost billions (maybe 100s of millions), but storage costs go down every year while the value of ads/data/lock-in and such go up every year. Similar logic for youtube. (1) Buy a leading video-sharing site; (2)bankroll HD streaming because you have the deepest pockets (3) Own online free TV entirely.
That's who Google is, good or bad. How funding works. What products get built. What infrastructure is necessary, possible, affordable. All interlinked. Wikipedia & Google were founded at the same time. Within 5 years (circa 2006) Google was buying charters and fiefdoms. Wikimedia, meanwhile, was starting to take flak for raising 3 or 4 million in donations.
It's kinda crazy that Wikipedia is comparable in scale to FAANGs when you consider these disparities.
I recall reading HackerNews used to have that problem, unsure if it still does.
The way he structured wikipedia, from back-end infrastructure to ownership/governance structure was just the logical way of doing the project. Times were different. Online culture was different.
I don't want to overinterpret the man, or put words in his mouth... but... I got the impression that Wales thinks that if he was starting Wikipedia now, he'd just do it asd a startup and also succeed.
To me, this is almost sad. Besides being an awesome encyclopedia, wikipedia is existence proof for something of scale outside the norm. Something that isn't a corporation. A lot of things are deterministic to the structure of an organization.
For example, take the current postpostmodern war over truth and stuff: platforming/deplatforming, freedom of speech, censorship, bias, manipulation, narrative = power issues, etc. Wikipedia is at the very centre all these problems. Whatever difficulties Twitter is experiencing should be 100X worse for wikipedia. Meanwhile, Wikipedia is withstanding far better, and with far more integrity. I don't think this is a coincidence.
Dunking on wikipedia's budget/spending is popular. Meanwhile, Wikipedia uses <1% of the resources/budget of Twitter. They are operating @ >100X efficiency compared to a realistic for-profit equivalent. That's a flying shuttle.
We know that Wikipedia, Linux & The Worldwide Web are possible because they exist. We literally wouldn't know otherwise. Theory couldn't have gotten us to this knowledge. Each is existence proof for other ways of doing things. They aren't necessarily roadmaps, but I'm a big believer in existence proofs. What Jimmy made is 100X better, more important and non-inevtiable than what Zuck made. The thought that he wants to be Zuck bums me out.
In terms of financial and organisational success it would probably largely beat what it is now. It terms of benefit to humanity, it would be much worse.
Company + for profit + laws means access to information has to be much more tailored to the laws of each place. "Let's remove tianamen's article or lose your chinese license" kind of things.
I'm for one am glad for the current wikipedia we have, despite it's numerous flaws. I still donate every year, although I wish Wales could stop having it spend its money the same a startup or FAANG does.
Stackoverflow is a decent example. Very capable founding team. They explicitly tried to be like a commercial wikimedia. They do embrace quite a lot of openness, notably creative commons... learning from wikimedia successes.
RE "I wish Wales would:" Another consequence for how wikipedia is structured is that Wales isn't the Zuckerberg of Wikimedia. Power is a lot more dispersed.
RE spending/flaws and such: I feel like wikimedia is held to an extremely unfair standard. Who/what should we compare them to?
Wikimedia spend $70m per year. This is probably less than Quora or stackexchange. FB & Twitter (IMO more comparable in terms of scale/importance) spend $55bn & $3bn. Twitter spends 45X more than Wikimedia. Facebook spends almost 1,000X compared to Wikimedia. The bang-for-buck is insane.
Also in terms of flaws in rules/judgement calls. A lot of people are highly critical of wikipedia's "deletionism" related MOs. What articles/edits stay in. How good the rules & procedures are for this. What "camp" has power, and how they treat the other camp. I get that this stuff is contentious.
Meanwhile on Twitter or Facebook, the rule is "I decide." "But it gets us clicks" is the killer argument. Nothing is transparent. Wikimedia is doing a much better job, respecting user & editor rights far more, being a lot less self righteous. Of course it's not perfect, but come on. The "norm" is Facebook's content policy, Twitter's safety department, or Apple's App store approval room. Wikimedia is the one example of being better than that... and for that everyone is always yelling at them.
Then again, WikiTribune was a for-profit.
Hi all, I've been doing a bit of research into possible apps that could be causing this and found two potential culprits that I am currently investigating.
The first is Mitron TV, an Indian TikTok alternative which was made available again on the app store June 6th (https://indianexpress.com/article/technology/tech-news-techn...).
The second is Say Namaste, an Indian Zoom alternative which was launched on the app stores June 9th (https://indianexpress.com/article/technology/tech-news-techn...).
Both fall into the timeline of huge increases, have millions of users and may be using '1280px-AsterNovi-belgii-flower-1mb.jpg' to check the users internet connection - especially for Say Namaste to ensure video connectivity. I've reached out to some developers at both companies and will report back. Let me know your thoughts.
EDIT: I have also noticed the dates match the reopening after lockdown for the whole of India: "This first phase of reopening was termed as "Unlock 1.0"[13] and permitted shopping malls, religious places, hotels and restaurants to reopen from *8 June*." (https://en.wikipedia.org/wiki/COVID-19_lockdown_in_India#Unl... )
Tom
Looks like they posted shortly after yours on the ticket that they found the culprit. Guess we'll find out tomorrow if we were on the right path.
I don't know much about Android development or APKs but it's not exactly "reversing." from what I understand the profile/debug converts the .dex files from the APK to .smali which is human readable.
>Ravn is your portal to the most private messenger as well as Korrax our proprietary token. Stay up to date with Korrax and other Cryptos and join the crypto group chats.
>Messages, images and docs are never stored on a server (after delivery), they’re only locally stored on your own phone. Ravn is not tied to your phone number or email, you only sign up with a username that isn’t searchable or discoverable.
Simpler times.
For almost a week it remained the most requested image the post on which it appeared, the most popular.
It did make me uncomfortable, though. Fearing that my rankings would plummet or so.
My takeaway is nothing new: there are weirdo's online.
My blog did have the word "Peanutbutter" on one or two posts. And the word "sex" on another. Maybe at some point both words showed up close to one another when experimenting with some "random articles" sidebar or some "you may also like" list.
I'm sure there is a more enlightened fix.
it's not theft if you leave it out for everyone to use.
But you can totally sue anyone for anything, and that makes for entertaining headlines - even though if plaintiff lost promptly
It is possible this principle applies to other countries and other things than pools.
I wasn't sure if it is the same or similar principle in Russia or a different one that requires active care for a burglar. Unlabeled chemicals causing liability for a burglar seems extreme to me
However my favourite example is the law that allows any bee keeper to enter any private property if they are pursuing fleeing bee swarm.
This is a pretty puzzling idea to me. How could linking something be theft?
To explore this, I shall try a metaphor. Imagine you're on a big social media website (lets call it Programmer Olds) which has an oddity in that 99% of its users use adblock. You then post a link to another small (ad supported) website on your Programmer Olds page, causing a large number of people to click through and download the page using large amounts of bandwidth (for no monetary gain to the site) and possible DDOSing the site.
Have you commited theft?
The difference here is that while a lot of users use adblock, there are some that don't. These users can still see the ads. Additionally even though it's a small website, it may lead to new readers that stick around or the content itself may even be sponsored.
The equivilent to hot linking a picture would be like taking the content of a blog post without really linking to the source, because there's no chance of conversions there. If you're linking to the site itself then there's a reasonable chance that users can convert.
So I suggest that it's theft just because the chances of readers being converted is nil while you're using their bandwidth.
That's because you're responding to an entirely different issue. "Hotlinking" isn't linking to something, it's including a resource that is hosted elsewhere. It's putting <img src="https://concordDance.whatever/images/big_image.jpg"> on my website without asking you. Now if my site ends up on the front page of HN, that could cause a lot of traffic to your site, potentially overwhelming your server or increasing your hosting bill. It's not nice, and rightfully frowned upon.
In both cases the site loses bandwidth for no gain due to your actions.
If I tell the customer they can go next door to get a panini, I'm not stealing anything. Maybe that restaurant is packed right now and they'ed rather not have an extra customer, but there is a reasonable expectation that they would generally want customers or at least have a means of turning away unwanted customers otherwise.
On the other hand if I break into my neighbor's restaurant, make a panini, then bring it back to my restaurant to serve and make money off of, all without permission from the neighbor, I am most definitely stealing. Even if I doubt the neighbor will mind because he let me come over and make myself a panini once, I can't unilaterally act off that assumption.
Of course when someone doesn't want us to hotlink to their assets then don't do it.
[0] https://commons.wikimedia.org/wiki/Commons:Reusing_content_o...
Or maybe wikipedia is already mostly static.
also, I wonder if HN is inadvertently ddos'ing the ticket system ?
1. https://www.jwz.org NSFW!
I think browsers did drop the path from it at least.
I saw the described image but after I visited the site directly I couldn't see it any more when redirectly via hacker news. Saw it again when I opened an incognito tab.
other good ones were about:1994 and about:mozilla
hey, about:mozilla still works in firefox
https://www.google.com/search?q=firefox+gran+paradiso+robot&...
1. Is there a collective biological term for scrotum and it's contents that is not general like "genitals" is?
Sorry to ruin the fun y’all but there’s images I won’t even mention that I can’t unsee and make me feel seriously ill when I do see them. I don’t want anyone else to feel that way without warning.
The Wikipedia page for https://en.wikipedia.org/wiki/Goatse.cx is text only and without any ascii art.
I'm amused that https://simple.wikipedia.org/wiki/Goatse.cx also exists.
I remember when I was about 15, before pop-up blockers were really a thing, someone sent me a link to that and it would keep opening popups with that image and you couldn't close all of them :-/
Sometimes people look back to the internet of the 90s with too rose-coloured glasses IMO.
I am personally most amused by #38659
Edit: this is a bit unfair, if its a specific app they should be convinced to cache just to avoid unfair resource usage, but hotlinking in general should not be seen as a problem
It's a waste of donors money if someone is using this image as some kind of "is this thing on" test using hacked computers...
But this goes beyond that - it's some blind check of internet connectivity for the app, and doesn't get shown to the user. We're pretty sure of that, given that with the amount of noise that task generated, if there was an app featuring that image at least one of the ~ 90M daily "views" would've been someone reading these posts.
Now, given we want to be nice, we didn't just blindly block the traffic, although making requests without user-agent is against our UA policy https://meta.wikimedia.org/wiki/User-Agent_policy
Things were done very differently back in the day. This problem would have been fixed real quick.
Today I'm sure it would be fine; instead I'm frustrated by my inability to create webp images larger than 16000 pixels tall (i was trying to write a data-saver proxy for reading webtoons)
Why would anyone do such stuff is, as usual, beyond me...
PS. "First!"
com.app.rcn/smali/com/app/rcn/utils/InternetSpeedCalculator.smali: "hxxps://upload.wikimedia.org/wikipedia/commons/1/16/AsterNovi-belgii-flower-1mb.jpg"
edit: defanged the link to maybe save the wikimedia team some bytes
Per the comments, right now the top suspect seems to be the app "Josh" or another TikTok clone because of how traffic surged immediately after the TikTon ban:
https://community.ntppool.org/t/recent-ntp-pool-traffic-incr...
Only 10^-6 of the "legitimate" requests would be affected, but a whole lot of the "undesireable" requests would see it...
There was a better idea posted in comments - serve a picture with a very short explanation and an email to contact.
Just don't give random unsuspecting people blinking images as a rule.
Same advice I gave a w3c.org admin who was lamenting how much traffic people generate by not caching xml schemas. Yes, you have to serve the requests. But you don't have to try to serve them in 100 ms. If a human is on the other end, 1-2 seconds is just fine. If a human is not, then the human will surely notice when their batch process goes from 3 minutes to 10 minutes because it fetches the same schema 200 times.
I guess a couple seconds won't matter unless the server is already redlining it and the tarpitted traffic is a small proportion.
If you care about the traffic because you're already having trouble with that many simultaneous requests, then you are definitely not going to solve that problem by increasing the response time by a factor of 10.
But an important property of reverse proxies is that once the proxy sees the last byte of the response, the originating server is no longer involved in the transaction. The proxy server is stuck ferrying bits over a slow connection, and hopefully is designed for that sort of work load. If the payload is a static file, as it is in both of these cases, then it should be cheap for the server to retrieve them.
Worst case, if you start running out of sockets because you're sleeping, sample the socket count once a second and adjust sleep time to avoid hitting the cap. Also, you could use that sampling to drive decisions about keeping http sockets open or closed.
I should add, select on millions of sockets is going to suck; so you'll need kqueue/epoll/whatever your kernel select but better interface is.
Sample code from Stack Overflow being used by some major app is the most likely candidate. It's also possible that the image fetch call is a vestigial appendix that doesn't even display the image, which will make tracking this down extra challenging.
https://www.wsj.com/articles/the-internet-is-filling-up-beca...
[0]: https://upload.wikimedia.org/wikipedia/commons/thumb/1/16/As...
A few thousand requests from clearly identifiable as coming from browsers and with a referer header from news.ycombinator.com would not exactly interfere with this and in the grand scheme of things isn't a huge burden in terms of network traffic.
I think it ended up being a sort of mobile-based botnet with a bizarre target, which luckily was deduced from some of the headers sent (they all had a random common header).
I'd bet that this is the flower of the week for them.
edit: queries?
Looks like we will know soon.
https://newshimalaya.com/2021/02/09/%E2%9A%93-t273741-invest...
I was sure I'd seen this website before, and sure enough, it's scraping and rehosting almost everything that's posted on HN...
That being said, though the image wasn't hotlinked directly, they expressed concerns of DDOS and the possible costs the Foundation has to incur from each load (they even pointed out that it's "fair and reasonable" to point donation link to them).
I would be interested to see how the licensing issue will be handled, though. The photographer licensed this photo as GFDL/CC BY-SA 3.0 [2], and hotlinking may break the term of these licenses.
1: https://www.mediawiki.org/wiki/InstantCommons
2: https://commons.wikimedia.org/wiki/File:AsterNovi-belgii-flo...
You could even serve another image in its place to this UA, with some text and an email address to contact. You'd probably find out pretty quickly what it is from users of that mysterious thing. A throwaway email address is probably best
Really good idea :-)
Why on Earth would you do anything like that, instead of just renaming the image so the URL stops working (and banning the old URL unconditionally?)
Most indian ISPs, even mobile ones are extremely cheap to not matter a 1mb.
If it's a speed test, they will eventually use another image.
(On a 500MB/mo plan you start noticing)
I guess we will find out tomorrow.
Ironic. Dang could save himself from spam, but not others.
I suppose the contents of the image/"message" could change every day, but presumably that would be very obvious in the edit history of that file[0], unless Wikimedia Commons were suppressing the fact that the file is constantly changing. If they are part of the conspiracy, though, you'd think they would have taken down the task from Phabricator too.
[0] https://commons.wikimedia.org/wiki/File:AsterNovi-belgii-flo...
https://pageviews.toolforge.org/mediaviews/?project=commons....
Also: What's your vector, Victor?