Google now shows why it ranked a specific search result
searchengineland.com
searchengineland.com
Not sure about something Google themselves made, but I've been using this one called "Unpinterested" for a few years and it works as intended.
https://chrome.google.com/webstore/detail/unpinterested/gefa...
Fuck Pinterest for how they've ruined google image searches.
google.##.g:has(a[href=".pinterest."])
google.*##a[href*=".pinterest."]:nth-ancestor(1)Right now I just have my own custom filters and since I'm in grad school this is taking a much lower priority than I would have liked it to, but I have a goal of making user accounts with customizable filters by the end of the summer.
[0] https://hadal.io
0 - https://addons.mozilla.org/en-US/firefox/addon/ublacklist/
Fuck SEO spammer.
But if you blocked each bad result, I'm not sure there'd be any good left.
For example, I search "a bougainvillea without pruning" and all the results are "how to prune bougainvillea". Try: "How do I search the web without e commerce results?" and all you'll get is "how to do ecommerce better" pages.
I recently caught a phishing site of my Govt. Tax website like that [1].
[1] https://twitter.com/Abishek_Muthian/status/14097614168657469...
Whenever we try to simplify AI decision into an “explanation” we throw out 99.9% of the information. In fact, all we did was make the AI more complicated by requiring that it spit out a “reason” alongside its result.
Why do they find long necks attractive? The answer to that sort of question usually turns out to be "because animals who randomly happened to be a little more attracted to long necks were more likely to breed with the long-necked animals, so their genes became better represented in the gene pool thanks to the long-necked animals' better survival power".
I think the answer is usually actually "because it is indicative of a healthy mate". If you need to be well fed and healthy to grow tall, or to maintain a giant set of ridiculous feathers, then it will probably be an attractive trait a few generations later, since it will predict success.
But the ocean is a big place. Tree tops in the savannah may not be as important resource.
For instance it's long neck traps a lot of air in it's breathing cycle, requiring larger lungs for the same oxygen capacity. It has twice the blood pressure of a human to enable pumping up the long distance, and as such a giant heart. It has a complex circulatory system to control the change of blood pressure at it's brain when it bends over, and handle the high blood pressure in it's legs. And so on and so forth.
Long necks aren't the only possible adaptations to get to taller trees though. For instance, tree climbing is another.
We do see convergent evolution into the "tree" and "crab" body-plans. Dozens of unrelated species have evolved along similar lines, filling similar niches by evolving similar forms.
There were long-necked animals in the past -- a well-known type of dinosaur comes to mind -- but the niche isn't so large or successful as the "eat a bunch of grass" niche.
DNA in nucleus - source code on disk
Messenger RNA - source code in memory
Ribosome - JIT compiler
Protein - tiny running program, think unix utilities. Also, not saved to disk.
There is source code for millions of such programs in the DNA.
The compiler (Ribosome) recognizes any combination of the base constructs of the language (C, G, T, A) as valid, but the compiled program may not work.
But the presence of random mutation is extremely well documented, yes - how else do you explain why identical twin humans don't have identical genomes? A great majority of corruptions have no effect, because biological systems are highly resilient to change; the analogy you want is less "delete a line of code" and more "delete a single machine from AWS". Some corruptions are so bad that they cause the organism to fail completely at the embryo stage. Some corruptions have an effect which is detrimental in one environment but beneficial in another; the whims of nature decide whether such an organism survives. Some corruptions are simply beneficial, and if they're sufficiently beneficial then they take over the gene pool in not many generations.
Multiple mathematical approaches were taken based on myriad theories of information encoding, compression, and redundancy. For a brief period of history, the question of how DNA worked wasn't one of biology, but one of statistics and information theory.
Unfortunately, once they had the technology to observe DNA directly, they discovered that almost all of the theories were just wrong, because the encoding from DNA trigraphs to amino acids is basically nonsense. It's not structured, or organized, or elegant. Like so much of biology, it's a hot mess of accident and coincidence, and massively, inefficiently redundant... But, that redundancy has an interesting property. The most commonly used amino acids in protein formation are encoded by a large number of DNA triplets, and in general, one edit to those triplets results in a triplet that codes for the same amino acid.
In hindsight, it is perhaps not surprising that the encoding schema from DNA to amino acids is edit-resistant.
(What cooks my noodle regarding the redundancy is: do we see the redundant DNA encoding because those amino acids are most commonly used, or are those amino acids most commonly used because a large number of DNA triplets code for them so they are statistically likely to be more prevalent? ;) ).
There’s a math book that also has a catchy title: Mathematics and Sex, by Clio Cresswell
I just listened to interview with Rama Ranganathan about "evolutionary physics".
https://news.uchicago.edu/story/big-brains-podcast-explores-...
I couldn't really follow their discussion. From your post I'm guessing you'd get a lot more out of it.
Your questions about redundant encodings seems related to Ranganathan's assertion that biology can be both efficient and adaptive. Whereas with human design, efficiency and resilience is a more of a tradeoff.
Any who...
Usually, but not always.
Add to that an astounding number of iterations (random point changes), the fact that changes are guaranteed to be ... well, the metaphor breaks down quickly, but it would be something like "syntactically valid in the language being used, and no larger or smaller than a certain size of diff", and a sort of error correcting/weighted (as opposed to absolute/binary) behavior as a result of each change, and it starts to seem a bit more workable.
But even that is barely scratching the surface of how this stuff works. Code is far from a perfect analogy.
I think it's also important to realize that making new stuff, like creating a new type of protein, is very rare in nature. Most evolution is just number tweaking: make the neck or beak a bit longer, increase melanin levels, etc. And the way those "variables" are encodes seems to lend itself to natural variations, which is (one reason) why humans have so different height.
Some mutations are so bad that they are fatal to the organism (don’t compile). Where a gene isn’t that important, the mutation will just confer a fitness change.
this is mediated through a variety of mechanisms. Some mutations just alter the performance of a protein, or the expression dynamics (when, for how long, and how much) of the gene expressed as protein. Many proteins don’t act directly on things but are part of complexes of proteins that perform a function, or are receptors that trigger other things, and thus many genes are “meta” in this way.
Imagine a complex sequence if gene expression during vertebrate development. At one point during elongation, a series of enzymes alter how this process starts and stops. Small variations in the performance and availability of the genes involved in this will nudge the outcome of elongation to result in a longer (or shorter) necked animal. If that is helpful for that individual, the gene will survive and flourish.
This gets more complex as in animals w sexual selection we have multiple versions (alleles) of each gene. There is also a huge realm of epigenetics and modulation (methylation, promoters, antisense, etc).
It’s really stunningly beautiful and amazing.
This is why genetic diversity is important. Populations aren't just more copies of a single genome. They are in fact reservoirs of genetic diversity which makes ecosystems more resilient to change. This partly explains why there are so many species of insects. As the recyclers at the bottom of the foodchain, they are the boots on the ground that do the heavy lifting and are subjected to the first shocks. They evolve into many species not only to specialize, but to guard against rapid changes at the bottom of a food web.
We know the size of genome of an animal, we can estimate the maximum possible number of mutations throughout history. And then calculate if it is possible to create such genome with this number of mutations or not.
For example, if a mutation happens once per K animals, if one successful mutation happens once per L mutations, and there have lived M animals then maximum number of successful mutations is M / K / L. We can compare this number to size of a genome and verify whether it could be created as a result of an evolution.
If a mutation offers an advantage, even a tiny one, over time and generations is will be selected more than the non-mutation.
edit: although some genetic mutations do of course result in organisms that will certainly die, possibly before being born
You might be interested to know how common (naturally) aborted pregnancies and eggs that don't hatch are.
This is why mutations can often not cause immediate "compilation failure" since they just alter a parameter of the environment a specific morphological system is growing in, and the change is usually just noise which the system already has a large buffer of variability it can use to adjust.
Genetic encoding has existed for billions. A lot of time and randomness have driven it to become resilient to changes in a way we struggle to yet even comprehend.
you took the one example there is only one of.
besides it isn't necessarily AI decisions. these are statistical results, e.g. showing you which parameters have strong correlations. (some would argue this is all there is to AI)
but aren't starfish symmetrical either way? in the same way we have a head, but our head is also symmetric to itself. actually they have 5x symmetries :) would this be considered rotational invariance? idk
So the question is rather why so few mammals evolved long necks. Maybe there's something about mammalian physiology that makes long neck mutations usually unviable?
That said, this is not totally unlike many human explanations. In fact, increasingly, if a decision can be explained entirely as a product of decision making policies, it's likely to have been automated already.
Eg. If I ask you why you chose to give this analogy, you will probably be able to provide a plausible explanation. But, that explanation is probably constructed after the fact. What we decide to say or write "comes to us" and we don't really have access to how that works.
This is also true of judge's rulings, for example. They refer to rules (laws, precedents and legal concepts), and rulings must be justifiable in these terms... but the giraffe metaphor still plays.
I suppose this is equivocation in a sense, but I'm not arguing that human and AI decision making are equivalents whether or not they share this core feature.
To which I counter: yes, humans are black boxes too. But they're all similar black boxes. We all share the same brain architecture, the same firmware. We have specialized hardware in our brains to simulate each other, predict other people's reactions. This is what allows us to work in groups and form societies.
Additionally, as societies, we've learned to force predictability on each other. In the past, high-variance individuals - ones that are too unpredictable - were shunned or outright killed. Today, we isolate most severe cases from society entirely, by putting them in prisons or mental hospitals, and isolate less severe cases from high-impact responsibilities - e.g. by gating pilot licenses, driver licenses and gun permits behind various kinds of physical and mental health evaluations.
For the day-to-day needs, humans are very predictable indeed, and their explanations are good enough. We know how to compensate for unreliable introspection. That's not the case with modern machine learning - the way DNN models "think" is completely alien to us. We can't work with this. We also don't have to, those algorithms are our creations. We can make them explainable. And we need to.
That said, I'm not sure that we can make algorithms explainable. We may get better at interrogating them, and the analogy to human "black boxes" may get spookier, but I suspect that they're unexplainable in a "hard" sense. An explainable algorithm would have made different decisions, and probably won't be as good at making them.
For AI, stating that we throw out 99.9% of the information is wrong or misguided depending on interpretation. A lot of people are fixated on ML 101 belief that models are black boxes and impossible to unravel. This really isn't true, and significant progress on explainability tooling has been made in the past several years. Most of which (by which I mean none that I'm aware of) require changing the underlying AI model.
2 examples: 1) The well known 20 questions genie, akinator. Not really ML, but works similarly to decision trees. You have a space of information. The underlying tree model learns what the most relevant questions are to divide the space to optimize entropy for information gain. 100% explainable. Surprisingly accurate for many people. https://en.akinator.com/
2) Shap values for AI models are one example of tooling to make ai models more interpretable. They show you which features mattered for a single prediction, and by how much in which direction. If we imagine a decision tree again, we can assume that this bears a strong correlation to the path followed through the tree and the associated information gain at each step. If we imagine a complex image recognition neural net, we can highlight which pixels in an image contributed to the outcome the most. It wouldn't be terribly surprising if shap is the tool being used by google here specifically. Focusing on a few important features rather than considering every single is ideal.
If I, for example, give you an extensive fecal analysis, psychological, infared, radiological, psychic medium prediction, and survey results from 5 normal people who got to look at the animal- and I asked you to identify if it was a cow or not... It would be appropriate to just say "survey said cow, so its a cow".
So it shouldn't be difficult to explain why the site is ranked higher. But I don't think Google will ever disclose that secrets.
My point was that an AI giving an explanation for its output will almost certainly not explain how the model was generated or trained, because that is the secret sauce. It can only give an explanation from within its own domain of knowledge, unlike a human who can make an informed decision based on facts and then explain it from a higher-level viewpoint, or from the viewpoint of the person asking for the explanation.
For a concrete example, imagine an AI for determining if a given video is of someone shoplifting. I give it a video and it says “yes, suspicious hand movements (78%), avoiding staff (65%), change of clothing (45%).” Here the AI has said why it thinks shoplifting is occurring, and a savvy user could guess what types of videos the AI was trained on.
Buuuut, we are left wondering all sorts of important questions. How was the training data collected? Did it represent a diverse set of shoplifters and patrons? What movements are considered suspicious? What constitutes avoidance? What is the false positive rate of the AI? What was the motive of the people creating the AI…
So to go back to my initial analogy, The AI can explain its internal process of coming to a decision, but it cannot explain the human environment in which it was created. In the same way, we can use science to construct useful narratives for “why” specific features of an animal fit with its environment, but we fail miserably to explain “why” the entire environment came to be.
What I've seen a alot of the last few years is sites with simply copied content republished on new domains. And I don't mean just some text, I mean the whole website.
Take this search for example: https://www.google.no/search?q=nrk+site%3A.it
It show's a bunch of Italian domains with mostly spam based on Norwegian content. This content regularly ranks high when searching for Norwegian cotent.
Overall google result quality declined a lot from my perspective (and colleagues also agree) in the past years.
the info-boxes and sidebars are often just plain wrong (wrong product, wrong answers) - have to add quotes to every other query but these also don't seem to work - websites that used to be indexed disappear and I get lot's of noise and outright spam websites with redirection spam (I've tested on Linux with different browsers, unlikely that it's hijacked).
Not sure what's going on - I guess they made a lot of progress in NLP with BERT and other models and nobody internally want to see the drawbacks?
reminds me of msdn, where I always search in english through some search engine, then get the german version presented with a SMALL button to turn it to english.
From what I've heard, the biggest problem with building a general-purpose search engine seems to be the index—indexing the entire web requires data centers many times larger than most can afford, and deciding what is relevant in such an index requires excessive training data. But while working, I typically only want my search engines to search the same 200-odd sites (GitHub issues, Stack Overflow, MDN, language docs, high-quality technical blogs, etc.) It's too many to know which "site:" query to use, but it's a small enough set that it should be a manageable problem for an upstart search engine.
It also seems like this is likely true in many domains. Taking on Google directly by trying to build a general purpose search engine seems to be a non-starter (most competitors just crib off of an existing player, and don't do much better). But a bunch of sharded search engines for specific domains could theoretically outperform Google/DDG/Bing within their respective domains, because the human using them has already narrowed the domain down dramatically. They wouldn't dethrone Google (general search is always going to be needed) but they could compete for a segment of traffic.
Until content itself is cryptographically signed, there's no verifying where it come from or how true it is, if it ever were.
At which point you have to determine which signatories to trust. Which is the same issue we currently have - which publishers (ie websites) should the search engine trust (ie rank highly)?
So really what we need is a more robust method of attributing publisher quality. Except that perception of quality varies drastically between different consumers. Also, most solutions you might propose quickly start looking an awful lot like a distributed web of trust but no one has ever managed to pull that off at scale so far.
Sad to see NRK being abused like this tho
I've seen this rotating around different country-level TLDs, .de & .bg spam domains were filling the results previously.
I suspect some big-name spammer rents domains for a short period while they're pending expiration.
A company perhaps also involved in some shady prescription med domains:
https://www.fda.gov/consumers/hosting-concepts-bv-dba-openpr...
https://www.fda.gov/consumers/health-fraud-scams/hosting-con...
Hacker News Post: https://news.ycombinator.com/item?id=27928454
Voting would likely require being logged in which is not so good for the more privacy orientated search engines.
Could flip the idea on its head and look at ways to verify the owner of a site/page (again, with privacy issues associated), and use that as a quality metric.
To what end? A verified real identity could post crap. Meanwhile an anonymous actor could publish gems. It comes down to how much trust you place in a given publisher; that currently goes by domain. The only thing verified identity buys you is the ability to trust snippets from that person published across a variety of untrusted domains.
Meanwhile, analyzing per-account feedback patterns along with other session behavior significantly raises the bar for abuse.
The general idea would be that people are harder to fake as a signal than links.
A lot of sites associate social profiles with their domain/page. Something like a topical author Pagerank would be interesting, with endorsements. That too, of course could be bought and gamed. Pagerank was a game changer when it arrived, essentially acting as a 'vote' for a page, where some votes are more important than others. Associating votes with people (like the OP thought) is interesting.
I don't think it would be as hard as you think. There are already "captcha farms" in India and Bangladesh - "Identity farms" would be just as simple to set up.
You would have 4 options for each item, ( Upvoted, Downvoted, Not Voted, and Blocked). If a SEO manipulator is Upvoting a bunch of spam content that you didn't vote(that you've seen) for, down-voted, or blocked, they would quickly un-pair with you. The next step for SEO is for them to start voting identically and predicting to how you would vote so that a few paid ads didn't skew the results to much, but isn't that mission accomplished[0]?
This isn't done because it's grouping the users and not the content? It doesn't scale? And yet thematically this seems close to the model used by Google FLoC. If Google would just make a sub-reddit like system for each cohort so that we could submit and discuss content, I have a feeling that would be very successful and people would opt-in.
Agree with the edition of news results.
It becomes very hard to find information about specific topics. Most often you find articles with "theory X has ben deboonked!" at the top
White Family image search (hilarious)
Hillary + anything autocomplete - Google shows nothing negative
various certain outlets article titles verbatim don't appear
Google bias has been on display for quite some time. And Yes Dang, I'd be honored to show evidence if you give me a forum.
There's ongoing antitrust lawsuits about them pushing their own services (e.g. Docs) in front of competitors, and some fuckery with Shopping. Their defense leans heavily on the statement that they didn't do it on purpose or manually, but that their algorithm considers their services the best. For that they need to lift the veil a bit on their algorithms (and iirc there is an outstanding court order for that one already).
To me it seems a rehash of this years old controversy:
https://www.tweettabs.com/googling-european-people-history-t...
Thanks for the article on that!
One of the first results I get on Baidu is the cover of a book called European Art and the Wider World: 1450-1550 [1]. It shows what I assume are the three kings/magi presenting baby Jesus with gifts. Each of them has a different skin color and style of dress, demonstrating how European art portrayed people from the wider world. I don't think a Chinese search company has any particular motivation to try to influence people's perception of European racial makeup, so I would guess that these results occur naturally because the phrase "European Art" is mainly used to contrast with non-European art or people, as in the book cover. When people are writing about European art in a generic context, they'll usually be more specific and say "French art" or "Dutch art" or whatever rather than "European art," so that phrase is less likely to associated with images of European art that aren't showing such a contrast.
[1] https://m.360buyimg.com/mobilecms/s750x750_jfs/t8641/58/7295...
Whether or not any of the specific instances are Google actually interfering or not the slew of evidence is all in one direction, hence I'm not sure why anyone doubts that these huge media companies are interfering in what we see, huge media companies have always done that, it's just that the newest ones have even more power to do so.
[1] https://reclaimthenet.org/google-removed-churchill-image/