Given the current Australian government's cosy relationship with a particular media company that currently dominates the media landscape here, I don't think it is coincidence.
Given the current Australian government's cosy relationship with a particular media company that currently dominates the media landscape here, I don't think it is coincidence.
Basically, this law would prevent Facebook from deploying just about any non-trivial change to its product without first doing a detailed analysis of how it would affect the Australian news business, in order to determine whether a notification is required.
See sections 52D and 52W of the bill: https://parlinfo.aph.gov.au/parlInfo/download/legislation/bi...
Guidelines intentionally kept vague so that some bureaucrat can slap a huge fine and collect the rent?
I wonder if that rent-seeking attitude will accelerate or curb the current brain drain Australia faces.
And 2020 put pause to any Australian brain drain and given how well we’ve handled the pandemic, is likely to be seen significant increases in net migration.
We were a large consumer of news before Facebook. We will be a large consumer of news after Facebook.
Australia does have competitive federalism, and so many of the localised decisions have been from the states, but the major decisions for seeding - closing borders, acquiring vaccines, etc, along with fiscal backstopping - are federal.
The federal government (under ScoMo) has never polled so well as it has during covid, and for good reason.
(NB: I am no blind Coalition supporter, but they have made decisions which are very popular, and no amount of directing attention to the states would absolve them from blame if we had a situation more like Europe or the US.)
Not really sure why the politicians get the slack. Which measurement was highly effective that could also be done in a country with neighbors.
Ps. I'm not pro Facebook. Just curious
... All of them? I'm not sure why so many people seem to think water is required to make a border effective.
You should assume it's ten or twenty times worse than the figures show, AT LEAST, in any countries that aren't fully transparent and rich enough to test widely.
Compare 477 to 16,000:
https://marginalrevolution.com/marginalrevolution/2021/01/su...
But politics responds much more to outcomes than causes
Honestly, I think that makes sense and it doesn't immediately strike me as a negative for either side. These are articles coming from trusted sources. There's no need to apply the anti-spam parts of the algorithm. News agencies get a more stable algorithm, Google gets to keep their secret sauce.
There is still an advantage to the incumbents. Those carousels are usually in prime real estate. Google would hold the keys to who is in the carousel though, so they could expand it without legislative changes. I like the flexibility, though I don't love handing Google the keys to more kingdoms.
Google and Facebook's algorithms should be required to be publicly disclosed. As a society, we should demand that we are able to see the algorithms that every web property lives and dies based on, that lives are built and destroyed by.
The fact that technology companies have been grossly negligent and irresponsible isn't a reason to not regulate them: It's proof regulation needs to be much, much stronger.
This is quite a bizarre claim as there is famously an entire category of problems that are hard to solve but easy to verify: P vs NP
Tell me, how did your brain come up with what you wrote? How do I validate that it isn't racist, sexist, or slanted towards encouraging violence and harm?
In other words the solution to this should be antitrust enforcement and decentralization of power.
male guest: "now first of all, let me just start by saying I'm not racist..."
female guest: "pfft..."
host: "ah see you made a noise there, but a lot of people accuse him of being a racist, so I think it's very helpful to know that he actually isn't one..."
Not to mention Facebook’s are even more difficult. Tangentially related, remember when you could use “View As” on your profile page to see what your profile looked like to others? It doesn’t work anymore, only works for Public and Yourself; you can no longer choose the person to view as.
It’d be great to test these algorithms. We can’t. They need to be designed and instrumented so this is possible.
I'm not sure a human-readable algorithm exists for ranking all the web pages in the world based on natural language input. In fact, I'm pretty sure such an algorithm does not, and potentially cannot, exist given the absolute failure of all approaches towards NLP that weren't based on absolute masses of text data and complex models.
Are you willing to make Google 10% as effective to achieve your goal of a human-readable algorithm?
Absolutely. If it can't be done responsibly and ethically, perhaps it should not be done.
This generally has worked well. On the other hand, actually attempting to manipulate search results based on automated handling of content is what has given us countless of censorship debates or simply failure where even uncontroversial content is removed or downranked because it violated some sort of strange rule because it had a 'bad word' in it. On Facebook recently clothing ads for the disabled people were banned[1], because turns out the ML system only cared about the wheelchair, not the person in it.
It's actually fairly straight-forward to build recommender systems on transparent, graph-based algorithms and it gives you the added advantage of not discriminating in strange ways.
[1]https://www.nytimes.com/2021/02/11/style/disabled-fashion-fa...
It's trivial to generate webs of fake, inter-related content and use that specifically to feed incoming links to valuable pages. Or to comment-spam websites so aggressively it ruins them. Or all of the secret deals between high-ranking sites to feed links even though the sites weren't related. There are countless examples of black-hat techniques to break PageRank.
I am sorry but you simply can't build a sustainable search engine without deeply understanding the user intent and the meaning behind the indexed pages.
there are also countless of adversarial examples to trick ML algorithms. In fact this is in many ways worse because of the 'idiot savant' character of ML systems, which are almost always oblivious to context and can be tricked in ways that aren't apparent from the design of the system.
In contrast to systems that are legible or even formally verifiable ML systems are entirely unable to provide any guarantees. When someone breaks pagerank at least it's apparent how they broke it. When an ML system mistakes a turtle with a fractal pattern on its shell for a gun nobody knows how to fix the system in any reliable way, other than feed it more data and pray.
One company controls 80% of what is found on the internet. They set rules, restrictions, penalties that are not public. They do not pass any sort of regulatory muster. They rip and tear through businesses standing in their way. They crush out a person's online existence through never explained reasons. They use every advantage they can to tweak a human's emotions, drive and needs to feed more and more advertisements.
You suggest those trying to use every advantage they can to rank higher unscrupulous?
Google's fight to keep search results crisp ended soon after they began selling advertising. Google long ago quit innovating search to be better for people, they've made it better for advertisers.
I agree that you don't need NLP to rank webpages (though it certainly helps), but you do need it to parse the kinds of queries given to search engines these days. The days of logical OR and NOT are long gone I'm afraid.
> It's actually fairly straight-forward to build recommender systems on transparent, graph-based algorithms and it gives you the added advantage of not discriminating in strange ways.
I think other commenters have addressed the PageRank issue, but I'd be super interested in papers doing the work you note above.
All you are doing here is convincing me that tech companies are just runaway trains with nobody at the controls!
Can you explain or understand the algorithms humans use to drive cars?
Explain to me step by step how you walk.
If a company makes a self-driving car and that car then drives badly, surely the response needs to be to incentivise the company to improve their engineering practices, eg, spend more on testing, or require more levels of review of changes, or whatever other organisational changes they need to make safer cars. You don't need to find an individual person responsible to create that incentive. And if you really do want to find an individual responsible it can easily just be the executives of the company (and the executives are probably pretty easy to find even a decade later).
Believe it or not, your car is not that primitive when compared to a self-driving one in terms of the number of things it does autonomously.
Machine learning is very widely used in the sciences and extremely beneficial to humanity in uncountably many ways and assuredly countless more to come. Of course technologies can be used for evil but so can nearly everything that exists. I believe your proposal comes from a desire to help or better the world, but to ban all non-human-readable algorithms is frankly ridiculous and demonstrates a naive understanding of the issue. It sounds a lot like the calls by the U.S. Congress to ban encryption.
Patients don’t care how cancer is detected. Patients care if the diagnosis is correct.
- In medical: your doctor should be responsible for your diagnosis and drug company is responsible for defective drugs, except when they get away with lobbying and hiring good lawyers.
- In physics: I'm not sure if it's as big of a problem as in social networks. But consider this case: If you cannot reproduce the result of an experiment due to a ML model being cryptic, that would lead to huge credibility issue in science.
My suspicion is that the concern with machine learning over racism is rooted in two things. The first is just the general modern trend of accusing anything you don't like of being racist, because everybody hates racism and wants to fight it. And the second is the fear on the part of people who make a living fighting racism that machine learning might actually put them out of a job.
Because machine learning is basically a paperclip optimizer. You tell it to maximize a thing, it maximizes the thing and minimizes everything else. Racism isn't paperclips, so the paperclip optimizer will optimize for smashing it in favor of making more paperclips. And then they're out of business.
Because when you look at the criticism of this stuff, it generally looks like this. ~12% of the population is black, only ~5% of the selected applicants are black, the algorithm is accused of racism.
But nothing is that simple, because all kinds of things like income and education level and so on correlate with race, so you have to take all of those things into account before you can tell what's going on. And taking into account all of the available data is how machine learning works.
Which isn't to say that you couldn't make an algorithm racist. Tell it to optimize for applicants with a particular skin color and it does. But then your problem isn't with the algorithm, it's with the jackasses who asked for that.
What to optimize for is a much more general and difficult question. (Hint: Not paperclips.)
Likewise, if the system is trained to duplicate human decision-making (like who gets loans), interesting things can happen: if the decision-makers unconsciously favored whites over blacks, the algorithm could wind up weighing skin color or stereotypically Black or Latino names negatively, meaning that the final model is explicitly racist, just because there is a correlation in the training data. That doesn't mean we shouldn't use deep learning, it means that it's not responsible to just fit the training data and ship without testing for such problems.
This isn't racism at all. It's just bad PR because humans take the implication that calling black people monkeys is calling them stupid, since that's the implication you would draw if a person did that.
An algorithm doing that is just recognizing that humans and gorillas are both primates:
http://www.aquilaarts.com/bushmonkey.html
And then it's a bug, in the same way that recognizing a black balloon as a balloon but a white balloon as a light bulb is a bug. It has nothing to do with race at all. The algorithm isn't racist against white balloons. The solution is a general increase in the amount of training data, which is what you want in all cases regardless.
> if the decision-makers unconsciously favored whites over blacks, the algorithm could wind up weighing skin color or stereotypically Black or Latino names negatively, meaning that the final model is explicitly racist, just because there is a correlation in the training data.
Except that this is exactly the thing that a paperclip optimizer will smash to bits because it interferes with the goal of making more paperclips.
Blacks don’t reach the intelligence and blah to be human. I think that’s what racists drive at when they call someone a monkey, and that’s why it’s so offensive.
It would also make your theoretical AI racist, as it identified blacks as not human.
Honestly, at the end of the day that is what is so difficult about much of this. It’s mostly subjective
https://www.theverge.com/2018/1/12/16882408/google-racist-go...
That isn't how racism works. It's like saying that an AI that misclassifies a bat as a bird is racist. It's not racism, it's just error.
And it's not a race-specific error, it's a general error for which someone cherry picked the instances that imply a racially motivated intent that doesn't actually exist.
Calling it racism is pointless and misleading because there is no race-specific cause or solution to the problem. The solution is completely identical to the one for the same error in the general case, i.e. get more training data.
I don't get to how you go from this statement, to then again explaining exactly how racism is embedded in algorithms. By using the biased data we have in the real world...
To fix that you have to cause more black high school students to go to college and study computer science and then wait two generations until their proportionality in the installed base of qualified computer scientists reaches parity. There is no magic wand that makes it happen overnight.
But concentrating on the places where it can't be solved instead of the places where it can will make it take even longer.
There's existing a term for people with this view:
An apt comparison.
If Google really has no idea what the impact of a change will be then it is fairly irresponsible to make that change given the real world harm it can cause. But I suspect in general it does have at least a reasonable idea what the effect of changes will be - that is why it is making them.
So the more reasonable version of this is that they need to submit human interpretable descriptions of the effect of changes based on reasonable evidence and validation of their models.
Google and Facebook partially relies on the obscurity to keep the fighting the spam battle. IMO we don't have the technology yet to have fully open ranking algorithms that are not quickly broken.
To think of it - similar to crypto around WW2.
Google's best asset for ranking is their user data. Even if you had the exact algorithm, you couldn't game it without massive amounts of user traffic. (At least not for popular searches.)
You could get rid of all their user data and it would still be a great search engine.
The reason I'm asking is that as these things grow in complexity, it's quite possible that even if you join the team that works on these systems it will probably take you a pretty long time to understand how they really work. Their actual behaviour is likely to still be mysterious a lot of the time because they're driven by data.
Is a high-level description in english OK? Do we need to see pseudocode? The source code code? Do they have to open source it? What parts, if it's tied to internal frameworks? If there is ML, do they have to disclose all their sauce there? The trained network / weights? The training data, if the alg alone is useless without a data set?
That documentation will need to be shared, and the implementation of the rule change will need to be delayed until the disclosure window has passed.
But yeah, the product manager view / documentation of intent sounds generally reasonable.
I do wonder how useful that would be to the news orgs in practice.
But on the other hand, a bunch of journalists will have a ton of never-before-seen information about how the world's most powerful companies affect every other company on the planet. That alone is going to be worth some major exclusives.
Also, by the mere nature of being forced to share it, Google and Facebook will have to clean up their acts, they'll have to assume any change they make that could open them up to legal scrutiny will be found.
The search algorithm tells you the order of search results for a particular set of terms. Except that as input you need to feed it a graph of the entire indexed internet, which is re-indexed periodically as the content on the index changes. How does knowing that benefit new companies? What, exactly would your hypothetical full-time guy/team, equipped with that index at huge cost, tell their company that would justify the time and expense? That they should write interesting content that lots of people consume?
Second, the general approach has been published and is well documented [1], as are its susceptibilities to attack [2]. So there's your algorithm, what does it tell you?
Third, general SEO isn't the problem, it's coordinated attacks that can poison all search results / ads markets if enough detail is known. Google invests [3] heavily to address these areas [4].
Finally, you underestimate how much of a firehose you'd have to drink from. It describes all of the internet.
[1] http://infolab.stanford.edu/~backrub/google.html
[2] https://en.wikipedia.org/wiki/PageRank#Manipulating_PageRank
[3] https://www.quora.com/What-does-the-Counter-Abuse-Technology...
[4] https://www.blog.google/around-the-globe/google-europe/meet-...
> Furthermore, advertising income often provides an incentive to provide poor quality search results. For example, we noticed a major search engine would not return a large airline's homepage when the airline's name was given as a query. It so happened that the airline had placed an expensive ad, linked to the query that was its name. A better search engine would not have required this ad, and possibly resulted in the loss of the revenue from the airline to the search engine. In general, it could be argued from the consumer point of view that the better the search engine is, the fewer advertisements will be needed for the consumer to find what they want. This of course erodes the advertising supported business model of the existing search engines. However, there will always be money from advertisers who want a customer to switch products, or have something that is genuinely new. But we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm.
Larry and Sergey themselves both believed that ad-funded search was problematic, and that a transparent search engine in the academic realm was "crucial".
Unfortunately, Larry and Sergey's price was clearly billions of dollars.
Publishing "the algorithm" doesn't have value for most of the internet, it is impractical to use even if it was, and bad actors would use it to destroy search quality to the detriment of e-commerce everywhere.
1. There is quite a difference between compulsory auditing (what the post you reply to refers to) and the government directly controlling industry.
2. In other industries this is quite commonplace and hasn't led to government takeover of industries (banking comes to mind. In their regulatory implementation on the Basel III accords developed in response to the 2008 financial crisis, both the UK and EU mandate government audits to ensure compliance with stress-testing and and leverage requirements; the US is also a signatory to these accords, but I am less familiar with their implementation into US law).
I'm not personally a huge fan of this approach, but I don't find the argument that government oversight is a slippery slope to totalitarianism that persuasive. In my opinion, a much a stronger critique of mandatory government audits is that they are often not that effective at preventing the negative outcomes they set out to prevent but still massively increase the legal complexity of operating in (or entering) a given industry without falling afoul of the law.
What is this company, out of curiosity? My guess is ABC, but I don't know.
I'm not sure the exact online share.
Kevin Rudd and Malcolm Turnbull are examples of what happens when you try and dictate terms with NewCorp and they turn on you with negative press.
https://www.theguardian.com/media/2020/nov/18/kevin-rudd-and...
On the other hand, the Liberal party is very hostile to the public service in general and the ABC in particular.
ABC's standards are that it's okay to lie as long as you retract it a month later in a tiny 10pt foot note.
On the other hand, I've personally reported a similar article inaccuracy to a News Corp writer and he replied in 10 minutes, issuing a retraction.
Similarly, I reported an article inaccuracy in a Fairfax website and they retracted in less than 2 days. No reply but as long as it's corrected I don't mind.
SBS is even worse, they actually have zero accountability for online operations.
What are these supposed inaccuracies??
This is not an organisation that cares about journalistic integrity. In fact they actively eschew ethics while their private sector counterparts reply in 1/50th the time or less.
Also, you're using an ABC-produced show as evidence that the ABC isn't ethically compromised? "We investigated ourselves and found we we did nothing wrong"?
Before you accuse me of being a shill, remember I've had retractions printed in News Corp outlets, too.