[0]: https://addons.mozilla.org/en-US/firefox/addon/amp2html/
https://chrome.google.com/webstore/detail/redirect-amp-to-ht...
Note to others: this is about the upcoming rewrite, available as Firefox Preview and Firefox Beta.
I got:
Don't f* with paste, Disable WebRTC, Dark Mode, a couple of image savers, NoScript, AdBlock Plus, Privacy Badger.
I recommend https://filterlists.com/ if you use any blockers, it's an easy way to subscribe to lists.
There's a big rewrite being done and the current stable Firefox for Android which supports basically all addons that the desktop version does, will be deprecated soon-ish. Preview has a broader support.
Beta and Nightly only support uBlock Origin, literally.
Preview supports six addons in all, but Preview isn't a promise of whats to come as they consider it a pilot.
They're not faking the URL; a signed exchange contains data that can only have come from the original site. It's a secure way of handling caching/CDNs/etc, and it'll be a net improvement for security that allows sites to put less trust in third-party servers and scripts.
Its a similar vein to how many of the objections to AMP were being white-washed with "its open source, if you have a problem with it why aren't you submitting a pull request??"
This is literally already happening with the info sidebar, the reason signed exchanges matter is so they can throw up the smokescreen about how its cryptographically verified to come from the original page, so why are people upset.
Any site that objects and refuses to implement this stuff will just disappear from the first page which is reserved only for Accelerated By Google sites. (Which is already happening for AMP links on mobile searches).
I've clearly missed something. Can you help me?
Perhaps I need to read them more closely.
The calculation appears to be that given the chance, some website controllers will choose to trade confidentiality of public pages for better load times. In business terms, this seems a pretty straightfoward win in many cases, so I can see why some would sign up.
A publisher is free to switch to AMP, but choice needs to be given to the user to agree or leave the site the same way it happens with cookies. I wouldn’t opt in and now I cannot block Google tracking at the DNS level thanks to this.
We are, unfortunately, a long way from this being normal. Even as third parties doing things like running CDNs or doing TLS termination for other reasons has been pretty thoroughly normalized. Though offering it as a Firefox extension could be an interesting exercise.
I think publishers view AMP as a question of their sovereignty and choice. Since it's their website that's being potentially served by Google, it's their choice to make. There's absolutely a lot of room to dispute if this is the morally correct stance, but I also think it's not wildly out of line with other questions publishers weigh in choosing what they serve and how.
How do signed exchanges break blocking Google tracking at the DNS level? You already need to have google.com unblocked in order to get a results page that serves an exchange from Google.
> cryptographically verified to come from the original page
Elaborate.
There will be no obvious distinction between a search result and google's own website, even though the "portals" will be showing cryptographically signed content (which will almost certainly be in a super restricted AMP-esque format that denies sites much control besides what the text says)
If google swaps one result for another, 95% of users wont even notice.
Right so exactly the same way the search results page works today?
> The page results will be a tightly restricted iframe-esque window inside of google results
This is completely independent of AMP. Portals work for non-AMP content too.
> which will almost certainly be in a super restricted AMP-esque format that denies sites much control besides what the text says
So you're suggesting that Google will create a new, even more restricted than AMP protocol to be viewed on the search page?
Will Google allow for example, Taiwan content in mainland China?
Is it like VPN/Proxy or has more knobs?
Google doesn't run in mainland china, so this is a bit of a strange question. But let's assume that Google did. How would AMP or signed exchanges change the accessibility of Taiwan related content on Google in mainland china? Are you saying that the current Google results page would show Taiwan related content, but signed exchanges wouldn't, or what?
> Is it like VPN/Proxy or has more knobs?
Neither. It's literally a way to say "a site had this html, css, and js, and cryptographically signed the blob so that we can re-host it and you can be certain that the person re-hosting hasn't modified it in any way."
My question is will signed exchanges serve content but where google search results are censored.
That said, Google already has control of every AMP page because the spec REQUIRES you to load a piece of Google controlled/hosted JS onto your page. That JS can change at any time without "signed exchanges" being aware.
Ok why is this an issue? Note that HN is a bad place for this kind of socratic method discussion, we'll both quickly run out of the ability to post replies. Assume I'm someone who doesn't share whatever values you share about the purity of the url bar or whatever. Why is you being unable to know whether or not the content came from Google's IP or mysite's IP relevant to anyone as long as it's the same content (which signed exchanges ensure)?
How is this different than today, where many sites use js from google, either as a cdn or part of the ads infrastructure? I guess you can block some of those, but blocking the jquery provided by google's CDN isn't going to work too well.
(And further, what kind of nefarious thing do you fear Google will do? How likely is it that they will do so, in your opinion?)
Well, the headline is one good example. That google controlled JS is EXACTLY how they removed access to the original URL...on somebody else's page that isn't theirs. "Signed exchanges" doesn't fix that either. It's also how they hijack the back button and swipe events for carousel navigated pages.
No, the Google AMP cache adds the header bar. That isn't added by the Google controlled AMP js. Let me repeat this: The AMP js didn't change. Google's AMP cache implementation changed. (if you disagree with this, please post the diff of the AMP js that removed the url bar, the js is opensource at [0])
> "Signed exchanges" doesn't fix that either.
Yes it does, in two ways:
1. It would prevent Google from mucking with the embedded page at all, like they do now.
2. It would remove the need for me to have the url redirect, since the url bar would point to the original site.
It's different because it's a requirement. nytimes.com is moving to phase out all third-party advertising data, so presumably they could design their page such that it only accesses their resources.
With a signed exchange, that would allow them to nicely compartmentalize and contain privacy to their site, if they aren't required to load and run some Google supplied JavaScript. The argument that Google already knows that someone visited the page so it's no big deal is not compelling, since there is a big different in knowing someone clicked to visit a page, and having carte blanche over loading your own code on the page in question.
Can you include the AMP rquiers JS inline such that it implement an AMP spec version, or do you need to load it externally? If you can supply it inline, that's great, and what people would want (as long as it doesn't load additional third party resources). If you can't then you're providing Google with an extra level of control that's not really needed, and that's what people are against.
> (And further, what kind of nefarious thing do you fear Google will do? How likely is it that they will do so, in your opinion?)
If we go forth only considering what we think people will do, and not limiting what they can do, we're destined to be upset with the outcome. If not from Google itself, then in twenty years when someone buys Google, or Google sells off a division that houses information, or there's a breach and it's exposed, or some other company rides on Google's coattails and uses the same precedence to get data but is less trustworthy.
The point is that some people don't want to share this information, and would choose not to do so if there was an easy way to tell when it was being gathered. Fighting against new methods that seek to make it implicit instead of explicit is the only real way to do that.
1: https://www.axios.com/new-york-times-advertising-792b3cd6-4b...
Then I'd direct you to Gregable's comment (who is a person who actually works on AMP) that
> the AMP project is actively working to move the origin (control/host) of the AMP Javascript to the publisher's own domain, as well as allow a version served on an origin owned by the OpenJS Foundation, rather than Google.
So while this isn't supported yet, the people working on it do ant that.
> If we go forth only considering what we think people will do, and not limiting what they can do, we're destined to be upset with the outcome. If not from Google itself, then in twenty years when someone buys Google, or Google sells off a division that houses information, or there's a breach and it's exposed, or some other company rides on Google's coattails and uses the same precedence to get data but is less trustworthy.
I'm unconvinced by such slippery slope arguments, given that the pushback were Google to do something like inject nefarious js would be swift. They've had the ability to do so for, well, 20 years now. They haven't yet.
as we don’t have proof that google did not do nefarious things, we don’t have proof that they haven’t. with such monopoly and power distrust is useful thing.
We do have proof that they don't do the specific nefarious things being discussed here: injecting nefarious js into otherwise useful things. That's easy to determine.
It's about power dynamics. If you get a consolidation of power, that's going to be open to abuse. Maybe not now, maybe in the future, who knows. Democratic systems have checks and balances in the public domain. Google doesn't have this.
Good! For what it's worth, I'm slightly pro AMP based on the idea, I'm just not entirely happy with the current implementation. Fixing it to be less dependent on a Google resource is a good change, IMO.
I use copious Google services, such as Gmail and Drive, and Hangouts (or whatever it's called this week), and Android, but I'm leery of becoming more dependent on Google. It's to everyone's benefit if there's healthy competition between all parties, and to my personal benefit if I don't find that someone's gotten access to my google account and literally everything is open to them (which is why I always use a username/password combination for sites I create accounts for instead of linking my Google account... even if I know my email is @gmail.com so it's of limited use, for now. Baby steps).
> I'm unconvinced by such slippery slope arguments, given that the pushback were Google to do something like inject nefarious js would be swift. They've had the ability to do so for, well, 20 years now. They haven't yet.
First, it doesn't have to be nefarious. The bar for Google deciding they deserve analytics for content they "serve" is much lower than the bar for actually doing something illegal. I prefer not to place options to do what I consider the wrong thing for business gain in front of companies when it can be helped. Hope for the best, plan for the worst, and all that.
Second, that was a single one of the scenarios I listed. The others notably did nt rely on Google doing or not doing the right thing, because the decision is no longer in their hands. If Google is no longer the authority deciding (because they are gone, or have a new parent, or the data was taken), what Google would choose to do is irrelevant. That's why it's important to some people to reduce the information being collected. It's impossible to know what it will eventually be used for in the long term, so the prudent thing is to limit it, and/or compartmentalize it (that is, maybe I'm happy with nytimes.com knowing where else I clicked in their article, but I would prefer Google only know I loaded that first article).
PS: You work for Google. Do you work on this project?
You'll only ever retrieve Google AMP cache results from the Google search page, where they were already able to track if you made such a request, since the link you clicked has trackers in it.
So from that perspective, nothing changes.
> PS: You work for Google. Do you work on this project?
No, I work on mostly internal infrastructure. My interest in AMP is simply that I don't dislike the AMP "experience", it's fine. But more importantly, I legitimately don't get the HN hysteria around AMP. Returning to your concern, literally nothing changes with AMP vs non-AMP.
I don't get it. The most compelling concern I've heard is that it's annoying to have to couple parts of your infra to AMP-standard stuff. And I sort of understand that. But even that isn't different than previous SEO/ranking changes that required changes to the page.
I am not affected, I don’t use Google search. The problem is for individuals who use Google search and now don’t have an option to avoid in deep tracking. The difference between regular pixel trackers and multiple data points associated with every resource a site serves is immense. I work in ad-tech, not particularly in the identification side, but I started multiple projects in that end. From experience, a regular tracker can be fooled, but you cannot fool every resource request. One of the things I did to identify ad fraud bots was actually drive them to a site in which I controlled every resource. The resource request fingerprint for bots was easily distinguishable from real people. Moreover, some humans exhibited navigation patterns that were distinguishable from other humans. I remember I caught a QA person doing a shoddy job of testing the front end once due to it. That is the kind of power that Google is acquiring as more and more sites choose to use AMPs. It is scary to think that a single identity has that power.
I'm confused.
What they're pushing with Amp and the related technologies grants them a near-unavoidable man-in-the-middle position.
(FWIW I am careful to avoid Google properties at a pretty high cost.)
Not anymore than I already do bat an eye at cloudflare.
(It's probably worth noting here specifically that I do work at Google, so my risk profile is probably different than yours, for me personally and speaking solely from a trust perspective, I'd probably prefer it if Google acquired CloudFlare since I would get a net increase in transparency, but I can understand why that isn't a general position, and there are other reasons I don't think Google acquiring cloudflare would be good).
I share these fears to a lesser degree with Microsoft and of course Facebook. Apple seems to do a great job of safeguarding, but they could become sour if they don't remain careful. Stuff like Clearview crosses the line into directly-dangerous. CloudFlare is currently innocent in my eyes, but they've managed to centralize a lot more channels than I'd like to think about.
Uh, I keep seeing Google AMP URLs shared on social media, emails, etc etc. Which is quite annoying :(
If you click the browser share icon, or trigger the browser native share intent, the origin URL will be shared, not the AMP Cache URL. Only if you explicitly copy the URL bar will the AMP Cache URL be shared.
The Signed Exchange spec that AMP has offered sites for a year now allows them to have their own URLs displayed in browsers that support it. In that case, the google.com URL will never be displayed and thus can't be accidentally shared.
All AMP documents on the AMP Cache contain `<link rel=canonical href={origin url}>` and Google recommends that social media prefers the canonical URL. This is useful outside of AMP as there are often multiple URL variants for any article. The sharer and sharee may not ideally get the same version. As an example, a mobile vs. desktop article.
An AMP page can be identified by examining only the first few bytes of the HTML. The `<html>` tag will contain either the `amp` or lighting-bolt emoji attribute, ie: `<html amp>`.
Technically an AMP document must pass AMP Validation to be truly AMP, so there are documents that match the above condition which aren't valid AMP. There are multiple ways to validate. A starting place is https://validator.amp.dev/
Also, "the google.com URL will never be displayed" is a world with an internet I don't want to be a part of.
Also, others might share an amp link from their mobile devices, which I then end up clicking in a desktop slack/mail/messages app, and there we go again with the amp virus even on desktops.
I see quite a few AMP cache results being shared on twitter (for example).
>where they were already able to track if you made such a request..
It's still possible to get raw Google results (with the right extensions/browsers), though I wonder for how long.
From then on, every asset it loaded via Google servers. Google now controls the entire internet. Google does this so it can serve its ads and track all users. It's as if I would only use Google for my internet surfing.
I don't use Google because I strongly believe it is an evil company, but if websites use AMP, then I am forced to hand over my data to Google even though I don't want to. Right now, I can block Google servers entirely. But if the entire web is served via AMP, I can't do it. And that's the whole reason AMP exists. So everything I do (or at least as much as possible) goes through Google servers.
You really do not see our concern? Really?
I’m curious, what do you value about the services you consume? I like transactions where I know what I am giving to the service provider. Do you really want to push away from this reality for the benefit of a few MS load time leaving a search page? That’s essentially what you’re arguing for.
I don't think you may realize how much of your online activity is already tracked by google / facebook / instagram.
Google's javascript is everywhere, including explicit tracking with analytics, and lots of CDN loads for endless lists of things (js libraries, fonts etc).
Their properties also track you, google search, youtube, email. They also make software you might use (chome / android / google maps / google play store).
If you think something about signed exchanges let's google track you, and they can't now... please examine these assumptions.
Folks who come up with these super complex schemes (google will use javascript loaded into AMP to take over and track you) ignore that google ALREADY tracks them.
And folks who say they don't use any google products (no android / google maps/ play services / chrome etc etc) are often either lying or don't understand how many third parties load google analytics into websites, or load recaptcha bot protection etc.
I wrote a reply to joshuamorton were I expressed my concerns.
If google said, we want to track people, and brings android, chrome, dns resolvers, network infrastructure, google cloud compute, AI systems, google analytics which these media sites voluntarily, google play services etc to target and track you - they probably could.
EVERY single person (including you) who claim they don't use google, if you dig down, they often are lying and do. And if you don't, some of the people you email or interact with do, so indirect profiles can be built.
AMP solved a need for a lot of users, which is the janky, slow ad filled websites that media sites in particular had become. So there is an actual end user reason people like AMP - it's a better user experience in many cases. This is where AMP is ruining the web gets hard to support. For most folks they don't perceive they are giving up a lot more in terms of privacy, and they are getting a lot.
Their data governance team wouldn’t allow it. You are basically describing a system they could only introduce with the permission of the government. I don’t care if the government is tracking me honestly, I can’t fight that. I just don’t want Google tracking me for the purpose of influencing my spending habits, emotional state, or perception of the world. That is my main beef with their advertising capabilities.
At the same time, the AMP project is actively working to move the origin (control/host) of the AMP Javascript to the publisher's own domain, as well as allow a version served on an origin owned by the OpenJS Foundation, rather than Google.
On publisher origin, the plan-of-record does not involve any validation of the contents of the AMP javascript files. When an AMP Cache (eg: Google) crawls one of these AMP documents, the same is true - the contents of the javascript files will not be relevant to the decision of whether or not the document is considered valid AMP. The files will likely not even be crawled by the Cache.
However, when the AMP Cache serves one of these files, it will rewrite them to the latest* version for serving to users. This is necessary since the javascript runs in a somewhat privileged context in search results.
https://github.com/ampproject/amphtml/issues/25873
* There is also a mechanism for publishers to opt-in documents to a "Long-Term Stable" release, rather than the latest evergreen version: https://amp.dev/documentation/guides-and-tutorials/learn/spe...
Lastly, and there is still some discussion around this, it is likely that Signed Exchanges may be able to load the publisher's own version of the javascript in the future, even in search results. This is because the execution context of the javascript is different for Signed Exchanges.
Third parties like Google, right?
Signed Exchanges mean the publisher signs the content using their private key. A third party can provide delivery like a CDN, but they cannot modify the content, or the signature would no longer match. The useragent (browser) enforces this. This gives the secure control of the content back to the publisher, unlike the trust model of CDNs or the AMP Cache.
This assumes the user agent is actually an agent of the user, and not the AMP provider, which is demonstrably [1] not the case.
[1] https://github.com/w3ctag/design-reviews/issues/467#issuecom...
If a signed exchange includes a URL, that URL must be signed for the browser to respect the field.
[1] Which is unverifiable, we just have to take your word for it.
To make sure I understand, does this mean that in principle a third party other than Google can deliver the AMP pages? Is google working to facilitate that AMP hosting is open to everyone and calibrating their searches point to any and all alternative AMP hosters?
Also how would you reconcile your comment with that of madeofpalk who appears to be treating that possibility as a hypothetical idea that hasn't happened, and which would be unpraticable due to needing to trust third parties?
All AMP pages exist at non-Google URLs. They are just cached by the link aggregator (typically a search engine), so the link aggregator can prerender them without deanonymizing the user to the publisher until the user clicks the link.
AMP served by CNN: https://amp.cnn.com/cnn/2020/05/27/world/france-shooting-sai...
The same page served out of Bing's AMP cache: https://www.bing.com/amp/s/amp.cnn.com/cnn/2020/05/27/world/...
> Do you have a ballpark estimate of what percentage of total amps are delivered by non Google domains?
All of them (100%) are delivered by non-Google domains to Google, Bing, and other caches.
> Also how would you reconcile your comment with that of madeofpalk who appears to be treating that possibility as a hypothetical idea that hasn't happened, and which would be unpraticable due to needing to trust third parties?
madeofpalk's comment makes perfect sense if you understood what I wrote above. Why should CNN or Bing be told that you have searched for a particular news article on Google before you have clicked it? The page has to be served from the link aggregator the user is browsing to maintain the user's privacy when prerendering results.
I don't intend to ask whether AMPs (hard to resist calling them 'AMP pages') exist somewhere on non Google servers. Obviously third party content that Google is presenting exists somewhere off Google. And obviously it has to be formatted in a way that's compatible with AMP, and it makes sense that that is going to be done off Google domains. I at least knew the gist of that already, and I regard the detour into that explanation to have been a non sequitur. The point is that Google presents AMPs and it serves it's cached version of them from Google servers, on a Google domain. The beginning, middle, and end of the experience of searching for finding and consuming that news never has to involve leaving a Google domain. It's not open in the sense of involving interaction between servers that aren't controlled by Google, until you make that extra click to go from a cached Google version of an AMP to the version that sits on the domain controlled by a third party, at which point going to the third party has been rendered optional and largely unnecessary from the point of view of the user.
This next part is super important: the fact that I'm asking about openness and interoperability, or the lack thereof, in this sense doesn't mean that I'm failing understand the technical advantages with caching and optimization. I regard those as derails that don't wrestle with the issue of openness that's being raised. The point is that the connection between consumers of content who start on Google, and the third party content provider, increasingly depends on Google in a way that shifts nearly the entire experience of consuming content onto Google's infrastructure.
>All of them (100%) are delivered by non-Google domains to Google, Bing, and other caches.
This is the starkest example of a question not being answered but replaced with a different question. I asked 'what percentage of total amps are delivered by non Google domains' and you replied by answering a different question, what percent of non-Google amps were delivered TO Google and other caches, noting that it was 100%. Which of course it is, but that's because that's a tautology.
By contrast, it is helpful to note that there are caches other than Google, like Bing and 'others', which, in contrast to much of the rest of your comment, I feel actually is a pertinent and fair response to the question I'm actually asking. But those aren't content providers, so unless Bing or Google are content creators that were delivering content to themselves, it's tautologically true that 100% of that is going to be delivered to them by third parties, which has absolutely nothing to do with openness. If I'm using magic words correctly, I guess what I want to ask is what percentage of AMP traffic to cached pages is served to users by Bing and others that aren't Google.
> If I'm using magic words correctly, I guess what I want to ask is what percentage of AMP traffic to cached pages is served to users by Bing and others that aren't Google.
If I search on Bing, the results will be prerendered from Bing's AMP cache. Reread the GP comment, and see if you can understand why that is so.
> Can a third-party other than Google deliver an AMP page?
Yes. Examples: Bing runs their own AMP cache and also delivers AMP pages. LinkedIn and Twitter also link to AMP pages, but they don't currently run a cache. IIRC, Twitter links to the Google AMP cache and LinkedIn links directly to the AMP variant on the publisher origin. They could run an AMP Cache. Cloudflare ran one for some time, but shut theirs down recently.
The AMP Project maintains a list of known AMP Caches here: https://github.com/ampproject/amphtml/blob/master/build-syst...
And provides some guidelines for running one here: https://github.com/ampproject/amphtml/blob/master/spec/amp-c...
It's non-trivial, but absolutely supported.
> Can a third-party other than Google deliver a Signed-Exchange?
Yes. Cloudflare generates them for their customers who opt-in via their "AMP Real URL" product. "Generates" in this context implies delivering them. To date, I'm unaware of any large scale implementation that is delivering Signed Exchanges for third-party origins other than the Google Cache though this may change. The tech stack absolutely supports this.
Fifteen years ago, if you asked me how google search would look, I would have responded “mostly the same, maybe they’ll have cool features like asking me which meaning of ‘converse’ i wanted: the shoe brand or the logical relation.”
Instead, they’ve only subtracted functionality from the query engine (no more domain blocking), discouraged you from clicking through to sites by automatically scraping and rehosting them as “semantic” results, and now they’re trying to actively acquire 100% of the outbound traffic. Fuck google.
It blows my mind that anyone thinks that’s an acceptable idea.
What’s shown in the URL bar has long been divorced from the HTTP request(s) that are made.
> It blows my mind that anyone thinks that’s an acceptable idea
It blows your mind that different people have difference opinions and values than you?
Also, a provider uses a CDN at their discretion. Giving them the ability to invalidate or update cached records at times of their choosing. Or remove the CDN entirely if they choose to.
This is Google using their weight to be anti-competitive and fall further down the anti-trust rabbit hole.
I know it seems impossible but nobody has to use Google.
Are there any copyright implications of amp? They're essentially republishing your property.
Does the original publisher get their ad revenue?
Given context and full knowledge of the issues involved? Yes, absolutely. For this particular issue anyway.
> Or maybe it’ll come from a CDN rather than origin.
That’s not a problem, since the URL and where it points to is still decided by the content owners.
>Or is it that you don't want snippets to appear in the search results and just want a list of links without any evidence for why they might be good matches for your query?
Is that how you feel about regular search results?
> Is that how you feel about regular search results?
Regular search results have snippets.
>Regular search results have snippets.
Right, and I was asking about the search results, not the snippets that accompany them. That is to say, the part with the blue title, green link, and the few lines in black displayed from the page that are displayed ten at a time, not the snippets that accompany them at the top of the page. Unless you were just using 'snippets' as a general term to mean the same thing that I mean by search results, in which case you were just repeating the content of my own question back to me.
They still know when their content is being consumed. They just don't know when their content is being searched for until the user clicks their link, exactly like a snippet. Does that make sense now? My point was that search engines already show cached portions of the page. Read the parent comment of my first "snippet" comment to understand why I was making that point.
If I click one of the links, then yes, I absolutely expect the publisher to know I did.
Google scraped and stole all the traffic for covid19 since 3 months just like it did for other topics
Exactly. Search your hearts and use the net with your values as though they’ve become a force to be feared by the swill that is invading our liberties (google).