AMP pages displaying your own domain
webmasters.googleblog.com
webmasters.googleblog.com
* does not show all comments, often ones I am actually looking for
* does not let me collapse comment sections
* uses the default white background theme which burns my retinas if I am looking at my phone in a dark environment
* shows overlay ads for the Reddit app that cover about 40% of the screen for no goddamn reason
* requires 2-3 separate actions to get to the original page
Yet I cannot find a browser extension or setting to tell AMP to fuck off. Honestly AMP might be what finally gets me to switch search engines after many years of using Google.
Surprising how little I've noticed the change, after using Google Search for over 15 years. I try queries on google.com maybe once or twice a week if I don't find what I'm looking for on DDG. If it's anything media or product related, I feel like I'm on an old, crowded MySpace page. DDG feels more like the old Google.
Also, links in the Twitter app default to AMP as well.
In contrast, say, Urban Dictionary is undistinguishable from the real thing.
Searx can be extended so it would be possible to create a plugin which rewrites AMP links into non-AMP equivalents, where available. It can already do things like Open Access DOI rewrite (Avoid paywalls by redirecting to open-access versions of publications when available) so the ground work has been done. I'm currently working on improving (and fixing, where necessary) the image search engines and will probably start on such a plugin if nobody else beats me to it.
AMP is straight up broken technology. Imagine if you subscribed to a print version of the NYT but instead of getting the Sunday edition you got a ransom note looking summary of some of the articles from Clipper Magazine. Would you be OK with that?
Also, there is no mechanism to limit who is allowed to serve your content for you.
I see no technical reason why the content has to be prefetched from Google instead of your own server.
It's also confusing for users and administrators. Want to block access to a website in your network? Guess what: Your block will not be effective because Google will proxy the data unbeknownst to the firewall.
Btw your account seems rather active for an account with the description "Inactive. Deletion Requested." :-)
JS based analytics (google or otherwise) is generally a better option for detecting actual usage. Yeah, you lose maybe 2% of actual users. You also lose 99% of the various bots. You still have to filter google's and bing's bots that execute JS though.
I'm OK with publishers knowing less about the people seeing their content.
I think the real reason is that Google wants to build a walled garden, but doesn't want the walls to be noticeable. Even with AMP, they display a header that looks like a browser's address bar [1]
Also, on that page Google admits that it uses AMP Viewer to collect information about users:
> Data collection by Google is governed by Google’s privacy policy.
Which is probably their real motication for creating AMP.
[1] https://developers.google.com/search/docs/guides/about-amp
That's what AMP already did. This spec is better because it ensures publishers retain control over their own content, and doesn't confuse users by showing "www.google.com" in the URL bar for content that didn't originate from Google.
What confuses users is Google displaying a fake address bar [1] or browser displaying the wrong URL.
[1] https://developers.google.com/search/docs/guides/images/amp0...
Your browser controls the contents of the URL bar, not Google or the publisher.
Page and DNS prefetching exists, HTML exists, why not just link to the page on the original domain?
Exactly, this is the real reason why this abomination came into existence - all of this is masked as work for greater good all for those poor kids with limited network speed. As end effect everyone will suffer - user will never leave google ecosystem, he will remain on search page without even knowing about it, creator will lose control over his own content
Why should the user‘s privacy be protected toward the content provider instead of the search provider? The search provider already knows more about me.
Now there's a solution that preserves the preloading and validation benefits of AMP caches but maintains the original URLs, in a way that's cryptographically sound, in the process of being standardized, and controlled by the publisher. This gets launched much faster than one would have expected. And suddenly everyone pretends that the AMP cache URLs were never a problem and this is some kind of a power-grab.
It feels a little dodgy to me this standard and a bit embrace extend but I'll see how it plays out and reserve judgement until we see this happening in the wild and how well it works. Personally I'd like to be informed in the browser chrome that it was being served via this mechanism rather than me visiting the original site.
Can you maybe see that people feel the browser is now lying to them about where the content is coming from?
And you're also back to the situation where you can't preload the content in a controlled manner or privacy-preserving manner, nor have the page-speed guarantees since the version being served to the user is not the version that Google crawled.
It's kind of the opposite. The cache is where the actual benefits come from. That's not the part you want to get rid of. The AMP spec was just a vehicle for making the caching possible in a secure manner.
This model would theoretically allow the validation, caching and prefetching to be done for all (signed, so opt-in by the publisher) HTML pages. Which is another one of the historical top complaints about AMP: why can't light, fast-loading, mobile-friendly HTML get the same treatment in search results.
> Can you maybe see that people feel the browser is now lying to them about where the content is coming from?
I can see that they are feeling like that, I just don't understand how they arrived there.
How is this different from a e.g. company X's website being behind Cloudflare? The browser didn't contact the actual server that company X hosted the content on. Instead the browser contacted a server run by Cloudflare that could prove cryptographically (via TLS) that it was authorized to serve content on behalf of the actual site.
The browser security model stops them from doing this, but presumably in this new world they could allow this to work and not host the content in the carousel themselves.
I think the argument about content suddenly becoming "slow" and no longer AMP validated if it's not served from the AMP cache is a poor one.
Finally I'm willing to postpone judgement but I did just explain why people feel that Google is embracing and extending the web if you can't understand why people are worried about this that's not something I can help you with ;-)
Cloudflare does not have the same scope, power, monopoly or scale that Google have - I can change CDN provider if they start doing weird stuff, no problem, but I can never really get away from Google.
A few people have pointed out the privacy-preserving aspect of AMP. I'm not sure I get how that's the case. Is this referring to the fact that the page is not being pre-loaded from the content owner's own webserver? The main privacy violators on the internet are Google and Facebook. How is loading something from Google cache protecting my privacy?
Worse still, if someone posts an amp link on Twitter or a chat client Google now gets to know when I access a specific website even though they are an unrelated third party[1].
Edit: [1] In practice this was probably already the case since Google Analytics is so popular. But still.
If you make a search query, but have not clicked on any results, you have a privacy expectation that the web servers of the search results you have not clicked on will not know you performed this query, your ip address, cookie, etc. For example, if you search for [headache] and then close the window, mayoclinic.com knowing that you made this query would probably be a surprising result.
With naive preloading, you would preload a search result from that origin. Your browser would make an HTTP request to the site and that site (sending an ip address, the URL you are preloading, and any cookies you may have set on that origin). So, this approach would violate your expectation of privacy.
Instead, if the page is delivered from Google's own cache, the HTTP request goes to Google instead of the publisher. Google already knows that you have made this query, and are going to preload it (the search results page instructed your browser to do so in the first place). The request will not have any cookies in it except for Google's origin cookies, which Google already knows as well. Therefore this type of preload does not reveal anything new about you to any party, even Google.
AMP has been doing this for a long time in order to preload results before you click them. However, until Signed Exchanges the only way to do this was that on click the page would need to be from a Google owned cache URL (google.com/amp/...). With Signed Exchanges, that can be fixed. The network events are essentially the same.
Note that once the page has been clicked on, the expectation of privacy from the publisher is no longer there. The page itself can then load resources directly from the publishers origin, etc.
To your last point, if someone posts a link on twitter to an AMP page on a publisher domain, and then you click it, your browser will make a network request to the publisher's origin. Google will not be involved in this transaction in any way. If someone explicitly posts a link to an Google AMP Cache Signed Exchange, then yes this will trigger a request to Google but this will be far less likely going forward as these URLs will never be shown in a browser. For example, try loading https://amppackageexample-com.cdn.ampproject.org/wp/s/amppac... using Chrome 73 or later. This is a signed exchange from one domain being delivered from another. You'll never see that URL in the URL bar for more than a moment, so it's unlikely to ever be shared, like I'm doing now.
At its root, I think my objections to AMP boil down to a few things:
On a technical level:
1. It's buggy and weird on iOS.
2. I'm not convinced I care about a few seconds of loading time enough to justify the added complexity of making this kind of prefetching possible. Additionally, this seems like a stop-gap that will be rendered unnecessary by increasingly wide pipes for data.
On a philosophical level:
3. It gives Google way too much power over content.
4. I want the option to turn it off completely because of points [1] and [3], and because I fundamentally want to feel in control of my internet experience.
Edit: The point about SXG making AMP URLs less likely to get copy/pasted to other mediums is a key benefit I hadn't considered and will likely make avoiding AMP outside of Google search easier.
I totally want that time saved if possible.
It looks like you were trying to make some deeper philosophical point, but you'll have to be clearer because your statement makes no sense.
Google has to know what you're searching for to compute and show the results. So there are few additional privacy implications from the preload.
And your last case is exactly what will no longer happen. People will now copy-paste the original URL rather than the cache URL. Click on the link, and you're taken to the original site.
What will stop Google from down-grading 2nd class URLs (ie, not hosted with google) to page 2 results?
It's effectively the same thing as having no AMP at all, yet they cleverly got everyone on board with this tactic.
Edit: I just skimmed through this... this looks _WORSE_ than having Google show their domain. This is some of the sneakiest most deceitful garbage I could have ever imagined.
Just no way. Need convincing? Look at the animated gif half way down:
https://3.bp.blogspot.com/-Xqfy7IhiTzc/XLY7goySWzI/AAAAAAAAD...
Because yes that's true, although cryptography it's maybe half true.
Now, they want to remove the remaining user interface element that says they’re spying on me!
Also, this makes it even harder to ad block their junk at the network layer (is foo.com down, or is this more amp bs?)
> The Google AMP Viewer is a hybrid environment where you can collect data about the user. Data collection by Google is governed by Google’s privacy policy.
With replaced URL it will be more difficult to spot.
[1] https://developers.google.com/search/docs/guides/about-amp
IMO, while the URL problem was a big issue, the bigger issue is that AMP's restrictions and limitations gives your users a neutered user experience in the final end. As others have pointed out, if it wasn't for Google's implicit requirement to implement AMP (e.g. to get into their carousel and other locations), AMP would have been DOA.
[1] https://amp.dev/documentation/guides-and-tutorials/optimize-...
Converting web pages into AMP isn't something you can automate, but supporting signed exchanges is. You need certificate authorities to support the flag and web servers support the protocol, but if this catches on then the only thing you'll need from the site owner is the decision on whether to allow it.
(Disclosure: I work for Google)
Sometimes, Google needs a gentle nudge from users saying "we don't like this" and hope they reconsider (I doubt it).
Let's Encrypt's response:
I think it’s likely too early in
this draft’s development for Let’s
Encrypt to prioritize implementation.
It looks like it has a ways to go
within the IETF before it would be
an internet standard.
https://community.letsencrypt.org/t/cansignhttpexchanges-ext...[1] https://developers.google.com/search/docs/guides/images/amp0...
* I like that when AMP is used for ads then the ads are fully declarative. Advertisers getting to run custom javascript, even in a cross-domain iframe, isn't great.
* I like that AMP allows sites (currently primarily search engines) to trigger preloading in a way that doesn't leak information to the site that is being preloaded.
* I like the way things like "sorry AMP only allows us to use 50k of CSS" can give developers leverage to push back against bad site designs.
* I like that it centralizes some measurements: instead of every ad provider using their own custom polling system to determine if the ad is on screen they can all subscribe to events triggered by a single well written system. This doesn't affect the amount of tracking (there's lots either way) but it makes it hurt the user experience less.
On the other hand, I don't like that:
* AMP uses a ton of JS, and if all you want is a simple website it's going to slow things down in the non-preloaded case. For example, taking a random post on my site (https://www.jefftk.com/p/trycontra-implementation and https://www.jefftk.com/p/trycontra-implementation.amp) I see a median speed index of 1.611s on non-AMP but 2.051s on AMP: https://www.webpagetest.org/result/190417_XB_22673cb98ce390a... https://www.webpagetest.org/result/190417_PS_1a60378762d87fb...
* A lot of people that don't want to implement AMP are doing it because then they get more search traffic. I understand how there isn't currently a non-AMP way of doing preloading in a way that doesn't leak information to the site (see above) but I think Web Packaging should be extended to support this in the general case and allow publishers to use AMP only if they want to.
* The interaction between AMP and content blockers isn't great. If you have a content blocker set to allow some JS but not all (for example, no third party JS) then it's not going to run the AMP JS or the contents of the <noscript> block, and AMP pages will render with 8s of white screen before the CSS times out. This is a pain, but I'm not sure what the right way to fix it would be. (I wish content blockers were smart enough to figure out which <noscript> tags to run, but that's probably asking too much.)
If you wanted to expand on how AMP seems like an attempt at a walled garden I would be interested in reading it; I haven't previously read any explanations that made sense.
If they don't/won't, no matter what your justification for why is (you believe it will provide speed, security, whatever), that's one of the walls.
Sure, you can not use it, but does that limit your ability to be found on the internet? If yes, then there's that wall again.
They're in the extend stage of Microsoft's favourite strategy.
Google clearly doesn't treat AMP and non-AMP pages the same way: only AMP pages are eligible for the carousel in Google search, and there's a little icon.
Once there's a way for non-AMP pages be safely preloaded I would be very surprised if Google search didn't start doing that, though. (Speaking only for myself, not the company.)
Disclaimer: I work at Google, nothing related to AMP or search.
I don't have links to hand but everything I've seen shows real dropoffs in users as you increase the time. Once you're looking at low numbers of seconds you're looking at significant numbers of users simply abandoning the site. Half a second extra is not insignificant, and the user experience changes a lot between things that feel instant and things that have a noticeable wait.
[0] For instance, Mozilla considers the current specification to be harmful[1].
(Unless you're using CloudFlare)
[1] https://github.com/mozilla/standards-positions/issues/29#iss...
[0] https://github.com/mozilla/standards-positions/issues/29#iss...
Not the public reason, but absolutely the private reason.
If Google, Apple, Amazon, Microsoft, or whatever publicly traded company makes a move its for money and power and preferably power, since that yields even more money.
AMP on web and email is the perfection of embrace, extend and extinguish
https://blogs.bing.com/Webmaster-Blog/September-2018/Introdu...
I believe if you refresh the page it triggers a request to the original site, which will probably then choose to give you the non-AMP version of the site.
Also, Mozilla members rally around a ton of stuff here on HN. That's why you see so many posts about Rust despite the fact that it's not really that popular. That's also why the top comments on stories about MS Edge switching to Chrome where lamenting the fact that they didn't choose Firefox, despite the fact that hardly anybody uses Firefox.
https://www.zdnet.com/article/former-mozilla-exec-google-has...
I have push notification disabled but it wouldn't be surprise for me if they asking to subscribe for push notifications on the first page view.
Current era of content websites is a disaster except few cases like medium and maybe reddit with a discount.
AMP is an only solution for general users who just want to google a cooking recipe or latest news in their town.
I don't care how AMP works. It's a power grab. Done.
And now we're trying to shoehorn it back in?
It used to be that a local caching squid proxy was a great way to make load times of various "front pages of the Internet" bearable on a shared low bandwidth uplink (local/national news sites etc typically being served from the cache/lan).
New ssl/tls kinda-sorta breaks that (there's no middle ground - either install intercepting cert that catches everything, or abandon caching on everything. Either cache CNN. com and medical records, email(webmail) and Facebook messages - or neither).
AMP might be a bridge too far - but some kind of (semi) public "signed, not encrypted" would still be a good fit for hypertext applications/documents - because of the caching benefits.
[1] As excellently outlined and contrasted by Fielding in his thesis: https://www.ics.uci.edu/~fielding/pubs/dissertation/top.htm
Also, Google controlling AMP means that Google decides what analytic systems and ad networks are allowed on the AMP page. With Google having its own ads and analytics business, doesn't this tempt them to make life little easier for its own products and little more difficult for competitors'?
If you look back, Google has made multiple attempts at "improving" the url. It starts to become clear what they are trying to do now.
Ideally people would develop fast sites on their own, but apparently they need the help of Google.
Users have no control outside of not using google. If google were to provide a setting for the user to never see AMP, I would have less issue with this. But they don't
Instead, they basically force publishers to use this because if they don't the news carousel will not show their article. It just gives Google more control over the web for minimal at best benefits
And AMP is a pain in the ass. It's sold as being "just HTML" but it isn't, really. You can't even use an <img> tag, it has to be <amp-img>. So you have to generate two versions of every page. Achievable for large companies but if you don't have a lot of resources that's a big overhead. As is so often the case, it helps concentrate all web traffic to a smaller and smaller number of sites/publishers and shutting the rest out. That's not good.
If there are some benefits to it why shouldn't those benefits be standardized? Is Google preventing the standardization of AMP?
This is false. as a user I cannot easily opt out of using AMP
Web packaging and Signed exchanges seems benign and beneficial, you can sign a particular page inside a package (let's say a zipped folder of some kind) and now anyone can cache that data and show it, while both the browser and the user knows that it's safe to display it. Since the AMP format is similar, it seems quite beneficial to now have all your AMP content support this feature. And anyone who made some of their pages AMP can use that same process to support other Signed Exchanges (such as p2p networks or CDNs) . This is great since it makes distributed caching much easier.
The bad part is that google search uses this signed exchange format not to show the actual URL but rather put it in an iframe inside chrome (and only chrome). The real question is whether we will be able to use this functionality outside search, if I have my own site and show a large iframe with signed exchange page, will I also be able to change the browser url bar? mmph, probably not.
Try it for yourself. Using Chrome 73 or later (you probably already have this), and a mobile browser (either a phone or mobile emulation), try the query [amp dev success stories].
It will only use signed exchanges in Chrome because currently only Chrome supports signed exchanges. The search engine explicitly looks for the browser to state that it supports signed exchanges in an Accept header, like any other new technology.
Yes, any page can use this. So, for example if you went and fetched a signed exchange from https://amppackageexample.com/ (or any other site that supports one, this is just an example), you could then serve that from your own server, more or less just like any other file (the less is that you need to set the right Content-Type header, but it otherwise works just like serving an image or a zip file).
Then, if a user visited the URL on your site https://yoursite.com/cached-copy-of-amppackageexample.com/ then the browser would display https://amppackageexample.com/ in the URL bar, as though that URL had 301 redirected, but without the extra network fetch.
Google search does exactly this, just loading a cached copy of the Signed Exchange, and any other cache (or even any website) can do the same.
Once I can create webpackages and deliver them to clients a lot of thing I want to do become hugely easier and nicer.
I know it's not exactly easy to follow but the only implementation repo I can think of to follow right now is the Chromium repo.
I also had a look in the blog and the "progressive web apps" might be the right thing to look at. There's probably something subtle that's different but I think I can use these to solve the actual problem I have.
https://developers.google.com/web/updates/2019/03/nic73?hl=h...
edit - damn, I don't think this is right at all. Frustrating as it seems pretty perfect but I have to serve from my own domain for 30s before a user can install it :( I just want a single file way of delivering web content! It seems like all the features are basically there, just with restrictions to focus on different use cases.
Isn't this whole exercise really just adapting public key signatures on top of old school caching?
With a http proxy you ask for an url, the proxy fetches or serves on behalf of the owner. This adds some circumvention around the way tls/ssl breaks that type of caching. But it should still be able to do a head-like request for a current signature - with no need to download the content again if it is unchanged?
Doing this on every page load breaks either user privacy (by making the origin fetch before the user clicks) or the preload performance gain itself (by blocking load while waiting for this round trip).
I hope at least Mozilla doesn't adopt this technology and will show the true URL.
This technology is complicated. Browser vendors have to implement all of this only to please Google.
Guess what they’re going to choose.
Last week I blocked Google from my domains (blog: lucb1e.com/!130), hopefully others will follow suit and degrade the search quality until people get better results (at least for some more obscure content) elsewhere, or perhaps until Google notices we are really not okay with their behaviour.
Blocking is based on user agent, they seem to set that reliably and the IP addresses change. You can do some reverse lookup magic but this was way easier than looking up every single IP that visits my site.
Turns out people were right to be suspicious. This is hot garbage. You can no longer ask a user "What URL does your navbar say you're at?". It is no longer a source of truth. They will actively be lied to.
For a long time already it's not being connecter to a particular physical server. Now it's the next step - to be completely decoupled from the server and just mean content instead.
If Google hosts the website and is masking the resulting url, they're able to have more visibility than Google analytics. They'll likely give this AMP some SEO boost temporarily and that will get web admins to adopt the technology.
It's just like reCaptcha, which is used to track users across the web (requires google.com + gstatic.com urls to load, which drops its own cookies or scans existing ones), blocking recaptcha will break core web functionality... and recaptcha v3 is even worse.
AMP doesn't load in a privacy sensitive way. It's on Google's servers and it takes many seconds to load if you have JavaScript disabled.
Also, the feature only works on Google Chrome and possibly Edge, which gives another point to the article below.
https://www.zdnet.com/article/former-mozilla-exec-google-has...
AMP is a fundamentally bad idea that needs to disappear.
Edit: Mozilla has marked Signed HTTP Exchanges as harmful.
how is this different than using your own domain, but pointing it to a github.io page? Or using medium, but with your own domain (but still being served from medium's servers)?
Is it just google you're adverse to, or the entire idea of someone else hosting your content?
2) I want full control over how I publish my sites with real web standards. AMP is not a web standard, it's a Google format that they are strong-arming people into using.
3) Mozilla considers Signed HTTP Exchanges harmful. This technology is as bad as what Microsoft was doing with IE in the old days.
4) I don't publish on Github pages, but if I did, I would still have a choice over which servers I put the sites on.
5) There shouldn't be a single company (or few companies) that dictates how we publish online.
6) Shame on the people who are splitting the web with this fake-opensource technology. There's even a Google engineer over here referring to the Web like it's a Google product. https://news.ycombinator.com/item?id=19631136
My point is, you cannot just blindly trust anonymous comments to be who they say they are, it’s an easy way to get yourself in trouble.
If they go through a content network like Cloudflare, you can't even tell who's hosting the site by looking at the IP address.
It drives home the point that websites are abstractions that have no necessary relationship to any particular physical hardware. Network tools may or may not tell you a bit more about the source, depending on if there are any leaks in the abstraction.
However the last question is a fair point - nobody complains about CloudFlare's caching of your web page as you designed it.
The critique of AMP is that it receives privileged placement in search results, and that content authors are being pressured into adopting this de-facto Google-controlled spec, where they host your content and control its presentation. Anything that furthers AMP helps Google in this effort.
Showing the name of the "signer" in the address bar, instead of the server where the content is actually hosted goes against decades of browser UI design.
Good on Mozilla for marking it as harmful.
Does it though? If you use Cloudflare or Akamai or Cloudfront or Netlify or etc. etc. then what shows up in the URL bar is not the server where the content is actually hosted. Well, it is the server where it is hosted, it's just one of the many domains hosted by that server.
Just because Google invented it doesn't make it bad.
It changes the meaning of the address bar from "this is who I'm talking to" to "this is who (at some point in time) signed this content".
I don't mind that you can sign and verify content, that's fine and useful. I'm just not a fan of changing the address bar's meaning.
What I'm saying is that this does not change the meaning of what's in the URL bar. It's the same as before. It tells you who published the content originally.
No, it tells you the origin of the document. If you are the creator, and you choose to put your content on server X it will tell you "I've got this from server X". Whether that server is a reverse proxy or a shared webhost or a dedicated server in a DC or a raspberry pi running on your desk doesn't matter - it's the designated original that you, the owner of example.org chose.
That's what it always meant, and it changes when you do a redirect, and it shows you the current URL even if there is a canonical header of http-equiv. I can put a reverse proxy on my host and proxy example.com to example.org - the address bar tells you that you're reading example.com, not example.org, as it should, because you're connected to me, not to example.org.
How is it more secure? If, as you say, DNS can be spoofed easily - I can easily get a certificate issued with the required extension and make a "cryptographically signed package".
Spoofing DNS to clients is much easier than spoofing DNS to certificate authorities. Otherwise domain-validated HTTPS certs wouldn't mean much.
Do a trace route on any domain and you'll see that the server isn't the one that give you the answer, but some intermediary. Sure in that case when you did the request, the content is fresh and the server answered RIGHT NOW, but that cache still get the content from the server, it's just a bit older.
A browser already doesn't show you what server delivered the content. That would be your wifi AP, cell phone tower, or ISP node. The internet has already long established that we can trust content without trusting intermediaries.
There are two elements that are important: integrity and privacy. The content integrity is protected via a digital signature, the "signed" part of "signed http exchanges". The signature proves that the document hasn't been tampered with.
Regarding privacy: The intermediary (a search engine in this case) already has the content being delivered as a result of crawling it. It also knows the user clicked on a link to get that content, and knows the user's ip address. Even without AMP or Signed Exchanges, the privacy situation is the same. Once the page is loaded, all further interactions with the origin are normal https traffic, so later requests are not different in privacy either.
What this enables, for search results, is the ability to load the bytes of the content before the user clicks a search result. If the browser prefetched those bytes with the origin's awareness, then the user's privacy with respect to the search query would be violated, making prefetch problematic. With this setup, documents can be prefetched while preserving user privacy and after the user clicks all browser behavior continues as normal from that point forward.
By forcing web publishers to host their content on a Google cache, they lose their server-side logging and the ability to determine how they set up they way they serve their own sites.
Also, why do you artificially slow page loads on AMP pages to 8 seconds when JavaScript is disabled? That is a privacy issue.
You misunderstand the 8 second CSS animation in the AMP boilerplate. Here's the code (simplified):
<style>
body { animation:-amp-start 8s steps(1,end) 0s 1 normal both}
@keyframes -amp-start{from{visibility:hidden}to{visibility:visible}}
</style>
<noscript>
<style amp-boilerplate>
body{animation:none}
</style>
</noscript>
See the noscript section: if javascript is disabled, the CSS displays the body immediately. If Javascript is enabled, but for some reason the AMP javascript fails to load, after 8 seconds, the page is displayed anyway. The page is probably somewhat broken without the javascript loading, but the 8s is a fallback, not code to slow down non-javascript browsers.An 8-second delay seems like an intentional "bug" to coerce users to turn on JavaScript (and advertising).
That is not the intention. If javascript is disabled entirely, Google Search won't even load AMP pages. The scenario you describe of a user loading an AMP page directly without javascript enabled is somewhat rare.
I don't have javascript blocked, but I do have Google's tracking blocked via standard tracking protection (which is now a built-in feature in most non-Google browsers), which means <noscript> tags are not triggered, and I get the 8 second delay due to non-loading JS resources.
I don't think my setup is as rare as you make out.
Just from the text of the pages you visit they can build a profile around you. What your interests are, how much of an article you're likely to finish, whether you're the type of person to highlight text as you read, etc.
Unless you live on an island with a poor satellite connection AMP is useless as anything more than a corporate user data collection tool.
Although your point is well taken that there could be ways to sneakily track users eventually despite the aforementioned measures, and potentially even without javascript being required (though I doubt that share of privacy-concious users will ever raise significantly - most people simply don't care).
Their AMP cache happens only on their search service. They already know which links you click... having an AMP cache on top doesn't give them MORE information than they already get. The use of that cache also make sure the website doesn't get more information because it's preloaded.
If the publisher chooses, they can send logging to Google Analytics, but this is not part of AMP.
The typical argument otherwise is that the AMP javascript is loaded from Google's cache, however these javascript resources allow for a very long cache lifetime (1yr if the page came from the Google Cache), so relatively few page loads will actually end up fetching them from the network for most users.
Edit: These resources are also on cookieless domains.
Christ this is thin as a privacy argument.
They might not now, but could ‘t Google start creating unique URLs on each page, allowing them to track you that way?
Is there anything preventing Google from changing this later?
> The Google AMP Viewer is a hybrid environment where you can collect data about the user. Data collection by Google is governed by Google’s privacy policy.
I assume they collect information from HTTP request the browser sends when requesting an AMP page.
[1] https://developers.google.com/search/docs/guides/about-amp#a...
A cell phone tower or ISP node is ideally just infrastructure, "plumbing". Google seems to be trying to advance their strategic position in that direction. Rather than just being one search engine among several, they are trying to become part of the infrastructure. This could prevent future privacy solutions (and even prevent competitions between search engines).
No. Incorrect. Completely backwards. Factually wrong. You just failed your networking-exam.
Those things you mentioned would be transparent networking nodes forwarding your TCP-packets and they have nothing to do with any layers above that.
The fact that you don’t even know this completely invalidates any other point you may have.
So this Amp exhange technology changes nothing in this regard. It's like Google provides its own Free CDN, it is just not done in a traditional manner.
Which is plain wrong. I care.
When the URL-bar says I’m looking at company.com, I expect my browser to have used my OS’s DNS-resolver to look that name up, connect to the IP-given and nothing else.
I certainly don’t expect it to send traffic to certainly-not-the-nsa.com which are MITMing my traffic and tracking/monitoring it.
If I can’t trust my browsers URL-bar to exclusively and accurately reflect what is actually requested, it is effectively lying to me, the user, it’s owner.
And then suddenly all URLs are phishing URLs because Google made URLs no longer matter or mean anything.
Completely unacceptable.
The proposed scheme is just another way to extend this kind of relationship that the publisher builds, a new mechanism if you will. There is nothing in there that requires more or less trust from your part than before.
You're complaining that need URLs to reflect what is requested - in fact, I argue that you want the URL to tell you what is being served. But this is not what's currently happening.
URLs are already lying to you.
I doubt that you WHOIS-lookup all DNS resolved-IPs to verify that the IP presenting a cert is assigned to the organisational entity that you want to connect to, and have a whitelist of those entities that you actually allow your browser to connect to. Because that's what currently required to make sure you don't go through CDNs and other intermediaries between you and the publisher.
In which case the URL serves what was requested.
What AMP does is provide google.com content and lie to the user and says it comes from company.com.
Which isn’t true, and it only does so for users coming from google.com. Where I’m sure google will be happy for the additional tracking data.
This is NOT the url the user was lead to believe he requested. This is not what everyone else is served.
This is malware.
> With this setup, documents can be prefetched while preserving user privacy and after the user clicks all browser behavior continues as normal from that point forward.
But Google can already preload and show cached version of the page without this spec. The only difference would be that address bar shows "google.com" instead of publisher's domain. There is no need for this specification.
Only if you load the page from a Google SERP, in which case, Google would already know if you visit the page. If it's loaded from a Bing SERP, it's served from a Bing server, and the same for Baidu and other AMP caches. This is far more privacy preserving than preloading a page from some third party web server that the user might never visit.
Yeah, I didn't think so.
Does that mean that the Google+ button is coming back? Seriously? Why not just serve the content and leave it at that? Is the tiny bit of extra data you get from a unique "share on Facebook" URL worth it?
> The Navigator.share() method invokes the native sharing mechanism of the device as part of the Web Share API.
The fact that it has seen such adoption is testament to Google's ability to influence with it's rankings alone.
Sure, having a very quickly opened page is nice, but on the other hand, features are limited. That might or might not work well, depending on what kind of content you have, what engagement you're looking for.
Just you wait until you notice you can't go to town in your car any more. Only teslas are allowed into city.
But anyway, that's off topic. I understand that it's a pain for developers but for users like me who are often on a bad connection it's a life saver.
AMP introduces a very non-Appley top bar within the browser, adds new swipe semantics that can be confusing, breaks "tap status bar to scroll to top" behaviour, breaks reader mode (although this is inconsistent), and generally looks out of place. The best way to describe it is like a GTK or KDE app running in macOS. It's clearly not a "native" experience and doesn't really look or act like any other webpage in mobile Safari.
Type "amp sucks" into a search engine to find out more.
One could imagine replacing the serving layer with something like BitTorrent or IPFS.
Also, sites that link to other sites can preload the linked site's content into the user's browser, without leaking the user's IP to the linked site, so if the user doesn't follow the link, nothing about the user is revealed to the linked site. That sounds like a performance and privacy improvement wrapped up into one. I'm finding the rest of this discussion thread extremely disappointing as it seems like most of the posts here are just "amp=bad and amp people like this so it's also bad".
What's really needed is a way for the browser to lock a given web app/package to a specific version (and hash), so that even if the signing key becomes compromised, the app can't auto-update to a newer version containing malicious code.
Combining this with something like Certificate/Binary Transparency would allow browsers to check that they are not being uniquely targeted with a specially altered version, and you could set a policy saying "Only auto-update to a newer version of this web app if its hash has been published in a log for more than a month (and/or endorsed by signatures from N out of M other organisations I trust)".
Another idea it could be used for a wifi "drop box" (drop station?) when there's no internet connection around. That isn't uncommon at some popular spots up river into the woods in the US.
The idea is that as people enter the area, they can update the drop station automatically for things like news or public posts with whatever they've cached recently.
I'm pretty sure I read about this idea before the spec was drafted but I couldn't find or remember the site, something like vehicle-transported data.
One thing to note is that the specification currently limits the lifetime of a signed exchange to 7 days. It's possible that by exploring some of these use cases, especially offline, the spec could be improved with respect to some of these constraints.
It's a very narrow spec designed just for AMP, basically.
It's so obvious why AMP is on the main google.com domain.
They're collecting your shit.
They're doing the same thing with ReCaptcha, also on the main google.com domain.
Break this shitty company apart.
This fixes the main UI issues with how AMP is currently used in google search - mainly the url not properly showing where the content is.
If secure exchange is treated the same as AMP pages in google search, I.E. SXG content will be preloaded whether or not it's AMP, it would get rid of the second complaint of AMP - that google's preloading of the content is an unfair playing ground and that the only reason it's fast is because it's preloaded.
If SXG is treated the same as AMP in the carousal then that would fix the last and most serious complaint about AMP.
As far as I can tell, google do seem to be moving in that route, so this should be applauded not derided (the original fiasco that is AMP non-withstanding)
The new feature is that Google's browser displays your domain, obscuring the fact that Google is doing the serving. The change is what is displayed, not the server.
What Google actually means here is "We make AMP pages _appear_ to come from your own domain".
That's something entirely different.
This whole thing is just more doublespeak.
When I had a website with embed videos from other sites, I had user contacting me because the other sites had some problems. They couldn't tell the difference between megavideo/youtube/dailymotion content and my site, so they came to me and blamed me.
So what this means is that not only Google bullies you into putting your traffic under their control, but now, any problem on their part will be blamed on you by the user.
I hadn't even considered that. Add to this Google's notoriously absent customer support department and you have a recipe for a lot of frustration.
So Google’s browser now directly lie to the user about what’s being loaded?
I bet the SSL mark is still there though?
How can anyone trust this Googlan horse?
This is why you don’t make a browser and control major web-assets at the same time. These lines should not be muddied.
Faster than AMP, more open than AMP, and all the benefits of AMP.
Conceptually, you can think of a signed exchange as a 301 redirect to a new URL which has already been cached by the browser (so there is no 2nd network event). The cache was populated by the contents of the signed exchange, assuming the signature validates.
Also, their algorithms mistook live streaming of Notre Dame fires as 9/11 incident. How can live stream be a past incident?
https://abcnews.go.com/Business/youtube-mistakenly-flags-not...
"YouTube Live is an easy way to reach your audience in real time. Whether you're streaming a video game, hosting a live Q&A, or teaching a class, our tools will help you manage your stream and interact with viewers in real time."
Youtube live is supposed to be streaming of live events. My point is that their algorithms incorrectly decided a live stream was a past event. Even google/youtube acknowledged it was incorrect to tag a live event.
Amp is designed to tighten googles grip on the web, nothing more.
The only way Google could proactively "solve" this problem was by creating a "standard", and then also offering to absorb end user traffic for sites that adopted the standard. FWIW, AMP is an open standard not solely owned or contributed to by Google.
I've yet to see a Google engineer, executive, or "fanboy" address this question adequately.
This thread will be no exception. Queue the crickets.
If visibility is influenced by AMP then Google benefits, users using Google services likely benefit, web developers suffer, users not using Google services to view the content continue to suffer (because companies will continue to maintain two versions of the website, a bloated version with 100 external tracking requests that will be shared on twitter/reddit/facebook/hn/etc, and an AMP version that will only appear on Googles services), and the internet as a whole suffer. Whereas if visibility is influenced by page speed+external requests then everyone would benefit.
- AMP is a transparent and unambiguous standard that leaves no uncertainty as to whether you are somehow "performant enough" to qualify for the simple but limited visibility boost (referring to the news carousel)
- AMP prevents important usability problems beyond performance, like page content jumping
- AMP can enable advanced/extreme performance optimizations by default that are somewhat rare in practice (eg. only loading images above the fold) or isn't really possible to do safely/properly without a spec like AMP (eg. preloading content before the user clicks the link without unpredictably disrupting the website's servers) or sometimes avoided due to cost (eg. fast global caching with Google's impressive CDN). Important for users in the developing world.
Addressing your other points:
- Users who don't use Google services don't suffer. AMP is not Google-exclusive, all the major search engines (like Bing, Yahoo, Yandex) are stakeholders in the AMP standard and are free to support AMP. AFAIK there is nothing in the AMP standard that favors Google over other search engines or any other platform that might support AMP.
- Not sure how web developers suffer more from AMP. I'd think web developers would suffer more from trying to wrangle their bloated website performance independently rather than use a standard toolkit that enforces best practices and enables difficult/expensive optimizations out of the box.
- It's not clear to me how the internet as a whole will suffer, but I suspect this is just general hyperbole and not a specific point.
amp.dev is owned and controlled by Google. ampproject.org is owned and controlled by Google. The core AMP team are Google employees.
How can you possibly say it's not owned by Google?
https://blog.amp.dev/2018/09/18/governance/
The TSC is independent and, at this point, the committee code commits are almost 4x the volume of Googler commits.
2. The privacy policy on amp.dev is Google's.
3. Both sites are hosted by Google.
4. The license in the amphtml repo says copyright Google.
5. The OWNERS.yaml file in the amphtml repo list 3 people, all of whom work for google.
6. Per the contributing code readme, contributing code requires signing this Google CLA: https://cla.developers.google.com/about/google-individual
7. Looking at the last few merged PRs nearly everyone involved is a Google employee. I realize this could be a coincidence but I'm not going to analyze the whole repo.
8. The TSC is 3/7 Google employees.
Regardless, until Google issues a legally binding release of the project to an independent organization it is owned by Google. The TSC and AC could be removed at Google's whim.
The performance will be greatly affected if you run some cancerous theme with endless JavaScript calls. But both of the mentioned "engines" have changed the way I see blogging with WordPress.
Best of all, this is accessible to your average user as well. DigitalOcean can spin you up an OLS instance in a minute or so...
It's absurdly slow, uses tons of unnecessary JS, and it is a privacy nightmare because now I can't just use server-side GDPR and ePrivacy guideline compliant analytics anymore, but either have to give up analytics entirely, or have to use privacy-obliterating Google Analytics.
And if a user ever loads the page with JS disabled (which all my sites are designed to support), AMP breaks and just shows nothing at all for over 8 seconds.
On a mobile device in India? Nonsense. Your page load time is dominated by latency, which the AMP user doesn't see because it is preloaded from near caches.
> uses tons of unnecessary JS,
Which of the JS is unnecessary? The JS to load images allows AMP not to preload images below the fold, which is absolutely necessary for speed and for being friendly to data plans.
> now I can't just use server-side GDPR and ePrivacy guideline compliant analytics anymore
Explain. You still get first party tracking that gets fired when the user clicks to your page and can get user consent via data-consent-notification-id.
> And if a user ever loads the page with JS disabled
In that case, it's the SERP's fault for showing the AMP page instead of the non-AMP page. In the normal JavaScript-enabled scenario, the SERP would be stupid to show your non-AMP page.
My test device is a Huawei Ideos X3 on a 56kbit/s throttled 3G connection. The same effect also applies with a Pixel 1 on the same connection, or either of the devices on a modern 3.9G LTE connection. (Tested on O2 net in Germany, works reliably better than AMP even and especially while on a train — if you've ever tried using O2 on the intercity train between Hamburg and Münster you know that every third world country has better internet than, I've seen 8kbps with 13 seconds latency there)
> Which of the JS is unnecessary? The JS to load images allows AMP not to preload images below the fold, which is absolutely necessary for speed and for being friendly to data plans.
AMP uses megabytes of JS for that purpose, I do the same in under 1kiB (even including an intersection observer polyfill). And my CSS is much much smaller as well. Part of why I get a 100/100 in all pagespeed and lighthouse tests, including when simulating mobile connections, while AMP pages get only 60/100.
> Explain. You still get first party tracking that gets fired when the user clicks to your page and can get user consent via data-consent-notification-id.
I want JS-free analytics that do not require tracking or any consent (GDPR allows collecting some information without consent, same with the yet unreleased ePrivacy directive with which AMP is not compliant anyway).
What? Where are you pulling these numbers? Also, what do you mean by hot cache? I'm starting to suspect that you don't even understand that the AMP page (the JavaScript for sure, and often the entire HTML and above-the-fold images as well) is already on the user's device, while your page is not.
Obviously, this part is affected by the AMP js being in cache or not.
Still, often my own page can load faster than just this user-visible part of loading the AMP version.
AMP works best when the user visits almost only AMP pages (so the resources stay in cache), and the user has a high-latency high-bandwidth connection.
But that's almost nowhere in the world true, in reality most people have relatively low latency with low bandwidth.
Your claim that your page loads faster also reeks of wishful thinking. Pretty much every AMP page I have loaded from a SERP loads instantly, not just fast. For someone on a worse connection, the page will have started loading before the user clicks on the link from near caches versus have not started loading at all from a far server. In the rare case where the AMP JS is not in the browser cache, it will be after loading the first result.
If you say pretty much every AMP page you've loaded has been instant, please post the specs of the devices and network you've been using for testing.
Additionally, if the latency between the device and the nearest server is over two seconds, the latency to a far server as well as the click latency don't even come into play anymore at all, instead the number of connections needed becomes much more important, and bandwidth also becomes a much larger factor.
Your claim that HTTP/2 would have worked towards better latency on lower connections is also false, on bad mobile connections HTTP/2 actually increases latency, which was a major reason for QUIC aka HTTP/3 in the first place.
And as I've mentioned, you've been testing the wrong thing by not understanding the whole point of AMP (safe preloading).
> instead the number of connections needed becomes much more important
A page preloaded from an AMP cache needs at most one TCP connection, usually zero if it uses QUIC.
> and bandwidth also becomes a much larger factor.
Which also works in AMP's favor because the device doesn't need to load your custom JavaScript or potentially unoptimized images, just the tiny HTML and optimized images above the fold. The weight of this (and the associated gain) is tiny, which is why bandwidth is a relatively unimportant factor.
> On bad mobile connections HTTP/2 actually increases latency
You're mixing up dropped packets with high latency. That's neither here nor there because Google's and Cloudflare's AMP caches both use QUIC — my point was that latency is the key factor that all modern web speed technology has attacked, including AMP.
Basically this. AMP sets a hard upper bound for how fast your webpage can be. Have a purely static HTML+CSS blog but want to get the page rank boost from AMP? Just add reams of unnecessary Google Javascript to what should be a very simple site.
yeah, there's a few sites out there that are faster than amp. but most of them are not, and before amp the trend was certainly not to make anything lighter or faster.
I'm frankly shocked that people can't see this land grab for what it is.
Peerweb helps sites automatically offload all resources (including streaming ugc video) to a decentralized p2p network.
ETA: ~1yr
It also at least helps slightly address one of the complaints of publishers, which is that cookies and some analytics will work now.
But it still doesn't address the biggest complaints of publishers.
I'm guessing Google cares a lot more about the user experience than the publisher experience, since users make up most of the traffic and all of the ad consumption, so this is certainly good for them!
Split Alphabet already.
Cache-Control: no-transform[Working on AMP at Google]
If I understand this right then this seems to open up some doors for some new email phishing scams.
Someone can buy a lookalike domain name using similar-looking UTF characters, send out a bunch of email spam with an URI that looks like the original, and once the user visits the webpage it instantly loads the AMP and suddently the URL is authentic. There will only be a very quick url change from the punycode url to the original that I doubt many will notice.
I'm not entirely a fan of this, though it could have uses in other places;
Debian for example, could use this to enable HTTPS-like security without requiring mirrors to upgrade their security to the late 2000's.
If this is so, and it only works for google search's page and not some generic web strategy, this is a massive breach of browser/web-content separation, bigger then even the auto-sign-in to google from chrome.
Conceptually, you can think of a signed exchange as a 301 redirect to a new URL which has already been cached by the browser (so there is no 2nd network event). The cache was populated by the contents of the signed exchange, assuming the signature validates.
There is no "reaching into the document" from the previous click or anything weird like that.
Implementing AMP while keeping my website URL is great. Glad to hear that Cloudflare is already supporting this.
Will implement it now in my sites.
Don’t use Google search.
That’s all. That’s really all there is to it. You should give it a try!
One example: https://news.ycombinator.com/item?id=19679136 , luckily someone also provided the non-amp equivalent...