Improving URLs for AMP Pages
amphtml.wordpress.com
amphtml.wordpress.com
AMP should have been purely an open source library implementing a specification, not a way to opt-in to becoming a sub page within Google
If I go out to a new URL it should then cut away the relationship with the old site, even if AMP makes that transition a bit quicker and gives the new site tools to load the page more quickly
I read the post top to bottom and it never mentions anything but changing the url. It doesn't say that top AMP banner thing is going away.
That's how they're able to display the original page's URL in the browser without that practice being extremely misleading.
If the bar is going away Google can help themselves by saying that specifically. I'm not the only one that hates that thing.
In particular, one obvious way to implement this would be to have everyone sign a static webpackage that includes a JS loader that pulls resources from www.google.com/amp/ via XHR/CORS (using ServiceWorkers I guess?), which would only need to be signed once, and send that over to Google. Then the infrastructure continues to work the way AMP does today - you just upload actual AMP-compatible web pages to your actual website, and Google downloads and caches them. Probably this webpackage would just include a single script tag that pulls the loader itself from Google, so that Google can apply updates to the loader without having to bother you again.
The webpackages spec talks more directly about signing the actual data, which would be way better for the web, but it's not obviously better for Google and for publishers already invested in AMP, so we should be cautious that maybe Google does not mean this and means something that's easier for them and retains the lock-in.
And in terms of functionality, does the AMP web version have all the functionality of the normal site? Often this is something I see missing on AMP sites. A top-bar gives you the ability to get back to the full site.
I think it's the other way around - webpackages include Expires: headers for each resource that are signed by the origin, so the browser decides whether to trust it or not, or maybe to render it anyway but show a staleness indicator. But if you're using a CDN, the CDN implements caching on its own (hopefully following the origin server's instructions or the customer's instructions in a control panel, but no guarantee) and all headers are entirely controlled by the CDN.
> And in terms of functionality, does the AMP web version have all the functionality of the normal site?
On the technical side, this is now the origin's decision, not Google's - they can include whatever functionality they want or whatever links they want in the webpackage. It's not really different from having a mobile site that has a "View full site" link, or worse, doesn't have one and also misdetects your browser.
I think I agree that in practice it matters a lot what people choose to do, and whether Google adds restrictions for what sorts of webpackages count as "AMP" for the purpsoe of special treatment on the search results page.
(Also, if we think this is worth indicating, it should be done by the browser itself, not by the web page voluntarily including some CSS.)
AMP is a pest. A textbook example of all things wrong with the web today. I'll spare you my thoughts because it would only be filled with hate, foul language and insults. Good luck with it though.
Use the Web as it was intended, pure HTML/CSS and pages will be blasely fast.
Hmm, so maybe in some future static-only sites will be able to sign a bundle with offline keys and not use TLS at all. Or maybe we just sign static bundle with a TLS key for our origin and upload the bundle to Google and other web caches. As in maybe the internet can be distributed again.
I see lots of interesting potential in decoupling origin verification from TLS connections.
Web Packaging Format Explainer: https://github.com/WICG/webpackage/blob/master/explainer.md
I'm imagining a world where a static site generator + Let's Encrypt automatically generates these webpackages for you and uploads them to Google, and also DuckDuckGo, Bing, etc. And when someone creates an HN or Reddit post, the HN / Reddit server checks to see if there is a webpackage available for that page and if it's not too many kilobytes, and serves that when you click on the link. That seems like it delivers on the original promise/vision of AMP, and also gets entirely out of the vendor lock-in problem that AMP has today and into something genuinely novel and good for the web.
I think there's a need for another layer of security in certain types of web applications (secure messaging, financial, cryptocurrency). It sounds like this is a "nice to have" use-case: https://github.com/WICG/webpackage/blob/master/draft-yasskin...
I've almost been able to abuse Service Workers for this purpose. After the Service Worker is installed on first load you can intercept any requests (including the HTML and JavaScript) and verify a signature on them using a key embedded in the Service Worker. The big problem is it seems like you can't reliably prevent the browser from trying to update the Service Worker itself, which is obviously a deal breaker. Also shift-reload bypasses the Service Worker completely.
These two approaches are built for different threat models. Both protect you from tampering, but only TLS protects you from collecting metadata like what exact page you visited. Attackers can only observe which domain you visited.
My thought was that it as a static site owner it would awesome to keep TLS keys offline. And only use them when content is updated. That way a server compromise is not so bad.
The spec also encourages/opens up the possibility of exchanging those bundles over peer-to-peer networks instead of HTTP, which further mitigates the threat of over-the-shoulder metadata collection.
Poe's law is very strong here.
Why anyone think it's the slightest bit professional to use it in official communications is mind boggling. It was sort of trendy for a little bit on Reddit with a certain type of abrasive and annoying American, but thankfully it seems to becoming uncool again.
It is also an extremely imperialistic word. Americans use it, it feels as if it only refers to Americans, and Google is only listening to Americans.
"youse" falls into the bad pattern of also picking up "guys" or "all" as hangers on, in my opinion, defeating the purpose. ("youse guys" being the terrible patriarchic movie Mafioso cliché, and "youse all" a terrible Frankensteinian monster I've heard far too frequently.)
"yinz" to me looks and sounds more like a weird pharmaceutical than an English word, and y'all aren't going to convince me otherwise. ;)
But of course, my opinion is biased by geography and familiarity.
Enough, Google. Making small web sites is EASY, OK? No AMP needed: just write your content and, as if by magic, it is small and loads nearly instantly. If web sites are bloated and slow, close them and use something else. Stop hyperextending the web to make lousy programming practices the norm.
Google should just more explicitly rank pages based on load times and incentivize sites to fix their crap.
Apparently they already did this and it had no effect. It seems like anything less than a binary signal (in the carousel vs. not) is too subtle for publishers to understand.
I believe AMP pages have no JavaScript.
As my boss said when I worked to a division of RELEX - the main problem is that publishers is that they treat their technical experts with contempt.
But in reality, Google tries to racket the free and open web in order to squeeze even more juice to feed its insatiable corporate greed.
Bonus Quiz:
1. When I encounter an AMP link I... a) Click it. b) Don't click it. c) I don't see AMP links using FireFox.
2. When my friends send me an AMP link... a) I click it. b) I don't click it. c) Friends don't send friends AMP links.
3. Reasons I've switched to FireFox... a) I love RUst. b) I care about privacy. c) I hate AMP.
4. My ISP is... a) Google b) Chrome c) None of the above.
The answer is 'c'.
Thinking about it more, why is it even a problem if the pizza shop across the street knows I searched for pizza and wants to try and offer me a better deal before I pick up the phone and call someone else? If I buy into having my privacy violated by allowing sites and ad networks to track me anyway, wouldn't the better user experience there be to have services competing over me (my business) as soon as I start looking rather than once I've entered their store? Kinda playing devils advocate but now I'm actually curious.
AMP is a standard that restricts webpages to a subset known to load very quickly, which is especially useful if you have a mobile device, or are in a country with poor internet such as in the USA.
Here in this thread are people talking about how much they like AMP from a user's perspective.
> Why can't Chrome just prefetch shit from the actual servers
The linked article answers this: for privacy reasons.
> and let ISPs handle the caching?
HTTPS does not allow ISP-level caching. This is generally a good thing; I trust ISPs significantly less than I trust Google.
> Why does mother Google need to serve me all the content from...?
Performance, presumably.
> I know everyone working on AMP means well but why why why does Google insist on destroying the internet and entirely undermining TLS in the process?
I don't think they're doing that.
It is called HTML/CSS with zero JavaScript. Quite fast in dial up modems.
That's not all it does. No-one would object to it if that was all it does.
The more important part is that it hijacks content and serves it from other servers, and requires including a js file from a large corp in every page. That's a massive vulnerability waiting to happen, but it also gives complete control of the web to whoever controls that js.
They need to ditch the requirement for js, and ditch the requirement for framing with Google junk around pages. The web is an open ecosystem, that's its strength.
Also, google should not be using their influence in search to push changes which are profitable for them - that's abusing their monopoly position.
The "hijacks content" part is by itself unobjectionable, especially with this latest update we're discussing which fixes the URL bar issue. The other server will serve a checksum so Google can't tamper with the contents. That makes it just a free CDN.
Is the problem that you do not want Google to see your content at all? You'll need to use robots.txt to ban Googlebot. You can't simultaneously want to appear in Google search results and also not let Google see your website.
The required JavaScript is legitimately frustrating, I know. The AMP project has an article about why they did it that way:
https://medium.com/@cramforce/why-amp-html-does-not-take-ful...
Specifically, it's to prioritize resource loads. I personally don't think their explanation is very convincing. But whatever, maybe it's easier for them to do it this way or something.
I don't think Google is doing it to push changes which are profitable for them, though. I legitimately believe they're doing it to make the pages load faster and otherwise be better for the users. I don't even understand how it could be profitable in any other way.
Strategically, for google, owning the frame around the web and a bit of js on each web page is vastly more valuable than customers having faster web pages.
So no, speeding up web pages is not why AMP exists.
If anyone is interested PM me.
The linked article explains this. Do you really want third party sites to know what you're searching for without you actively deciding to click on their page?
> and let ISPs handle the caching
ISPs can't cache the contents of HTTPS pages.
My user agent /is/ me. If I want to prefetch results, I unerstand it means people know I'm looking at those results. I don't think it's really the privacy issue they are trying to make it out to be. And the counter consideration is that you are implicitly saying you trust Google with that info more than the actual service providers so it's not "do you want people" it's "who do you want" knowing.
Easy solution: off by default and inform users of the implication when turned on. As a user I don't actually want the internet prefetched for me mostly because it's a stupid idea catering to a subset of people who are so impatient it's baffling. If your site performs poorly on mobile let users complain to the creators so they can make better sites. Does Google really get the schtick when a mobile site has a poor UX? No. So why are they even investing effort into this?...
> ISPs can't cache the content of HTTPS pages.
Indeed. That's the point of HTTPS. As a user I don't expect that contract undermined even by a claimed benevolent Google just trying to altruistically get you content better with signed bundles that masquerade around as the original service. A few points, too. Often the static CDN content is served HTTP because it's not sensitive or it employes other means of restricting access such as GUID/signed urls etc., all of which can be cached. But even if that's a bad idea (I think people are moving away from that model, and http2) that just shifts the responsibility to the ISP and CDN providers to make sure that they have good fast connections between their networks. That's something that is in both their business models and it works just fine today. ISPs want to provide the best user experience on their network compared to others and CDNs sell their highly available global network to customers so a better CDN closer to more users means happier customers.
Why do that though when you can have your cake and eat it? Google's proposed solution allows you to prefetch results _and_ not allow the third-party to know you're looking at those results.
> the counter consideration is that your implicitly saying you trust Google with that info more than the actual service providers
If I'm using Google search, then yes, obviously I'm fine with Google knowing what I searched for. If I wasn't okay with that, I most certainly would not be using Google search. In contrast, I'm less likely to be okay with a random site in the search results page that I haven't clicked on knowing that I saw a link to their site in a search results page.
Note that if I were using DuckDuckGo instead and DuckDuckGo supported AMP, then my browser would prefetch from DuckDuckGo's AMP cache, not Google's. No additional information is being shared with any party who doesn't already possess that information. (DuckDuckGo already knows what I searched for. Me loading an AMP page from them related to that query reveals no additional information.)
> Indeed. That's the point of HTTPS. As a user I don't expect that contract undermined even by a claimed benevolent Google
Could you explain how Google's proposed solution here undermines HTTPS? Note that the OP talks about using the upcoming [Web Package standard][1] to distribute AMP pages. This standard would allow the integrity guarantees of HTTPS to be preserved even when the page in question is being served by Google's AMP cache rather than the original server.
re: web packaging, if a new standard is being developed that will set user expectations around content delivered by a third party then fine. But to boot I'm really confused why the content requested is not the result of loading the certified URL in my address bar (which TLS has conditioned us to be privy to). And so users will become accustomed to the implications of using a service that leverages the standard and there shouldn't be any long term confusion, I guess. But let's not ignore what it is: AMP is evolving into a CDN owned by Google only focused on serving Google search results which is only even feasible because it exploits Google's scale to further lock down its search monopoly and prevent information from leaking out that might even arguably benefit users. DuckDuckGo wouldn't get the same benefit from adopting AMP because they don't have loads of internet infrastructure to play with.
If you're using Firefox, for example, this change to the URL bar will only happen if Mozilla decides to implement the web package standard, after evaluating it and deciding its reasonable to display the URL of the original page for pages loaded using that standard.
> DuckDuckGo wouldn't get the same benefit from adopting AMP because they don't have loads of internet infrastructure to play with.
That may be true in the sense that DuckDuckGo can't deploy a global CDN as easily, since Google has more resources than them, but you could make the same case against almost any new web technology.
For example, you could argue DuckDuckGo wouldn't make as effective use of HTTP/2 as Google, since Google can afford to rewrite their applications to take full advantage of HTTP/2 server push much more quickly than DDG can. That's not a very good argument against HTTP/2 though IMO.
My main issue with it is that it the spec still contains [this line][1]:
> AMP HTML documents MUST
> [...]
> * contain a <script async src="https://cdn.ampproject.org/v0.js"></script> tag inside their head tag
The actual specification requiring JavaScript to be loaded into your page from any one particular server is unacceptable IMO. Hopefully this is fixed in future updates to the spec.
My other problem with it is much more minor; I just don't like the idea of serving a "special" version of my site just for mobile. I'm very much a fan of responsive design, and I'd prefer to just serve one version of my site that looks and works great no matter what device or internet connection you're viewing it with.
Ooooh, <raises hand aggressively> I know.
Google is fighting to minimize ISP incidental access to your internet actions, while aggressively expanding the amount of incidental access Google has to your internet actions.
Apparently they've "earned" their surveillance because their core functionality can't be divorced from their ability to conduct surveillance as easily as they can be separated in the case of ISPs.
"We embarked on a multi-month long effort, and today we finally feel confident that we found a solution: As recommended by the W3C TAG [1], we intend to implement a new version of AMP Cache serving based on the emerging Web Packaging standard [2]."
I'm just reading through this so I'm gleaning as I go, but it looks like the W3C TAG came out with a recommendation for 'Distributed and Syndicated Content' [1] that specifically addresses AMP by name, and recommends strategies to do this kind of content syndication in a way that preserves the original provenance of the data.
The Web Packaging Format [2] aims to, apparently [3], solve packing together resources, but, rather, HTTP request-response pairs, maybe HPACKed?, and signed and hashed for integrity, in a flat hierarchy, in a CBOR envelope, that nonetheless has MIME-like properties? I'm still digesting what's all involved.
[1] https://www.w3.org/2001/tag/doc/distributed-content/ [2] https://github.com/WICG/webpackage [3] https://github.com/WICG/webpackage/blob/master/explainer.md
Yes, that privilege is reserved for Google.
Google quoting privacy concerns is especially disingenuous, as they can track you and your activities across multiple devices.
> As we detailed in a deep-dive blog post last year, privacy reasons make it basically impossible to load the page from the publisher’s server. Publishers shouldn’t know what people are interested in until they actively go to their pages. Instead, AMP pages are loaded from the Google AMP Cache but with that behavior the URLs changed to include the google.com/amp/ URL prefix.
To me, this reads as "for our privacy, we don't tell the publisher what page has loaded" but that may be an uncharitable interpretation. I read the referenced blog post and it didn't clear up anything about the "privacy" issues.
Instead, the concern is sidestepped by the extra indirection: the user's user-agent will load the prefetches from the AMP cache.
How much of this is moot given that many browsers, including Chrome, offer speculative fetching as a feature, is debatable.
There's a difference between explicit prefetching (given by html tags, ie there's intent), speculative prefetching on the same origin (you already talk to them) and speculative prefetching across the entire net (you talk to somebody new out of the blue).
Preloading search results for faster display without a local Google-side cache means that more parties know that you (IP, User Agent, cookies) are potentially interested in certain pages due to a Google search (referer header).
With the AMP cache as currently implemented (and with the TAG bundles in a future version), Google gets to know that you just got the URLs A, B and C proposed by Google. Which is no additional information for anybody, at Google or elsewhere.
If this new scheme allows rolling back some of the less fortunate effects of AMP (the visibility of the AMP cache URL, the in-page URL bar emulation as a workaround to that), all the better.
AMP-enabled pages load faster, but on the other hand I have an ad-blocker and LTE that gets 10Mbps, so the improvement is negligible. Not worth breaking the web, IMHO.
And AMP is planning to use cryptographic signing so that the browser can verify that the CDN didn't tamper with the page.
In this future, it seems like the content will be signed with Apples TLS certificate/key and then the content can be distributed by anyone.
This could open the doors for a more distributed web too :)
It was not presented as a new and fairly dramatically different way of distributing content on the web, it was presented as a way to get around showing the URL in response to pushback from the web community.
I think if this announcement were written differently, I would've come away with a much better impression of the whole thing.
But the new door this opens in terms of content distribution are very interesting.
Shared caches might be a thing again... Who knows :)
Until you need to share a link, wonder why the page loaded slower than normal thanks to 100Kb of render-blocking JavaScript, or get phished or believe a spoof because it has google.com in the URL.
I really like the stated goals but shipping something with usability problems is a great way to get tarnish its reputation. Hopefully this new incarnation will live up to the original hope.
If Google allowed people to host AMP pages on their own servers and still show up on the Carousel, I would have no problem with AMP. As it is, though, AMP is a blatant power-grab from a corporation that's already got way too much control over the web.
Also I end up on AMP pages (generally the origin's AMP-reduced pages, not the Google cache thereof) way too often on desktop.
That said, yes, I still lean towards clicking on AMP sites and hoping that it's faster. But that's the irrational part of my brain, and the rational part would be happier with an additional 200-300 ms loading time in exchange for reliability.
In Google News, on a Google device running a Google browser on a Google OS, I can't even open a news article in a new tab because of a Google technology.
If that's not a happy path where things should just work, I don't know what is.
- AMP sites, much like most mobile sites, are frequently little more than less-useable versions of the full site
- AMP's forced Javascript transition and loading icon easily overwhelm any questionable loading speed benefits and break flow
- AMP URLs are unusable for copy-paste
- AMP adds further control to the Google web hegemony
- Most importantly, AMP is not optional. If Google was implementing AMP for the good of the user, rather than additional control and data on user activities, they would provide the ability to opt out of their invasive, hostile protocol.
If Google still enjoyed the trust they once did, AMP might be acceptable, though annoying. However, thanks to their continued campaign to monopolize and profiteer off internet users' data, any activities that might allow them to further do so must inherently be viewed in the worst possible light.
/rant
It is quite easy to make fast Web pages using standards without having Google messing up with the Web, pure HTML/CSS with zero JavaScript.
> while maintaining the [...] privacy benefits of AMP Cache serving
"AMP Cache serving" == hosted on Google's server. This makes this statement at best, stupidly oxymoronic, at worst, deliberately dishonest advertising.
> privacy reasons make it basically impossible to load the page from the publisher’s server.
Browsers (including Firefox[0][1]) already do this. There are no "privacy reasons" preventing this. The only reason not to do this is to present another justification for opting into their AMP Cache product.
> can take advantage of privacy-preserving preloading and the performance of Google’s servers
Also a contradiction of terms.
[0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Link_prefe...
If you are a news organization and Google won't let you be in a certain section without serving on Google, complain.
My understanding is that AMP exists to solve a problem. It obviously isn't the only way to solve that problem.
I don't believe this is the case, no. Unless how AMP works has significantly changed recently, it requires your site to be hosted on, and served from, Google's servers.
This hosting is called "AMP Cache". When AMP first launched, Google's servers were the only available "AMP Cache". They have since added the ability to set up your own AMP Cache, which seems like some attempt to appease people's concerns with AMP being Google-only, but doing so seems an utterly pointless enterprise because of the below (from AMP's docs):
> How do I choose an AMP Cache?
> As a publisher, you don't choose an AMP Cache, it's actually the platform that links to your content that chooses the AMP Cache (if any) to use.
So if you set up your own AMP Cache, this just means your site will be hosted on Google's servers, and served from Google search results on Google's, and probably noone will ever visit the copy on your own AMP Cache server.
Think of the Google AMP cache as an extended web site snippet on the google.com search, not as your primary delivery mechanism: To make this safe and useful for all parties (publisher, user, link service providers such as google search or twitter), AMP defines an html/css/js subset that is considered safe but still functional - so that (for example) analytics can still be made to work, which is important to publishers, but without working in the AMP cache hoster's domain context, which is important for their security.
Another AMP cache provider is Bing: https://blogs.bing.com/search/September-2016/bing-app-joins-...
I use an app for reddit, an app for hacker news, an app for YouTube, etc. If I stumble upon a URL of one of those sites I always want it to open in the app, since the UX is so much better there.