Show HN: Make your site’s pages instant in one minute
instant.page
instant.page
Temporary added it in-line for testing. I was already in the sub 100ms level but this just puts it over the top! Also updated all admin add/edit/delete/toggle/logout links with "data-no-instant". Pretty easy. Open developer preview and watch the network tab. Pretty neat to watch it prefetch! Thanks for creating this!
ps. Working on adding the license comment. I strip comments at the template parse level (working on that now).
pps. I was using https://developers.google.com/speed/pagespeed/insights/ to debug page speed before. Then working down the list of its suggestions. Scoring 98/100 on mobile and 100/100 on desktop. I ended up inlining some css, most images are converted to base64 and inlined (no extra http calls), heavy cache of db results on the backend, wrote the CMS in Go, using a CDN (with content cache), all to get to sub 100ms page loads. Pretty hilarious when you think about it but it works pretty well.
I'd expect this preload script to remember the pages it's already fetched and not duplicate work unnecessarily. :/
If you really care about speed anyway, you should already have setup your site to max out caching opportunities (Etag, Last-modified, and replying "not-modified" to "if-modified-since" queries) - I'd suggest the author should ensure the script does support caching to the broadest extent possible - hitting your site whenever appropriate.
ps. I do tons of caching for any images and videos though. I know those things never change so have weeks of caching enabled.
I'd be very curious to know what your performance looks like over time, especially as it relates to various improvements that you try out.
Seems to have no impact on any javascript, including ads. Pages do load faster, and I can see the prefetch working.
Just make sure you apply the data-no-instant tag to your logout link, otherwise it'll logout on mouseover.
Logout links should never be GETs in the first place - they change states and should be POSTs.
I just checked NewRelic, Twilio, Stripe and GitHub. The first 3 logged out with a GET request and GitHub used a POST.
Using a POST, especially if you're building through a framework that automatically applies CSRF to your forms, forecloses this possibility (unless you maintain a separate secret GET-supporting logout endpoint, I guess).
For example: Let's implement authentication, where a user logs in to your api and receives a session id to send along with every api call for authentication. The session should automatically be invalidated after x hours of inactivity.
How would you track that inactivity time, if you're not allowed to change state on get requests?
A GET request should never, ever change state. No buts.
Just because a bunch of well known sites use GET /logout to logout does not make it correct.
Doing anything else as demonstrated in this and other cases breaks web protocols, the right thing to do is:
GET /logout returns a page with a form button to logout POST /logout logs you out
Yes, I was a total n00b in 2001. But then, so was e-commerce.
The requirement on GETs is that it must result in no changes to the observed representational state transferred to any user: for any pair of GET requests a user might make, there must be no change to the representation transferred by one GET as a side-effect of submitting the other GET first.
If you are building dynamic pages, for example, then you must maintain the illusion that the resource representation “always was” what the GET that built the resource retrieved. A GET to a resource shouldn’t leak, in the transferred representation, any of the internal state mutated by the GET (e.g. access metrics.)
So, by this measure, the old-school “hit counter” images that incremented on every GET were incorrect: the GET causes a side-effect observable upon another GET (of the same resource), such that the ordering of your GETs matters.
But it wouldn’t be wrong to have a hit-counter-image resource at /hits?asof=[timestamp] (where [timestamp] is e.g. provided by client-side JS) that builds a dynamic representation based upon the historical value of a hit counter at quantized time N, and also increments the “current” bucket’s value upon access.
The difference between the two, is that the resource /hits?asof=N would never be retrieved until N, so it’s transferred representation can be defined to have “always been” the current value of the hit counter at time N, and then cached. Ordering of such requests doesn’t matter a bit; each one has a “natural value” for it’s transferred representation, such that out-of-order gets are fine (as long as you’re building the response from historical metrics,
> So, by this measure, the old-school “hit counter” images that incremented on every GET were incorrect
Yes they are incorrect. No Buts.
Two requests hitting that resource at the same exact timestamp would increase the counter once if a cache was in front of it.
The whole reason this is supposed to be the case is in order to enable such functionality as this instant thing.
In the old days it might have been acceptable to get away with a GET request but these days thanks to prefetching (like this very topic) it's frowned upon.
https://stackoverflow.com/questions/3521290/logout-get-or-po...
> In particular, the convention has been established that the GET and HEAD methods SHOULD NOT have the significance of taking an action other than retrieval. These methods ought to be considered "safe".
(there's an exception listed too, but doesn't apply to logout)
EDIT: I know of someone who made a backup of their wiki by simply doing a crawl - only to find out later that "delete this page" was implemented as links, and that the confirmation dialog only triggered if you had JS enabled. It was fun restoring the system.
What would be a good way to avoid prefetch requests in your analytics if you only derive analytics from server access logs?
For a non-JS solution, I guess an tiny iframe at the bottom of the html page that accesses a special server page with an unique stamp that causes the same "hit". the iframe loading would mean that the rest of the page is mostly loaded before it was closed.
How do you use this library with JS disabled?
You need to be emitting your analytics events from a rendered/executed page. Preferably with javascript, and a fallback <noscript> resource link to a tracking url can work here.
Considering a user hovering over a bunch of links and then clicking the last one, and doing this in a second. Let's assume your site takes 3 sec to load (full round-trip) and you're server is only handling one request at a time (I'm not sure how often this is the case, but I wouldn't be surprised if that's the case within sessions for a significant amount of cases). Then the link the user clicked would actually be loaded last, after all the others - this probably drastically increase loading time.
The weak spot in this reasoning is the assumption that you're server won't handle these requests in parallel. Unfortunately I'm not experienced enough to know whether that happens or not, but if so, you should probably be careful and not think that the additional server load is the only downside (which part like is a negligible downside).
I used to use a preload-on-hover trick like this but decided to remove it once we started getting a lot of traffic. I was afraid I’d overload the server.
About your first statement though, which server software do you use that still sends data after the client has closed the connection? Doesn't it use hearbeats based on ACKs?
Closing a connection to Postgres from the client doesn't even stop execution.
Unless you are focusing on the word server and assuming that has nothing to do with the framework/code/etc, then I can assure you it can be done. I’ve done it multiple times for reasons similar to this situation. I profiled extensively, so I definitely know what work was done after client disconnect.
Many frameworks provide hooks for “client disconnect”. If you setup you’re environment (more appropriate term than server, admittedly) fully and properly, which isn’t something most do, you can definitely cancel a majority (if not all, depending on timing) of the “work” being done on a request.
> Closing a connection to Postgres from the client doesn't even stop execution.
There are multiple ways to do this. If your DB library exposes no methods to do it, there is always:
pg_cancel_backend() [0]
If you are using Java and JDBI, there is:
java.sql.PreparedStatement.cancel()
Which does cancel the running query.
If you are using Psycopg2 in Python, you’d call cancel() on the connection object (assuming you were in an async or threaded setting).
So yes, with a bunch of extra overhead in handler code, you could most definitely cancel DB queries in progress when a client disconnects.
[0] http://www.postgresql.org/docs/8.2/static/functions-admin.ht...
I use nginx to proxy_pass to django+gunicorn via unix socket. I sometimes see 499 code responses in my nginx logs which I believe means that nginx received a response from the backend, but can’t send it to the client because the client canceled the request.
I admit I haven’t actually tested it directly, but I’ve always assumed the django request/response cycle doesn’t get aborted mid request.
If that's not how it works, it could easily be modified to add a throttle on how many links it will prefetch simultaneously.
Every site that can use a last mile performance optimization like this should already be serving everything from some form of cache, either from varnish or a cdn. So in theory, availability of the content should not be the problem.
There's also log analysis to identify the set of web pages visited by each employee during work hours, and an attempt to programmatically estimate the amount of non-work-related web browsing. This feeds into decisions about promotions/termination/etc. Prefetching won't get anyone automatically fired, but we'd still prefer it isn't a default.
I've heard a lot of stories of ridiculous rule-by-HR culture, but that's so extreme it sounds made up.
The other part about log analysis seems crazy, though, I agree with you on that.
That sounds batshit insane.
Of course, had I known about these practices in advance, I would have declined the job offer. But I didn't. I ended up quitting a few weeks later anyway.
IT would monitor all connections from all employees and send a report to upper management with summary statistics, on a monthly basis.
I was told this was the case by a fellow worker during my second day there, so I tunneled my traffic through my home server via SSH. When IT asked me why I had zero HTTP requests, I reminded them that monitoring employees traffic was illegal under our current legislation. Doing this in a university-like non-profit research center is hard to justify.
Couldn't you just say "I just dont use http anymore because this X company data is very valuable to me" ?
Invert the scenario: if they told you that you had to do work-based research on your own personal Internet connection, would that be OK? Any overage charges are yours to pay, no compensation.
In addition to being less user friendly, having the mindset that users/visitors to your website must live and work in ideal settings means that whatever you create will tend to be fragile and brittle because you don't try to take into account situations that you haven't seen before.
I mean that firewall was used to track every website you browse and other evil stuff.
I always let my team browse Facebook, if they wanted. One of my top people browsed it the most out of everyone. If you block a page, then they will just use their phone.
If you are going to measure, then measure outputs. Measuring inputs will make you and your team equally unhappy.
The webbrowsing of other people is their private matter. If you think someone is surfing the web too much (or taking too many coffee breaks, or leaning on the shovel for a minute or any other normal activity people do to take little breaks from work), it's on leadership to tell the person to get back to work, or generally, create a work environment where work flows more naturally.
Logs are for investigations in case of crimes etc.
I read so many things here on HN that are illegal in Germany. We have laws and powerful worker representation that prevents dehumanitzing stuff like that but, often, things started in the US find their way over here....
If your company is doing that, then do not browse on company time and/or using company equipment at all. Ever. They obviously don't (or for regulatory reasons, can't) trust you, so you should treat them as an adversary for your own good.
Remember: HR exists to protect the company, not the employees.
Where do you work that makes tolerating that level of idiotic behavior worth it? It the job super interesting or the pay above market rate? If not, there are much greener pastures my friend.
Companies that treat their employees like morons eventually push out everyone who is not one.
Does your company tell its' employees about their Orwellian policies upfront when hiring, or is it a public secret?
How do you deal with encrypted traffic, e.g. https?
Some companies simply filter web traffic via corporate white/blacklists, maybe you have some insights why an illusion of freedom has been chosen in your company?
I think only very few people browse the web from an employer who has rules about what types of pages can be fetched. As a web developer, if I can make a faster experience for 99% of my population at the cost of potentially annoying the HR department of some tiny fraction of them, I'm going to do it. And I won't feel bad about making it slightly more difficult for my site's visitors' management to effectively surveil their browsing habits -- not my problem!
P.S: Using HN at my office. Was looking at career pages of other software companies a while ago.
But quite an interesting little thing, especially useful for older websites to bring some life into them. The effect is very noticeable.
It’s a different product. The initially planned name for it was “InstantClick Lite”.
If you send a requested page, that's obviously fine. Normal use of websites is expected, but unilaterally instructing your site visitor's browser to download further unrequested content that's not part of a requested resource ...?
I'd be a little wary of using a script from an unknown person without being able to look at the code - I'd rather see this open source before using. Especially being free and MIT licensed, I don't see why it wouldn't be open.
The code is not obfuscated or minified, very easy to read.
Both are under MTU size of TCP packet.
However, I think if browsers had this, but off by default until seeing tags to enable it along with any exclusions, that would be great.
User happens to brush over the logout button while using the site. On their next click, they're logged out. Weird. Guess I'll just log in again. Doesn't happen again for some time, but then it does. Weird, didn't that happen the other week? What's wrong with my browser? Oh cool, switching browsers fixed it. You're having that issue, too? Don't worry, I figured it out. Just switch browsers.
It's like pointing to a list of best practices and saying "everyone surely follows these."
For example, someone changed their signature to `[img]/logout.php[/img]` on a forum I posted on as a kid and caused chaos. The mods couldn't remove it because, on the page that lets you modify a user's signature, it shows a signature preview. Good times.
EDIT: For completeness, I have to add, that I am also part of the group of people who have violated that concept. Maybe neither frequently nor recently, but I did it too :-/
It's nothing to do with REST. It's part of the HTTP spec and has always been, that "GET and HEAD methods should never have the significance of taking an action other than retrieval".
> It's like pointing to a list of best practices and saying "everyone surely follows these."
It’s not a ‘best practice’ it’s literally the spec for the web.
Browsers take a more practical approach than "well, it's in the spec, they should know better" which is apparently what you're suggesting.
It's the same reason browsers will do their best to render completely ridiculous (much less spec-complaint) HTML.
They're already broken, exposing that is a good thing.
If you violate the standards, your website doesn't work. Who knew?
> People who read articles by highlighting the text with the mouse would also probably hover over all of the links and would end up wasting bandwidth for no reason.
"For no reason" is obviously wrong, making the web snappier is a reason.
Maybe browsers should only prefetch links on bloated websites since their owners clearly don't mind wasting bandwidth.
By the way, if your standard contradicts a popular methodology, it's probably a bad standard.
You can't assume a methodology is good just because it's popular. That's how you get cargo cults.
But if a methodology violates the standards, it's almost certainly bad.
Standalone program on Windows XP (or maybe Windows 98?). It had it's own window where you could see which pages it was loading.
Does anyone know the name?
Separately to that, the communication between the user and your server when downloading your script is secured by your SSL. This can be secure even if example.com is not, so it should only be secure.
When you forgo SSL on your own server someone can also intercept your script in exactly the same way, they don't need to hack the website embedding your script. Now they are your consequences, your fault there's no SSL, and your problem may be affecting everyone who embedded your script insecurely.
[1] https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
You should NEVER load javascript over https on a page that was served over http.
It gives a false sense of security that doesn't exist.
Because the source page was served over http, the source page can be modified by an attacker, making the script be loaded under ssl makes you think its protected. But its not since an attacker could just modify the script tag to remove the https bit in transit then modify the script thats now being loaded over http.
In sort, forcing the script to use https gives you no gain when the page that includes it is served over http and it tricks you into thinking that javascript asset is secure.
Does it help much? No. Should you use it? Yes.
Who still has pages served over anything hit HTTPS?
Someone who wants to explore our emerging financial liability as we grow increasingly compelled by law to actively protect data in transit (that's SSL), at rest, and in distribution. Penalties for violations can be huge, as much as 20
million euros or 4% of the company's annual turnover.
https://blog.quttera.com/post/gdpr-and-website-security/More HTTPS == Better. If you can load this over HTTPS you should, no matter what circumstance. The browsers iconography will handle notifying people when the context is secure or not.
I can see children getting punkd by drive-by prefetch and reporting to teaching staff that X visited a neo-nazi site or, Y downloaded porn during class, etc..
"Prefetch did it" is probably not going to be apparent to most, and is going to sound like a weaksauce excuse.
Every instance of web filtering I've been subject to in my life just blocks the bad page and the admins expect people to have a few bad requests just by accident or whatever. You'd have to be constantly hitting the filter for it to actually become a real issue.
It's curious you mention checking where links link to, because I think that's also another user expectation failure. The url that appears in the status bar below (or what was once a status bar) is not necessarily the link's true destination. You can go to any google search results page, hover over the links in the results and compare with the href attributes in the <a> tags. They're different. It looks like you'd be going directly to the page that's on the URL, but you're actually first going to google and google redirects you to the URL you saw.
It used to be that checking the url in the status bar allowed you to make sure the link really would take you to where the text made you think it would take you, but that's no longer the case. It seems one can easily make a link that seems like it would take you to your bank and then take you to a phished page.
I would bet 99%+ of web users do not have a sufficiently detailed mental model of web pages that this is something they've decided one way or the other.
Google Analytics et al, allow custom events which are used to record mouse overs, clicks, et cetera on a majority of websites. I always just assume everything I do, down to page scrolls and mouse movements, is recorded.
The validity of that expectation died with the advent of web analytics probably two decades ago now. The wholly general solution is to disable javascript, possibly re-enabling it on websites you trust. Sites that break without js are oftentimes not worth browsing anyway.
It took a lot of tweaking to give it the right feel. imho it shouldn't start fetching to fast in case the mouse is only moved over the link. Loading to many assets at the same time is also bad. Some should preload with a delay and hovering over a different link should discontinue preloading the previous assets.
Perhaps there is room for a paid version that crawls the linked pages a few times and preloads static assets. Who knows, perhaps you could load css, js and json as well as images and icons.
Or (to make it truly magical) make the amount of preloading depend on how slow a round trip resolves. If loading the favicon.ico (from the users location) takes 2+ seconds the html probably wont arrive any time soon.
Fun stuff, keep up the good work.
[1]- http://web.archive.org/web/20130329183223/http://widgets.ope...
My guess is what would happen when "pre-fetching" a react-router link, is that it would prefetch the JS bundle all over again for no gain.
A problem with the loading instructions is that it reveals, to an unrelated site, every single time any user loads the site that is doing the preloading. That is terrible for privacy. Yes, that's also true for Google Analytics and the way many people load fonts, but it's also unnecessary. I'd copy this into my server site, to better provide privacy for my users. Thankfully, since this is open source software, that is easily done. Bravo!
<script type="text/javascript">
if (screen.width > 768) {
let script = document.createElement('script');
script.src = '//instant.page/1.0.0';
script.type = 'module';
script.integrity = 'sha384-6w2SekMzCkuMQ9sEbq0cLviD/yR2HfA/+ekmKiBnFlsoSvb/VmQFSi/umVShadQI';
document.write(script.outerHTML);
}
</script>> On mobile, a user starts touching their display before releasing it, leaving on average 90 ms to preload the page.
Just found that you have the source available too :) So overall, this is pretty cool.
1) What, if anything, is the downside here?
2) Is (Google) analytics effected by the prefetch? That is, does that get counted as a page visit if the link that triggers this prefetch is not actually clicked?
Tia
Client-side analytics like GA aren’t affected.
How? It would load the HTML just as often and not download it a second time as that would invalidate the usage
> this makes for additional load on your server.
True for users who hover over links they decide not to visit
Or am I misunderstanding something here?
I'm guessing you didn't read the linked article? It preloads after 65ms on hover, at which point it estimates a 50% chance that the user will click. Hence "loaded twice as much".
> True for users who hover over links they decide not to visit
Yes, that's the point.
Does this also work for all the outgoing links as well? I don't want to improve other sites rendering time at the expense of my own.
Very cool regardless.
Also, there’s usually no incentive to improve other’s sites pages load.
Off the top of my head, a good way seems to be to write better sites that don't include 10mb of javascript libraries.
dieulot, is there a small bug with the allowQueryString check?
const allowQueryString = 'instantAllowQueryString' in document.body.dataset
I think should be: const allowQueryString = 'instantallowquerystring' in document.body.dataset
If I have: <body data-instantAllowQueryString="foo">
then 'instantAllowQueryString' in document.body.dataset === false
and 'instantallowquerystring' in document.body.dataset === true
because html data attributes get converted to lowercase by the browser (I think).Anyone else having this issue?
* Well, there's this, but I assume it's not dependent on google... `Loading failed for the <script> with source “hxxps://www.googletagmanager.com/gtag/js?id=UA-134140884-1"`
For example, in Drupal every path (whether or not it causes state change) has 2 forms: "/?q=path/to/page" (when you don't have access to .htaccess or .conf) and "/path/to/page" (when you do, and you enable clean URLs).
But it is a clever idea. Applied carefully, it could give the impression of a speedier site. Of course, I see no reason to need a 3rd party for this... updating event handlers to operate this way shouldn't be outside the abilities of most web developers.
However, I find it very good the posted snippets include SRI, which sadly to this day almost every script CDN omits. The code is also small enough to just include it in projects, which avoids the external request entirely.
I wonder if you could do something similar simply by looking at the viewport and loading all the links within the currently visible part of the page. Might be overkill though and end up wasting the user's data.
> On mobile, a user starts touching their display before releasing it, leaving on average 90 ms to preload the page.
That is, if you wanted to prove that "X is better than Y by 1%" you would need a sample approaching 10,000 attempted conversions to have a hope of having a good enough sample.
But there's no guarantee that it will apply for all pages.
See: https://developers.google.com/web/fundamentals/performance/w...
There are other libraries that do similar things, such as InstantClick: http://instantclick.io/
I'd surmise the author is benevolent, but if this were to be turned into a business, some kind of data play seems like the trivial next step.
Enumerate through a list of pages on your site and use something like Puppeteer to simulate hovering over links on each page.
But you make an important point.
Not if you use the `integrity` attribute like the page recommends. If the file hash doesn't match the one given in the tag, the browser won't run the script.
https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
I've added it to my blog and a Django side project. Really speeds up page loads. Just need to add `data-no-instant` attribute to the logout link.
The talk (https://www.youtube.com/watch?v=V8oTJ8OZ5S0) was one of the most watchable performance optimization talks I've seen.
TLDR - they used a combination of the link prefetch technique, which works for HTML but is not fully supported by all browsers, as well as XHR prefetching, which will work for prefetching Javascript and CSS.
Testing on my dev box now. I’ll be rolling this out to a subset of users next week. Cool stuff!
You could write something to parse the page and download images, but I don't recommend it. You risk a significant performance hit for the client, and a lot of wasted bandwidth for both client and server.
click on subforums and threads.
> "Did you even read the article? It mentions that" can be shortened to "The article mentions that."
I thought CloudFlare was charging workers based on number of requests.
(My apologies if this question has already been raised)
It prefetches on touch event while a "click" is normally triggered by a touch release.
So now I can't even touch my screen without being taken somewhere else, because I don't know where all the active areas are.