If users might define a list of sites trusted for cross-site caching (like fonts.google.com and other known library CDNs), this could help cache quite some common resources without downsides to user's privacy.
fonts.google.com is, after google analytics, the most effective spyware around today.
You can make up theoretical models for how many popular libraries there are, or how many sites, and say this should or shouldn't be possible. I haven't seen any such models, but advertisers were definitely using this technique in the wild so the models that say it's impossible are all wrong.
Adding noise to a small sample of requests doesn't buy you that much entropy - the signal is a little noisy anyway.
Think about the 10,000s of JavaScript libraries out there, and the 100s of versions of each.
A good intuition is that although there's low certainty which site you have visited, there's high certainty which sites you HAVE NOT visited
I think is easier to see how you can fingerprint someone based on the set of sites they HAVE NOT visited
Edit: I'm not so sure anymore, I think that requires to test looots of libraries, it's not practical unless you are a really nasty ad company that tests hundreds of libraries in the background really sucking up your bandwidth
Imagine ad companies choose libraries that are roughly used by 50% of users, if they test you with 10 of those, they learn ~10bits of entropy to classify you, i.e. which of the 1024 classifications you belong
* Any random delay of up to less than 0.9 seconds is completely pointless since we would know that any resource that loads in 0.5 seconds has to be cached - the 0.4 seconds spent waiting is useless.
* A random delay of up to a second creates ambiguous cases: is 1.1 a cached + maximally delayed response or an uncached but only slightly delayed response? But, with a random delay, we're still going to have pretty different peaks in the expected latency graphs for cached and uncached resources - uncached ones would peak around 0.5 seconds and cached ones would peak around 1.5 seconds.
* A random delay much more than a second would start causing the peaks on our expected latency graphs to be closer together - at a 10 second delay, we'll have a peak in the graph around 5 seconds for cached resources and 6 seconds for cached ones. I'm a bit fuzzy on how many resources we'd have to load to get a statistically significant finger print. I suspect its not all that many - but I could be wrong. However, we're also talking about a pretty significant delay by this point.
We also have the issue of how to handle AJAX requests. You only get one shot to measure the initial load. But, after that you can make an AJAX request to re-load that same resource with caching disabled and measure that. I'm not a statistician, but, I suspect that that information will be pretty helpful in figuring out what was cached vs not cached. Of course, we could add in some random delays here too - but since we can measure this an unlimited number of times, I suspect this would be even easier to defeat.
So, my suspicion is, that these type of random delays are possible to defeat if you have a better grasp of statistics than I do.
But, we could also assume that that is not the case and that they can't be defeated - what does that get us? We've had to add all these random delays in to block the side channel which is going to hurt latency - the very thing we're trying to improve with caching. I also suspect there is a ton of complexity to consider if you want to avoid having all of the same random delays on repeated visits.
You can't reduce first-visit-to-site-B latency below what it would be if you hadn't visited site A and also not reduce first-visit-to-site-B latency below what it would be if you hadn't visited site A; no possible delay policy will help with that.
> they can control the latency of non-cached requests.
This is more of a problem, but it doesn't need latency as such - site B could compare request traffic to b.example versus cdn.example to see whether a request was skipped due to already being cached.
(To be clear, I'm not sure you can actually do cross-origin caching securely for web traffic, for the above reason; I was addressing the narrow question of how to pick a delay - namely, don't delay cached and uncached responses by the same amount relative to their naive latency, because the whole point is that their naive latencies are different and we're trying to make that not true.)
> To delete the HTTP Cache, you just have to either issue a POST request to the resource, or use the fetch API with cache: "reload" in a way that returns an error on the server (eg, by setting an overlong HTTP referrer), which will lead to the browser not caching the response, and invalidating the previous cached response.
https://sirdarckcat.blogspot.com/2019/03/http-cache-cross-si...
- if the resource isn't in cache, download it and note the time it took.
- if the current site has requested the resource before, just return it instantly.
- otherwise, wait the exact same time it took originally.
With this the first request a site makes for a resource will always take the same time and you have no way of knowing if that was the first time it was downloaded or if it's served from cache. It's obviously not quite that simple, you'll need to factor in which connection is in use etc, but it should be possible to keep cross site caching.
Millions of websites use vue (as an example) from the same CDN. There's no way to know which site it was cached from.
What a terrible thing they've done, making the internet slower for everyone on earth for no good reason.
Websites don't just "use Vue". They use some particular version of Vue along with particular versions of other libraries. My understanding is that once you put all of that together, this can create some significant privacy leaks.
Users need to trust the browser's producer anyway.
Kind of like how macOS (used to) ship with python preinstalled?
JS was considerably slow until Google decided to spend resources to make it faster.
Not that Python can't be made faster (though many architectural decisions of the language resist it), but it hasn't been so far, so there's little incentive to include it to a browser. It gives too little new capabilities, on top of JS.
OTOH, say, WASM gave many new capabilities, and has been included.
"Oh they fixed that in Typescript."
But they could have fixed it with practically any other language as well.
Like he makes fun of the fact that Array(16).toString() prints 15 commas. In the context of the talk it's funny but in reality, what would you expect. You made a array with 16 empty elements. Array.toString() calls toString() on each element and separates them by commas. Why is that unexpected?
He then shows Array(16).join("wat") which is the same as the previous except JS uses "wat" between elements instead of ","
"wat" + 1 is string + number coerced to string so string + string. string + is defined as concatenation
"wat" - 1. There is no override for - for a string so numeric minus tries to add a string to a number and returns NaN. Ok, why is that unexpected? When will this bite you. You shouldn't be adding numbers to string or strings to numbers. I've been programming JS for ~20 years, I don't remember running into any of these issues.
I've also never tries to add to arrays, add an object and an array, add an array and an object, nor add 2 objects. It's funny that the language does something but so what.
You wanna talk about a language that sucks try bash. Meanwhile I've had no problems shipping 100s of projects in JS. (also, C, C++, C#, perl, python, assembly, others)
Python is already >31 years old.. Python is even older than Java. It is so slow - can barely crawl compared to other languages. Maybe we should retire grand-daddy Python ?
Server side Python is often deployed in a venv which has the desired python version and all the dependencies copied with the code.
What am I missing?