Python 3.11 in the Web Browser
2022.pycon.de
2022.pycon.de
If users might define a list of sites trusted for cross-site caching (like fonts.google.com and other known library CDNs), this could help cache quite some common resources without downsides to user's privacy.
fonts.google.com is, after google analytics, the most effective spyware around today.
> To delete the HTTP Cache, you just have to either issue a POST request to the resource, or use the fetch API with cache: "reload" in a way that returns an error on the server (eg, by setting an overlong HTTP referrer), which will lead to the browser not caching the response, and invalidating the previous cached response.
https://sirdarckcat.blogspot.com/2019/03/http-cache-cross-si...
* Any random delay of up to less than 0.9 seconds is completely pointless since we would know that any resource that loads in 0.5 seconds has to be cached - the 0.4 seconds spent waiting is useless.
* A random delay of up to a second creates ambiguous cases: is 1.1 a cached + maximally delayed response or an uncached but only slightly delayed response? But, with a random delay, we're still going to have pretty different peaks in the expected latency graphs for cached and uncached resources - uncached ones would peak around 0.5 seconds and cached ones would peak around 1.5 seconds.
* A random delay much more than a second would start causing the peaks on our expected latency graphs to be closer together - at a 10 second delay, we'll have a peak in the graph around 5 seconds for cached resources and 6 seconds for cached ones. I'm a bit fuzzy on how many resources we'd have to load to get a statistically significant finger print. I suspect its not all that many - but I could be wrong. However, we're also talking about a pretty significant delay by this point.
We also have the issue of how to handle AJAX requests. You only get one shot to measure the initial load. But, after that you can make an AJAX request to re-load that same resource with caching disabled and measure that. I'm not a statistician, but, I suspect that that information will be pretty helpful in figuring out what was cached vs not cached. Of course, we could add in some random delays here too - but since we can measure this an unlimited number of times, I suspect this would be even easier to defeat.
So, my suspicion is, that these type of random delays are possible to defeat if you have a better grasp of statistics than I do.
But, we could also assume that that is not the case and that they can't be defeated - what does that get us? We've had to add all these random delays in to block the side channel which is going to hurt latency - the very thing we're trying to improve with caching. I also suspect there is a ton of complexity to consider if you want to avoid having all of the same random delays on repeated visits.
You can't reduce first-visit-to-site-B latency below what it would be if you hadn't visited site A and also not reduce first-visit-to-site-B latency below what it would be if you hadn't visited site A; no possible delay policy will help with that.
> they can control the latency of non-cached requests.
This is more of a problem, but it doesn't need latency as such - site B could compare request traffic to b.example versus cdn.example to see whether a request was skipped due to already being cached.
(To be clear, I'm not sure you can actually do cross-origin caching securely for web traffic, for the above reason; I was addressing the narrow question of how to pick a delay - namely, don't delay cached and uncached responses by the same amount relative to their naive latency, because the whole point is that their naive latencies are different and we're trying to make that not true.)
- if the resource isn't in cache, download it and note the time it took.
- if the current site has requested the resource before, just return it instantly.
- otherwise, wait the exact same time it took originally.
With this the first request a site makes for a resource will always take the same time and you have no way of knowing if that was the first time it was downloaded or if it's served from cache. It's obviously not quite that simple, you'll need to factor in which connection is in use etc, but it should be possible to keep cross site caching.
You can make up theoretical models for how many popular libraries there are, or how many sites, and say this should or shouldn't be possible. I haven't seen any such models, but advertisers were definitely using this technique in the wild so the models that say it's impossible are all wrong.
Adding noise to a small sample of requests doesn't buy you that much entropy - the signal is a little noisy anyway.
Think about the 10,000s of JavaScript libraries out there, and the 100s of versions of each.
A good intuition is that although there's low certainty which site you have visited, there's high certainty which sites you HAVE NOT visited
I think is easier to see how you can fingerprint someone based on the set of sites they HAVE NOT visited
Edit: I'm not so sure anymore, I think that requires to test looots of libraries, it's not practical unless you are a really nasty ad company that tests hundreds of libraries in the background really sucking up your bandwidth
Imagine ad companies choose libraries that are roughly used by 50% of users, if they test you with 10 of those, they learn ~10bits of entropy to classify you, i.e. which of the 1024 classifications you belong
Millions of websites use vue (as an example) from the same CDN. There's no way to know which site it was cached from.
What a terrible thing they've done, making the internet slower for everyone on earth for no good reason.
Websites don't just "use Vue". They use some particular version of Vue along with particular versions of other libraries. My understanding is that once you put all of that together, this can create some significant privacy leaks.
Users need to trust the browser's producer anyway.
Kind of like how macOS (used to) ship with python preinstalled?
"Oh they fixed that in Typescript."
But they could have fixed it with practically any other language as well.
Like he makes fun of the fact that Array(16).toString() prints 15 commas. In the context of the talk it's funny but in reality, what would you expect. You made a array with 16 empty elements. Array.toString() calls toString() on each element and separates them by commas. Why is that unexpected?
He then shows Array(16).join("wat") which is the same as the previous except JS uses "wat" between elements instead of ","
"wat" + 1 is string + number coerced to string so string + string. string + is defined as concatenation
"wat" - 1. There is no override for - for a string so numeric minus tries to add a string to a number and returns NaN. Ok, why is that unexpected? When will this bite you. You shouldn't be adding numbers to string or strings to numbers. I've been programming JS for ~20 years, I don't remember running into any of these issues.
I've also never tries to add to arrays, add an object and an array, add an array and an object, nor add 2 objects. It's funny that the language does something but so what.
You wanna talk about a language that sucks try bash. Meanwhile I've had no problems shipping 100s of projects in JS. (also, C, C++, C#, perl, python, assembly, others)
Python is already >31 years old.. Python is even older than Java. It is so slow - can barely crawl compared to other languages. Maybe we should retire grand-daddy Python ?
JS was considerably slow until Google decided to spend resources to make it faster.
Not that Python can't be made faster (though many architectural decisions of the language resist it), but it hasn't been so far, so there's little incentive to include it to a browser. It gives too little new capabilities, on top of JS.
OTOH, say, WASM gave many new capabilities, and has been included.
Server side Python is often deployed in a venv which has the desired python version and all the dependencies copied with the code.
What am I missing?
You can just drag and drop a file in, it’ll be imported in as the ‘data’ variable, and then you can run Python code/do matplotlib visualisations without installing anything.
https://blog.jupyter.org/jupyterlite-jupyter-%EF%B8%8F-webas...
Pyground is specifically written for a use case I wanted to optimise: to get data from a file on your local machine into a structured Python variable fast. I think with Jupyterlite you’d have to upload the file and then write your own code to read it/parse timestamps, which is just boilerplate. So if you're trying to do something like that and don't need anything else that JupyterLite offers then pyground might get you there faster. JupyterLite is way more flexible though.
Also you can use pyground to load the data, then do `import pickle; pickle.dumps(data)` in pyground, copy the output and then do `import pickle; data = pickle.loads(<copied output>)` in JupyterLite and you'll have the loaded data variable way faster than writing that code yourself and all the flexibility of JupyterLite :)
[1]: https://skulpt.org/
[1]: https://github.com/blockpy-edu/skulpt/blob/master/src/lib/ma...
it would be really great if we could have full python/pip support in the browser but with some vetting done (just realized we don't have such thing in pip, other than relying on pip lockfiles and pipenv)
There are some limitations around what Python code you can run, there’s some details in the readme about those.
What I really wish for is for ~all Python packages to work in the browser without manual porting of the underlying C/Rust/etc. being needed, since a lot of the interesting and useful libraries aren't pure Python, and manual porting is non-trivial.
I'm not sure what the best route to that future is, but I'm guessing it'd probably help if Python had a wasm runtime in its standard library[1], since then authors of libraries that use C/Rust/etc. might make cross-platform builds (perhaps by default).
Regarding this Pycon speech, it seems that it's related to the following entry in the 3.11 changelog[2], which the speaker was heavily involved with:
> CPython now has experimental support for cross compiling to WebAssembly platform wasm32-emscripten. The effort is inspired by previous work like Pyodide. (Contributed by Christian Heimes and Ethan Smith in bpo-40280[3])
But maybe Christian has more to reveal here? In any case, I'm hugely appreciative of all the work that is being done to bring Python to the browser!
[0] https://github.com/pyodide/pyodide
[1] https://discuss.python.org/t/add-a-webassembly-wasm-runtime/...
That way you'd be able to use other languages in the browser without needing users to download 20MB of a WASM-compiled interpreter just to run 1KB of code.
Right now, WASM interpreters only really make sense for teaching and exposition with REPLs, where the user probably won't mind large downloads in order to do something out of the ordinary. Shipping interpreters would instead make that ordinary.
Personally, I'm pretty happy that we have WASM at all, and I think there's a lot of work we can do (over the next years) to make interpreters that work well in WASM.
My install of Google Chrome is around 85 MB. 20 MB is a lot more than nothing.
Edge: almost 500 MB
EdgeCore: almost 400 MB
EdgeUpdate: about 20 MB
In the grand scheme of things, 20 MB is indeed nothing, because many browsers out there (that cannot be uninstalled without crippling the OS in some regards) are already pretty bloated, use bunches of plugins anyways and just generally have untold amounts of cruft in a variety of other software (e.g. just compare MS Office vs LibreOffice and look at how much space professional software like Photoshop or Blender or whatever takes up).What i'd like:
- to optionally be able to maximize the browser size install to minimize the amount of data that would have to be fetched over the network (e.g. one bundle for Python, one for .NET/Blazor, one for Rust, one for Go etc., based on what you need, maybe just all of them), the same way that plugins work
- to have these bundles support being toggled (or even downloaded, if allowed) on a case by case basis, as they become necessary (e.g. enabled in daily driver device, disabled and not installed in a Firefox/Chrome/... install inside of a Docker container for testing)
- somehow have the industry force everyone to slow down - e.g. you'd have new updates for all of these come out perhaps once a month instead of every other day, as you already do with front end plugins and needless updates of JS bundles
- so, since CDNs were crippled and browser caching is impossible for common resources across different sites, re-introduce that mechanism in some capacity with these bundles that'd be installed locally
- enjoy the ensuing hell that'd be like the Java Applet idea which was brilliant for rich content but had untold flaws in regards to sandboxing and other things (similarly to how Flash was later also killed off, really torn about that one)
So obviously we cannot win and will never have that.Alternatively:
- build more static sites
- realize that you can't implement everything you need without JS
- reinvent your own minimalist framework/library, badly; though hopefully use something like Svelte or Alpine.js for cases like that
What we'll realistically have instead: - unique bundles of resources for every site, no caching across sites, no proper ways to utilize CDNs in a cross-domain context due to fears of being spied on
- fast development velocity with lots of updates, continuation of needing to download hundreds of KB or even multiple MB of JS assets to keep browsing the same content in almost the same way
- the problem will only be made worse by developers pursuing larger WASM bundles, like Blazor in .NET, with few advantages for developers but many disadvantages for everyone else
- this problem generally will not be regarded as serious, because most don't care about how much bandwidth they waste, a la Wirth's lawThere's no reason browsers couldn't have a persistent cache for this. Think of it more as an integrated dependency manager and VM that just happens to use web infrastructure. No one would blink at downloading 20MB of dependencies anywhere else, after all.
If you're downloading it on-demand and caching it, why would it need to be integrated in the browser at all? Why not just stick it in a CDN and treat it like a regular file, the way it works right now?
The browser handling this as a special feature by default avoids cache isolation issues. It also makes it trivial to avoid leaning on even more third parties (CDNs, package managers) to run your code: as long as your browser is supported, code won't stop working in it.
That does open an argument for having a kind of global interpreter cache, though it could still be used for fingerprinting.
And yet whenever this comes up, someone always insists that you'd have to re-download the entire runtime with every request, as if caching wasn't a thing.
As far as shipping vs caching goes, I see no reason not to do both. Maybe ship with the latest version of a few popular languages including javascript already pre-cached and allow for downloading others as required. I don't know what would be more optimal. Maybe when you download a browser, you can select language support options.
My point is, this is just an implementation detail, it doesn't have to be awkward or inefficient.
in which version?
import numpy as np
import time
n = time.time()*1000
r = np.random.rand(10000000)**2
print('done in', time.time()*1000-n, 'ms')
In the browser I was getting around 200ms and on my computer I was getting around 65ms, so around 3 times slower.Also, if I try to allocate an array of length 100,000,000 then I get a `MemoryError` in the browser.
Does anyone know any ways to get around these limitations?
I imagine I need a WASM runtime installed, and to somehow get my shell to recognize that WASM programs should be loaded with it (for lack of a hashbang line), but is that actually doable?
You do have to create a js host file, load in your webassembly and then run it with node.
wasm3, WebAssembly interpreter: https://github.com/wasm3/wasm3
… and many more
Really, without all the details/nuance, it sounds like we shot ourselves in the foot, went down a wrong path because of it, and now we're full-circle back to where we started. Except now the web is drastically different and the browser is now our one and true only terminal to the Holy Server that is Google et al.
/end of rant
Are you afraid of JavaScript in your browser? Perhaps you are, and that's fair, but WASM is no more dangerous than the JavaScript that everyone already runs.
This should make it possible to run python in wasm in python.
Well, there was that one web browser written in Python, with built-in Python scripting support: Grail.
I'm not saying you can't do async work in any language. Rather I'm saying that likely the patterns you're used to using in some language that's not JavaScript won't work in the browser and you'll have to massively change your style.
Just imagine how awesome would it be to use the full power of Python to build a react like library.
It is a source of great disappointment to me that WASM cannot directly control the DOM.
But when you don't, a thin, glue layer of code following different paradigms than the other 90% of your project will crush your development agility anytime you have to touch that glue layer.
Right now you can define any scripts and run them, but behind the scene, the way the workers work is that they have a static list of dependencies they can handle and always fork a python process to run the code in that environement.
I was scratching my head about how to provide proper isolation and handling of custom dependencies for Python short of zipping the entire list of pip dependencies and unzipping it at runtime. For Typescript, this can be achieved easily using deno compile. Store the small output bundle and run it with the proper runtime restrictions. With the ability to do more or less the same Python, this is a huge game changer.
I would guess it's only for core python, without modules.
Brython was nice to use, although its dom syntax was a bit awkward.
This would mean to me they simply do not work or only to an extend. As you can reading the virtual filesystem seems not fully compatible.
Long ago we compiled Python 3.6 and published to WAPM: https://wapm.io/python/python (you can run it online there!)
I wonder if we could publish the new version also! I think we got the Python repl properly running so it would be interesting trying to have that working on the new version too
And type hints is something that JS lacks. How many times I've tried to copy paste typescript code to nodejs/browser console and fail miserably
This comes from Python dev with no idea how to make web app. Eg I’ve seen some package not maintained for years in PyPI while continue to work fine.
Also I think Python is strongly typed while js isn’t. I would think that’s a bigger problem?
Also "strong typing" is not well-defined across the literature as acknowledged by this one link from Cornell (which lists strong typing as having "the type of every variable and every expression is a syntactic property" and variables that are "used only in ways that respect its type"): https://www.cs.cornell.edu/courses/cs1130/2012sp/1130selfpac...
You also have lecturers at Carnegie Mellon teaching that strong typing means "Types must by explicitly converted" :https://www.cs.cmu.edu/~07131/f18/topics/extratations/langs....
Personally I think the historical definitions from Wikipedia are the most concise (it's paraphrased by me as saying "arguments to a function should have the same type as the parameters the function was defined with): https://en.wikipedia.org/wiki/Strong_and_weak_typing
In terms of 'strong typing' (with the definition of not being able to do any kinds of implicit type conversions), Python can be seen as weakly typed since there are forms of implicit type conversions like adding different numerical data types together. I would say it has less instances of type implicit type conversions than something like Javascript though, so if there was a metric of "strength" defined that's inversely correlated with the instances of implicit type conversions than I would say Python could be considered "stronger" (although a lot of Javascript quirkiness is removed through the introduction of static types).
All that being said, the argument against Python seems to a gripe against implicit and dynamic typing whereas Typescript has neither and instead brings explicit and static typing (the argument doesn't seem web specific).
> I worry that we’ll now see more broken web apps, given the lack of type safety in Python.
I just don't see why "broken web apps" would be related to "type safety in Python", and how JS/TS solved the issue.
Put it in other way, we could forget anything about types and not even having unit tests, and if the first time you deploy it, it works, with sensible deployment you should expect them to work "perpetually". That's why I mentioned that some unmaintained package in PyPI continues to work for years, because Python 3.x doesn't break backward compatibility. And if they should worry about compatibility with any dependencies, "pinning" to fixed versions or minor versions is better practice. (I know it doesn't solve all problems, and ideally people should keep maintaining them. But again the statement is about "broken web apps" and "type safety".)
P.S. Python as strongly-typed language is about say `1 + "2"` situation, where JS seems to allow but not so in Python (although Python's object/data model can handle them via `__add__`). Rebinding a name to another variable is ok in Python, but is a separate feature/characteristics.
Also, while I don't know TS, from other comments here it seems that type hints can be more powerful in some cases, if not strictly more powerful. I'm a believer in type hints when using Python and I find it very useful (and recent Python versions are making them more and more powerful.) It is more flexible than static typing but is still useful in performing static analysis, which has detected bugs in my program even when it passes unit tests. (Often it is related to sloppiness rather than "real" bugs though. Having good types make reasoning about the program more easily.)
Lastly, runtime type check is a thing in Python too. E.g. I think it is very common for the `__init__` to perform runtime type checking. And then there's some packages allowing you to define a schema where the library would perform runtime type check automatically given the schema.