WebAssembly: A New Hope
pspdfkit.com
pspdfkit.com
Excellent !!
It does not do Postscript, the "PS" in PSPDFKit are the initials of the creator, Peter Steinberger (really nice guy, by the way!). Two/three letter prefixes are common in Objective-C as there are no namespaces (I use MPW).
Maybe steipete can chime in?
[0]https://pspdfkit.com/blog/2016/a-pragmatic-approach-to-cross...
I do see that pspdfkit folk are active in the pdfium bug tracker, for example:
https://bugs.chromium.org/p/pdfium/issues/detail?id=641
That sounds like a production problem, but it could just as well be them trying it out.
Also using the pdfium via Enscripten is not news (https://github.com/coolwanglu/PDFium.js is 3 years old).
On commercial, battle tested side PDFTron offers PDFNetJS & WebViewer for years https://blog.pdftron.com/2015/11/10/pdfnetjs-html5-pdf-viewe... (http://xodo.com/app is the demo referenced in the presentation).
You can see this demonstrated in their demo where highlighting the text garbles it -- because they render without using browser technologies, they miss out on browser implementations of things like highlighting. (Another example: their "get in touch" link in the demo is not a real link and can't be right-clicked, indexed by search engines, etc.)
In the specific context of a PDF that is perhaps inevitable (PDFs likely need special text layout) but in the larger context of random native apps, you only get a "free" port to the web by cutting out web features.
With that said, let me emphasize again that in the right domains this is all super awesome. It just doesn't solve the problem of making good webapps.
I 100% agree with you that our current implementation misses out on lots of browser-provided native behaviour. But this is not an inherent limitation of WebAssembly.
PSPDFKit started as an iOS-only framework. We then abstracted out a shared C++ layer which has tight hooks with the platform (Android or iOS).
PSPDFKit for Web is still a young product, one feature we have on our long-term roadmap is to have the browser handle rendering text natively. Our WebAssembly Core can "extract" WebFonts for the glyphs from the PDF and position them using HTML tags. This truly gives you the best of both worlds. PDF files need to be rendered absolutely accurately, so for now we went with rendering to an image.
If you're interested in our cross-platform strategy, you might be interested in https://pspdfkit.com/blog/2016/a-pragmatic-approach-to-cross...
i can't comment on the highlight-garble problem you mention, because i don't see it (firefox x64 for linux); however, the link "problem" has nothing to do with WASM. it's a PDF link, not an HTML link, and those have always been a completely separate concept. the same "problem" would exist however you're rendering the PDF, unless you convert it to HTML (which i don't think is the goal of pspdfkit).
wasm is only supposed to speed up heavy computations like those needed to decode and rasterize the PDF, modifying the page is still done through javascript.
http://kripken.github.io/emscripten-site/docs/porting/connec...
"compile our 500.000 LOC C++ core to WebAssembly and asm.js"
This is absolutely terrible for the web :-( We've escaped activex, flash, java applets, and now we're bringing lots of legacy c++ to the web in a way that conflicts with the underlying tenets?
I.e. we can maintain less code, keep it fast, maintain high quality and identical behaviour across platforms.
My hesitation isn't about your code. It's about what we as a community are doing to the web platform.
Not trivial stuff at all.
Also, why do you call this "legacy C++ code"?
Wasm is designed around bringing non-web-native languages to the web. A good project would have been to enable new web native languages to complement JavaScript. Instead, there's an emphasis particularly on bringing C++ to the web. I think this is counter to the philosophy of the platform and harmful in the long run.
Otherwise, then, we should tell people they "must" use C to write applications and that othrr languages should be there only to "complement C."
https://en.wikipedia.org/wiki/LLVM#Front_ends
C++ just happens to be a popular choice because there are many useful C++ libraries, and the compiler already produces very optimized code, so the performance benefit is most noticeable. In particular, I'm excited to get the Lua interpreter, written in C, on the web in a performant way.
WebAssembly makes it even more common to run untrusted, random code on your system. The next RowHammer.js will be a lot more powerful.
Can't wait for this!
I wish we'd all agreed to severely cripple Javascript[1] early on and instead improved built-in form elements and stuff like frames, but we didn't so here we are. The web's already mostly transitioned from being for people to being for machines and the delivery thereof, so webassembly's not going to ruin anything. It's already ruined.
[1] Say, no ability to trigger a broad set of events that would ordinarily be user-initated, no ability to submit forms itself, no XHR or similar. Probably some other stuff I'm not thinking of right now. So, almost but not quite worthless. That would have been necessary if we'd wanted to keep the people browsing in control of the Web. Instead we chased "oooh, shiny!" and took that control away, practically speaking.
Which keeps the delivered document a document, and not an application.
> At least some large percentage of the code for some of these "modern" web apps is on my computer: that means a lot; and anything we can do to discourage code running in the cloud with data stored in the cloud--which absolutely implies giving complex turing complete control over local behavior and expansive state--is a fundamental win for users of the World Wide Web.
I consider it much better that only explicitly-requested actions (clicking a "submit" button, or a link, for example[1]) trigger submission of data and execution on a server than to let javascript trigger the same. Control over behavior and data is better retained, especially for lay users, if it's impossible for an ordinary web page to send every mouse movement and keystroke to a server—something that's in fact done, all the time. Retaining that control is well worth pushing web "app" logic to the server.
Programs that need local execution would just have to be delivered as... a local program, not a web page. Losing that distinction cost the average user control and safety when they're in their browser, and fundamentally changed what the web is. Some of us much preferred what the web used to be, not because we hate progress, but because progress that destroys something that was itself valuable kinda sucks. "Web apps" should have had some other delivery channel, if they had to exist—and clearly there's a demand for better cross-platform GUI application delivery. Making browsers serve that purpose has made the web fat, opaque, and untrustworthy.
[1 EDIT] and in fact I'd rather restrict submission of new data, that the server didn't already have, to form elements the appearance of which cannot be modified at all, like a "submit" button. No JS submit-form-or-add-querystring-data-to-URL-on-link-click shenanigans. On my ideal web links would be safe to click, always, with only form submit buttons carrying any risk at all (and even that far less than now, if, say, JS isn't permitted to add to or modify form data in a POSTable form, which is another restriction JS would have on my ideal Web)
The underlying tenets are a bunch of technologies designed for rendering documents + one pretty horrible programming language. Which we're now building full-fledged apps and games on. I think that's way more of an abuse of the underlying tenets than webassembly is. I truly believe webassembly will set us free in many ways.
We want the best experience for users and devs that we can deliver.
Multi-platform makes sense client-side, where you have to deploy to Windows/Mac/Linux/FreeBSD/Android/iOS on x86 and ARM.
You know your server arch. You know your server OS. As long as you don't use C or C++ (or you're carful to be multiplatform), all switching is a recompile away.
If you are asking why I think it may happen rather than why it would be a good idea: imagine that you are a compiler writer for language X. You can target WebAssembly and run the language not only on the web, but also cross platform on servers and desktops and potentially mobile, and you even get a cross platform standard library including networking, UI (DOM/HTML), OpenGL, and more. This will save a tremendous amount of work compared to targeting LLVM separately and implementing your own cross platform libraries (which requires expertise on each platform that the compiler writer is unlikely to have). It's like targeting the JVM but lower level (compiler writers like this) and runs seamlessly on the web like JS.
Browser WebAssembly implementations will compile really fast because it has to be loaded and compiled for use on websites, making it suitable for a really fast compiler for interactive development and REPL, while it will be possible to compile to fast code via LLVM.
why do you say "as long as you don't use C or C++" ? these two languages are especially good when it comes to compile a single code base to multiple platforms. Got a 150kloc C++ app that "just compiles" to windows, mac, linux, android, iOS, under x86, x86_64, ARM, and it even worked on PPC macs last time I tried.
You know, people tried to do the same thing countless times already, there is Java and it tries to do exactly that. We all know where Java's promise "to run everywhere" went.
There are JVMs for things beginning from microcontrollers to s390, but find one that can flawlessly run code that was compiled for a VM that is even one major releases old.
Serious software in Java usually comes with something saying "version 2 of this software runs on JVM 'A' 1.4.13 with patchset 'B' 1.2, classpath 'C' 1.5, and exact VM setting from supplied jvm.conf".
"At this point, we want to issue a special thank you to the WebAssembly teams at Mozilla and Google, especially Alon Zakai, for being so helpful. We did run into a few edge cases but with their help we were able to still make it happen and even improved the Emscripten tool chain a bit along the way."
If some percentage of realistic projects run into edge cases that are incompatible with either the spec or between browsers, it violates the whole "write once run anywhere" goal.
Exactly.
Expectations that WebAssembly will be more cross-platform than any previous attempt at cross-platform are probably overly optimistic at this point.
Though I'm hopeful, as well.
WebAssembly is how these things should have been done from the start. Leave the object model, memory management, and libraries to the languages built on top, where they can be changed without affecting the backwards compatibility story of the VM itself.
When I got into computers it would have been ridiculous to imagine carrying around a computer running a different OS in a virtual machine, within which we were running a virtual network of small virtual machines existed to run programs written in a scripting language that itself runs in a VM!
Sounds crazy, right? And yet it isn't rare to find a developer on a Mac laptop running Linux in VirtualBox, inside of which they run Docker containers to run applications built on node.js. This is not done as a joke, it is done because it is a useful development environment for an application that is supposed to eventually be deployed to a cloud.
Portable software has been a motivating dream of our industry since forever. It was a selling point of COBOL, Pascal, HTML, Java, Flash, JQuery, and so on and so forth. The idea of deploying a known client target to get rid of browser differences is inevitable. It is just a question of who will be the first to do one that is good and fast enough to get mindshare.
From: https://www.destroyallsoftware.com/talks/the-birth-and-death...
That maybe true of current implementations but the design docs[0] appear to suggest that you will be able to import platform features in WebAssembly.
[0] https://github.com/WebAssembly/design/blob/master/Portabilit...
Could you enlighten us about this? Is this something to do with the eternal C++ modules proposal that never seems to make it into a standard?
The thing that comes to my mind when I hear "versioned modules" and "C++" is the way you can use "versioned" namespace re-exports to avoid ABI breakage, like so:
// header for version 1 of library
namespace mylib {
namespace v1 {
struct Foo {
int num_apples;
};
}
using namespace v1;
}
// header for version 2 of library
namespace mylib {
namespace v1 {
class Foo {
int num_apples;
};
}
namespace v2 {
class Foo {
int num_apples;
int num_bananas;
};
}
using namespace v2;
}
But I don't think that's what you mean..NET Rocks! #1455 - WebAssembly and Blazor with Steve Sanderson https://www.youtube.com/watch?v=cCdF9-q4n5k
> Is this actually .NET in the browser?
> It's not the regular .NET Framework or .NET Core runtime. It's a third-party .NET runtime called DotNetAnywhere, which has been updated and extended in various ways to support being compiled to WebAssembly, to load and run .NET Core assemblies, and with some additional functionality such as basic reflection and so on.
> Can I build a real production app with this?
> No. It's incredibly incomplete. For example, many of the .NET APIs you would expect to be able to use are not implemented.
Me too! C# or Common Lisp or Python or whatever, choosing the best choice for the specific project and circumstances.
Try it out.
And what they showcase here is a PDF document renderer, not a PDF generator.
We currently deliver SDKs for all major platforms with Windows coming in Q4 and they all have the same high-fidelity renderer and deliver the same result.
With pdf.js, how the output renders also depends on your browser, and it just doesn't support some of the more complex graphics with knockout groups, patterns and gradients.
See Mozilla's presentation where they say that full support for these is not even planned: https://youtu.be/4yLzRoErOHw?t=20m27s
Our customers are in Enterprise B2B space. Some might still use IE11 (which we fully support). Some care that their PDF files do not leak and only want streamed access. Most care about annotation syncing. Some apply watermarks on-demand. We solve a different category of needs than the more consumer-focussed pdf.js.
The big question is who will make the first WASM stdlib iface so backends can implement it. Otherwise, you might leverage LLVM right now because you have libc.
Why so? Why can't you avoid using JavaScript?
Or is the way those are doing things mostly incompatible?
more powerful features, such as threads, are planned as well
What, no more async hell when coding for the browser?Edit: Re "Anything wasm can do, JavaScript can do." ASM.js was one thing and ok, wasm is the real concern. What about the "binary blobs" do you not understand. The usual downvoters who have vast interest in succeeding of this questionable tech.
Edit: As browser sandboxes get better, running "binary blobs" gets less and less risky. We're a long way from "curl file | bash".
I don't understand the assumption that minified and obfuscated javascript is easier to reverse-engineer than compiled wasm is. This is simply untrue. In both cases the user is, in general, running a closed source (in the GNU sense of "the preferred format for code modification") system that doesn't have their best interests in mind.
Adblockers are much the same war as antivirus was: you want to run untrusted code but still have some say over which computation it performs. Unlike with antivirus, sandboxing and isolation don't help. We're just going to be writing heuristics to strip ads forever (or until we standardize on a non-executable document format, e.g. html-without-scripts).
Take a real webapp. Say something like Google Docs. Can you read it and understand it? What about gmail? Or Facebook?
Practically, they're a binary blob running on your computer. Not only that, but they could be hiding malware, and even if you viewed source you wouldn't be able to tell it's there.
Now maybe if JS was meant to be slow, and if there was no such thing as AJAX, and there would be no way for JS to send/receive data from the server (say, by modifying a GIF url), then I could see a complaint that a browser shouldn't run arbitrary downloaded code on your computer without opt-in. But that ship sailed over a decade ago.
I think WASM can and will bring a lot of cool stuff to the web, just like JS did. What worries me is that we learned virtually nothing from the mistakes that we made with JS.
Just out of curiosity, what did you learn from JS's mistakes? What would you like to see happen in place of WASM?
Giving a lot of power to the server is a bad idea. Users should be able to audit, control and modify how the data is processed in a very granular way. Strong artificial limiters on how many resources something can use. Basically, it should be a fully tunable VM for each process. Readable and modifiable by the end-user, without a lot of hoops to jump through.
tl;dr: the more power that the end-user has over the code, the better.
But it seems unlikely things will go that way. People are motivated by what they can make webapps do, and will take the quickest way to get there. Doing it right, where the end user has the most control, is more work to design and implement, so people won't bother.
Unfortunately, doing the right thing isn't as profitable, so we've collectively chosen to fuck the users instead.