Introducing SIMD.js
hacks.mozilla.org
hacks.mozilla.org
I would suggest two alternate approaches:
1. First come with a standard SIMD specification across CPUs. In some ways like the GL specification (for GPU) or even the webrtc stuff (for p2p connections). Once you have done that, define a JS API.
2. Identify proper use cases. The example are quite bogus currently. They show a mandelbrot. Well, all this is done in the GPU. Image manipulation is done best with css level filters. Once we identify the use cases, we can them maybe add APIs to specifically address them if it's important for the web in general.
In it's current form, this is really a 'because we can' API. Where does this end? What if we just expose raw socket and networking API instead of a webrtc style API? What about pixel manipulation of images?
Also see: http://www.mail-archive.com/webkit-dev@lists.webkit.org/msg2.... The apple guys are saying they can do all this even without the API.
Are there cases where algorithms simply cannot be expressed in a way that automatic vectorization could transform them into SIMD instructions? Or is this just a way to avoid implementing automatic vectorization in JavaScript engines?
This applies regardless of how good the autovectorization really is: it's in a weird catch-22 kind of space where adding more and more features to your autovectorizer can actually reduce its perceived reliability, by making the answer to "will this vectorize?" harder and harder for a programmer to answer at a glance. <xmmintrin.h> has a lot of problems, but it's reliable, and at the end of the day that's what history has shown that game devs and video codec authors want.
Yes, but _will it blend?_
Nit: Most of the projects I'm familiar with (libav/ffmpeg, x264, etc) prefer to break out the SIMD into hand-written functions, instead of relying on intrinsics or even inline asm. This avoids problems with register allocation and code gen, consistency/portability between compilers, etc.
Otherwise, yes, autovectorization is hard, both for application developers and compiler writers. Application code needs to be structured in a very precise way, and the correctness of C -> SIMD transformations needs to be proven. Intrinsics and hand-written SIMD aren't going away.
Yup. Which just goes to show: reliability is king.
I'll admit some ignorance here, but it also seems to me that a JIT may also have some advantages WRT autovectorization as compared with a static compiler, since you can collect runtime information about aliasing and loop trip count before choosing to vectorize. But if the point is to make performance easier to reason about, why not start with the rest of the language before worrying about vectorization?
But there certainly have been such efforts! Standards bodies have added features like Typed Arrays, Math.fround, etc., and work is ongoing on Classes, Typed Objects, and Modules. All of those things make performance more predictable.
There are also better devtools all the time, which help you understand performance issues better.
And there is also asm.js which aims to make a certain type of JavaScript extremely predictable.
A final point - the unpredictability you mention is exactly why a SIMD API is needed. JavaScript is more unpredictable than C and C#, but even those have added SIMD APIs, because even in their predictable worlds, autovectorization wasn't good enough.
Prefer assembly. Intrinsics usually make a disaster of register allocation and you lose much of your performance to needless load/stores.
Lately I've been rather disappointed in how minimal the gains are in reducing register spills from intrinsics on modern CPUs, with their wide decode/issue, 16 registers, and dual load pipelines - by the time a loop is complex enough that a compiler spills, extra load/store uops are almost free from a micro benchmark perspective. The macro gains from smaller code and reduced cache usage are a bit bigger, but still depressingly minor for the effort expended.
But if you care about 32-bit x86 that's another story of course.
On some of these processors, 128 bytes of stack or so is not really "memory" (in the sense of being stored with memory), so spilling is not that bad.
100% true for C++ (though it would be more accurate to say "4 decades" if you want to count fortran autovectorization, which has been going on since the late 70's)
But, i'll point out, plenty of the time, they end up writing slower intrinsics than the compilers autovectorization did to the same code.
(Plenty of the time they don't, too).
Additionally, all of the problems you mentioned are due to specific issues in C/C++. In other languages, autovectorization is not just "relied upon", it's basically "part of the standard" (see, e.g., Fortran 95).
Given that all of the brittleness you talk about is precisely because of the lack of pointer safety, alignment issues, and all sorts of things that simply only exist in C/C++, where programmers have a lot of control, i'm not sure it makes sense to base your argument on the experience of a language that is very different from the one this API was designed for.
All that said, truthfully, IMHO, neither autovectorization, nor intrinsics at the level you are talking, make for a good programming model in most languages.
The intrinsics at this level don't get used effectively: Among other reasons, they codegen differently on different platforms that don't directly have the exact same simd semantics, which is "all of them" :P
I know you guys are trying to avoid this by limiting the ops available/etc. It is, IMHO, a losing game.
So you end up with the same problem: People write loops that are really bad on some platforms, and good on others.
Autovectorization knows what the target looks like, but doesn't trigger in some cases people want it to.
In the end, I think doing things like Halide is a lot more useful as a programming model than simd.js
simd.js is a usable implementation mechanism for some of those programming models, but i would not sell it as the programming model itself.
In fact, almost the exact set of intrinsics mentioned in simd.js were allowed for generic operations on vectors in GCC (you can create a vector 32x4 float in a platform independent way, do normal ops on it, and it will codegen down to lower level vector ops, without ever seeing xmmintrin). It was simultaneously not high level and not low level enough.
People resorted to the lower level platform specific intrinsics to get better performance, or wrote higher level libraries to get better abstract.
In any case, i'm sure it's faster than what you have now, and certainly an advance. I'd just be careful of thinking it's going to work all that well except for targeted use cases.
Why wouldn't we want the same in JS?
> 2. Identify proper use cases. The example are quite bogus currently. They show a mandelbrot.
That's just one tiny demo. See the github repo discussions for lots of debate on use cases,
https://github.com/johnmccutchan/ecmascript_simd/issues
This has been discussed very intensively, with input from people from Intel, Mozilla, and Google. Feel free to join in as well.
In general, while I sympathize to some extent with you and the Apple position here, you are fighting a strong trend in the industry. C has a SIMD API, C# has SIMD API, Dart has a SIMD API, etc. - all those were created for good reasons. Autovectorization would be great - hopefully Apple can prove it beats the API approach - but no compiler has proven it thus far, hence SIMD APIs in all those languages I mentioned.
They are good technical discussions. What I am looking for is use cases for people to use in websites.
[1]: http://en.wikipedia.org/wiki/Skeletal_animation
[2]: https://github.com/h4writer/arewefastyet/blob/master/benchma...
It's honestly baffling that we keep away from threads in javascript but do all these optimizations like SIMD which are very minor. I mean practically every CPU out there has multiple cores. Workers don't cut it because they require copying data.
Both of those would be harder, and less general. Game developers are asking for SIMD; they aren't asking for specialized APIs.
> It's honestly baffling that we keep away from threads in javascript but do all these optimizations like SIMD which are very minor. I mean practically every CPU out there has multiple cores. Workers don't cut it because they require copying data.
Threads aren't easy to "just add" to JS. Making a thread-safe GC perform as well as today's highly-optimized single-threaded GCs is hard. And not even counting the engineering effort required, none of the millions of lines of JavaScript out there is thread-safe. Of course we will need a way to do threads eventually, but it's much harder than SIMD.
That is entirely a problem of your own making. You decided to bet hard on single threaded dynamic Javascript being a suitable model for all end user software. It turns out it isn't, but it's too late now.
This kind of thing is exactly why Javascript isn't a good choice as a general purpose VM platform. Which you are presumably aware of because you aren't writing the next generation of Firefox in Javascript, you are creating Rust. But apparently what isn't good enough for Mozilla is good enough for everybody else…
Which brings us to the real problem with the web - everybody except the browser vendors is a second class citizen.
https://github.com/johnmccutchan/ecmascript_simd/issues/59
is interesting here, it mentions, among other things, the major SIMD-using code portion from an important real-world codebase (IMVU).
Might also be relevant discussion on the emscripten repo,
https://github.com/kripken/emscripten/issues?q=is%3Aopen+is%...
as some debate happened there while implementing code generation that emits SIMD.js.
And, we certainly do have real-world use cases in mind. See [0] for one example. We're also interested in using SIMD for video codec development which can't always be done on GPUs.
[0] http://blogs.unity3d.com/2014/10/07/benchmarking-unity-perfo...
And if, in the future, Apple can show that auto-vectorization in this domain is more successful than their attempts so far have shown it to be, then JS engines can always just go back to implementing SIMD.js via a simple polyfill which the engine can auto-vectorize, leaving very little baggage in the language or implementations.
SIMD is used a lot in all kinds of things dealing with video, so that's one area where it could find uses. I wouldn't mind being able to do (realtime) video processing in the browser. Some relevant reading that mentions SIMD.js briefly (though mostly that it wasn't really useful at all at the time of writing):
We got this years ago. You can do per-pixel stuff via the Canvas API.
Maybe it's not too late to take up painting.
More like "we don't have 64b ints, yet."
https://www.kickstarter.com/projects/gfw/espruino-javascript...
This is what programming is now.
I am so ungodly sick of Google pushing DART technologies everywhere. I really am starting to think their strategy is embrace, enhance, exterminate.
As pcwalton says, SIMD.js evens an advantage that was in favor of Dart, while at the same time making JavaScript a better compile target for languages like Dart and asm.js languages like C++. It's also a really clean API that fits nicely onto typed arrays. What's not to like?
It looks like Alon Zakai is working on something the "Emterpreter," which is a bytecode format that Emscripten could compile to to sacrifice some runtime performance for startup time. (JS parse time is quite a big deal for Emscripten applications)
I am hoping that this effort goes really really well. So well that browser engines start natively supporting this bytecode.
Then JS can actually die.
Also, am hoping 'JavaScript Shared Memory, Atomics, and Locks' gets accepted by TC-39. https://docs.google.com/document/d/1NDGA_gZJ7M7w1Bh8S0AoDyEq...
1. It is already standardized language and API. There exist a lot of code for WebCL.
2. WebCL engines can run on CPU, they can use all CPU cores, SIMD etc., while still being a part of web browser (no special drivers required). It will give us much better performance, than asm.js, SIMD.js, Google's Native Client or any other "unstandardizible" things.
That is why we have been stuck for 15 years without a lossy image format that supports transparency (meanwhile they put lots of effort into supporting an animated png format that was invented by Mozilla that nobody else in the world cares about or uses). Mozilla's NIH syndrome holds back the web.
As for image formats, there was a table going around Twitter that I can't find now showing implementation of new image formats by browser. Basically Google implemented WebP, Microsoft implemented JPEG XR, and Apple implemented JPEG2000 (IIRC). There is no consensus in that space.
There is no consensus because Mozilla have spent the last 15 years rejecting every proposed format. Mozilla (repeatedly) rejected JPEG2000 long before Chrome even existed, so you can hardly cite Chrome's lack of support as an excuse for your inaction.
Would really love to see the speedup on fixed tuple transforms and convolution algorithms.
When I saw JS speed wars originally, I started writing an adobe curve apply algorithm in JS, to apply .acv curves live using Canvas, but essentially hit CPU & shelved it.
https://github.com/t3rmin4t0r/io.dine/blob/master/lib/iodine...
http://notmysock.org/code/iodine/
I want to rewrite it using something like the newly introduced float32x4, assuming I can read out the images int o RGBA tuples.
Something like a blur would actually be possible once this is fast.
EDIT: That said, this is good news, in a sense.
(I know little about other instruction sets.)
Maybe that's a good thing, at least then the goal is that people who care about performance don't write JS. But I'd prefer to just have the compile target as numbers then since you can't read the result anyway.
For backward compatibility, just use and extend emscripten. As a bonus, this would allow the DOM interface to be reworked so we can get it right.
Another solution is to version your code <script type="llvm" src="foo.ll" version="3.6.1" />
I'm already of the belief that w3c needs to require versioning of js code instead of stuff like 'use strict' (if it's bad, just remove it in new versions and move on). For backwards compatibility, if there's no version, assume ECMAScript3.
SPIR 2.0 on the other hand already is based off LLVM IR 3.4
I like the idea of having a language specifically for compiling to but either way someone has to build the prototype implementation. This idea of building LLVM, JVM or CLR in to browsers might be good but code is what matters when it comes to proposing standards.
Also, quite a lot of people use floats for every data value.