Intel Bringing SIMD to JavaScript
01.org
01.org
I'm trying to figure out what happens when you port this to ARM NEON, and how you catch it with architectures that don't support NEON (they often lack them in Marvell and Allwinner).
CPUs that lack SIMD units can support the functionality (though not the performance of course), and there's even a polyfill library that can lower this API into scalar operations for SIMD-less browsers too.
One thing to keep in mind is that most programmers probably won't want to use this feature directly; it'll be used in libraries that expose higher-level APIs. It's still true that every feature we add increases overall clutter, but SIMD seems sufficently useful and sufficiently self-contained that it's worth the tradeoff.
https://www.dartlang.org/articles/simd/
The primitives are pretty generic, just a few new vector types based on typed arrays. Operations on those types are supported on CPUs without a SIMD unit, they're just slower, but not any slower than coding with non-SIMD operations.
I'm probably nitpicking here, but:
* All Allwinner SoCs have NEON[0]
* Most current ARMv7 processors have NEON. Of the current ARM cores, only Cortex-A5 and Cortex-A9 don't have mandatory NEON support (it's optional). Cortex-A5 is intended for embedded applications. Of the existing Cortex-A9 processors, AFAIK the only somewhat popular one without NEON support is NVIDIA Tegra 2, which is retired. Out of the third party cores, all Qualcomm and Apple ones have NEON support.
MPEG1 in javascript is pretty fast and runs at native speeds on phones
https://github.com/phoboslab/jsmpeg
but MPEG4/H.264 is not
In any case, the two are not mutually exclusive, different people are working on each.
Do you know what the implementation status and intention to implement is, among browsers?
Samsung made it work in Tizen, but sadly, it is not very widespread OS. https://www.youtube.com/watch?v=TurCVdaUTMY
Though it might still be the smartest thing to do given the poor state and lack of recent progress of GPU OpenCL drivers. Even desktop apps that would like to use OpenCL are just barely limping along. See eg. Blender.
BTW. GPUs use SIMD too :)
In the right situation, for example in node.js, this discourages anyone from doing anything that would block the single thread handling requests for very long and makes it easier to handle the C10K problem in a much more interesting way (where traditional servers create a thread per request or queue them to handle one by one from a pool, where node does it all from a single thread). In other ways this can be bad it causes (in the case of Node developers) developers to not do a lot of the heavy lifting in Javascript itself and to hand off to native code that can do long hard things that may block on real threads (database queries, persistence backends, file IO).
The thing is that some code is easier to write with threads in mind. If I know there are 4 processes on a system, I can make 4 threads and use effectively use all for processors (potentially) at the same time. This is not really possible with Javascript. This makes heavy apps on a device like the Firefox phone (which is all about Javascript only) that want to could really benefit from using multiple cores impossible to take advantage of them (unless the user uses webworkers).
There are also just an epic amount of legacy with apps that are inherently built around the concepts provided by threads. Javascript is basically becoming the bytecode of the web. Things like ASM.js and emscripten are pushing it this way. But it's not really easy to port all code and fit on a model where threads are not allowed.
Conceptually web workers are more like processes than threads. It's a shitty work around in Javascript since we can't easily have threads without breaking the way Javascript works (which would break the web potentially). That means we have to write code into different processes where everything is isolated. On top of that, you can't share memory and you have to message and marshal data between webworkers. This is awkward and impossible to target from something like emscripten making it near impossible to port something expecting threads exist (theoretically if you have a C app that didn't use threads but different processes, you fit better but this almost never heard of). Most operating systems provide threads (pthreads most often except on Windows).
My solution is to make a scheduler that is used by my fork of emscpriten. All functions use return on call semantics and a "scheduler" in Javascript will decide which function should be called next from a list of virtual thread "stacks" (arrays) or if it should yield to the browser (setTimeout). The "scheduler" is single threaded itself so it's not really "threads" but does make it possible to emulate them for porting from native code.
It will, however, run (...) on the platforms that support SIMD. This includes both the client platforms (...) as well as servers that run JavaScript, for example through the Node.js V8 engine.
...and:
A major part of the SIMD.JS API implementation has already landed in Firefox Nightly and our full implementation of the SIMD API for Intel Architecture has been submitted to Chromium for review.
...and:
Google, Intel, and Mozilla are working on a TC39 ECMAScript proposal to include this JavaScript SIMD API in the future ES7 version of the JavaScript standard.
So, yes, there's definitely an intention there to put it into V8/Node.js/ES7 (guess, it will be in this exact order).
Not knowing much, I think it'll be interesting to see how general purpose applications would benefit from SIMD if it's accessed from a higher level. Does that mean that if I want to loop through 103 items and run arithmetic operations on them I'd have to do the following, (let's say I'm multiplying each item in items[] by 2, and items.length % 4 !== 0):
var batch = [],
results= [],
i = 0,
j = 0,
len = items.length,
a, b = SIMD.int32x4(2, 2, 2, 2),
c;
var mod = len % 4;
items.forEach(function (item) {
if (i < mod) {
results.push(item * 2);
i++;
} else if (j < 4) {
batch.push(item);
j++;
} else {
a = SIMD.float32x4(batch[0], batch[1], batch[2], batch[3]);
c = SIMD.float32x4.mul(a, b);
results.push(c.x);
results.push(c.y);
results.push(c.z);
results.push(c.w);
batch = [];
j = 0;
}
});
Of course this is the interpretation of a non-CS graduate who taught himself JS, some of the stuff mentioned at https://01.org/node/1495 seems a bit over my head. It'd be great if V8 would (unless it already does) transparently handle creating SIMD-optimised code where one is looping through an array or the like instead.(edited to fix code, hopefully)
You'd want to convert your non-SIMD data into a big stream of SIMD data up front, then do lots of operations on it, and then after that perhaps unpack it. Most SIMD scenarios just keep data in SIMD format indefinitely.
(Sometimes a compiler can use SIMD operations on arbitrary data by maintaining alignment requirements, etc. That sort of optimization might be possible for the JS runtime, but seems unlikely for anything other than typed arrays.)
For example, we have a media library that implements HTML5 canvas2d, WebGL, WebAudio, video/image/sound encode and decode and a bunch of other media related compute-heavy tasks. We've done this as a native module because it lets us control all those APIs using the same code we would use in the browser, but we run that code in Node.js as a separate worker process where we don't care about things like IO latency. It works great.
Edit: grammar
I'm developing an MMO server in Node JS. So this is welcome in for example Vector calculations.
The reason I choose JS is that I can write "dumbed down" code, that just anyone with JS experience can manage.
Hopefully, JS will be just as fast as optimized C++ in the near future ... First you will ignore it, then laughs at it ...
And when you run something heavy in Node JS the VM can use many threads. So I think it's weird to say it's single threaded. Aren't all programming languages single threaded then?
The author of FFTS, for example, chose a different strategy on ARM than x86_64: http://anthonix.com/ffts/preprints/tsp2013.pdf
He found himself writing the NEON code in assembly entirely by hand because vector intrinsics didn't even expose CPU features he wanted to use—even in C, where vector intrinsics are CPU-specific.
Having access to SIMD is definitely better than not having it, but it really should be paired with good optimized implementations of things like BLAS and FFT libraries.
There may be performance differences of course where AMDs chips process the same instructions differently internally, but I would expect similar optimisation differences between different families of Intel's chips too.
They wouldn't bother if it didn't work well enough on AMD CPUs because if it worked significantly badly or not at all then people other than Intel simply wouldn't use it. Of course they'll make no specific efforts to optimise it specifically for chip designs that aren't their own, but that is not the same thing.
Dart's SIMD stuff works just fine on AMD CPUs which were produced in the last 10 years (the K8 line started in late 2003).
Anyone take a guess what those might be (honest question)?
In the past SIMD has been the primary way to accelerate audio and graphics related compute tasks, but with WebGL and shaders, JS users already have a very powerful vector processing unit at their fingertips.
https://www.youtube.com/watch?v=CKh7UOELpPo
He's also the guy wrote the proposal for ECMAScript:
https://github.com/johnmccutchan/ecmascript_simd
Also, using SIMD is way easier than using shaders. You just write something like this:
double average (Float32x4List list) {
var n = list.length;
var sum = new Float32x4.zero();
for (int i = 0; i < n; i++) {
sum += list[i];
}
var total = sum.x + sum.y + sum.z + sum.w;
return total / (n * 4);
}
Instead of: double average (Float32List list) {
var n = list.length;
var sum = 0.0;
for (int i = 0; i < n; i++) {
sum += list[i];
}
return sum / n;
}Adding 2 numbers (x4) is just the simplest thing you can do.
There are a bunch of other operators and methods:
https://api.dartlang.org/apidocs/channels/stable/dartdoc-vie...
Good intro:
https://www.dartlang.org/articles/simd/
GMMs in speech recognition:
http://dartogreniyorum.blogspot.com/2013/05/performance-opti...
(Be sure to check the follow up article for the updated numbers.)
(That said, there are plenty of other areas for using SIMD in JS)
Will we have to explain to our grandchildren why they can only write code in some flavor of Javascript? Will it make sense, then, to wave our hands at ActiveX? Or do we have any better ideas for the future?
Mozilla's mission to protect the web from all languages but Javascript is locking us into a future where there will be no choice and nothing better than Javascript, because there will be neither an audience nor hardware support for anything but Javascript.
No, because while JavaScript is the "common" language of the web, its not the only language you can code in on the web -- there are plenty of implementations of other languages with it as a compilation target.
The value of having a guaranteed-to-be-everywhere target language for the web, even if it isn't the preferred development language of every developer, are fairly obvious.
> Mozilla's mission to protect the web from all languages but Javascript is locking us into a future where there will be no choice and nothing better than Javascript
As long as JavaScript keeps getting better -- and, particularly, as long as it is spurred on in that by efforts which propose alternative standard languages for the web with compelling stories so that JS has to keep moving forward in order to be acceptable as the universal, guaranteed target language -- that's fine. "Nothing better than JavaScript" isn't a real limitation if JavaScript is a moving target.
JS is only a "moving target" in the sense that stuff is being added to it. If you could make a perfect language by just adding things, then we'd be fine.
But the nature of the language itself is not going to change, because that would break backwards compatibility. The type system, prototype inheritance, `this`, type coercions, etc. There are plenty of undesirable things in JS which we're stuck with (unless we break compatibility, in which case it might as well be a different language).
Yes. If you prefer low-level programming, then write C++ and compile it to Javascript: http://leaningtech.com/duetto/
You can have your low-level cake and eat your high-level Javascript too.
They want to improve web application performance. JavaScript is the only language that works on the web. So they are bringing it to JavaScript. That doesn't mean they won't bring it to other languages in the future.
It doesn't make sense to improve HTML5 game performance by bringing SIMD to Lua or Python.
The only way around the problem you raise is to make something truly better, and convince people to use it instead.
PS: this is called "open web" - locking everything and everyone to one language to try and halt innovation.
Dart will do that. Python, Haskell, or whatever are free to do the same.
The difference though compared to plugins, is that
1. Plugins must be installed by the user.
2. Plugins are binary and usually do not exist for all OSes/browsers (e.g., no more Flash for linux, Silverlight never worked on linux, etc.)
3. Plugins are nonstandard, created and controlled by a single corporation, typically not open source, etc.
1. Browsers are installed by the user.
2. Browsers are binary too. It seems Mozilla prefers C++/Rust to build them rather than JavaScript.
3. New web tech is mostly developed by Mozilla, who is de-facto paid by Google.
2. The point is that while browsers are binary, they can then run a huge set of portable apps. As opposed to all those apps not being portable.
3. Everything Mozilla does is open source, and not just Mozilla but also most browsers today are open source - Chromium and WebKit in particular. This is a very open space.
But, people have meanwhile shown that cross-compiling to JS is practical, from things like CoffeeScript to C++. This is opening up the space to new languages, but it takes time and effort as well - again, the speed depends on how many people volunteer to help out.
this is just exposing a yet another platform API in the browser - once all browsers are updated in the field, you can reliably deploy your new code; or you can do it even before, but with a performance penalty; I'm not seeing anything wrong with that.
User: "Cool! What do I need to run it?"
Dev: "Just install {X}!"
User: "Wait, it's not available for {Y} yet."
Dev: "Shame... Let's just wait n years until it gets released."
X = Chrome or Unity plug-in
Y = e.g. Safari for iOS.
PS: the whole idea of API-mega-mutant browsers sounds like shipping Node.js/Angular with all the packages instead of just letting them be installed in a flexible way via NPM/Bower/whatever.
Whereas a third party plugin just used to run "a nice game" will never get the same adoption, and users wont be bothered.
So your example can be rendered as: Dev: "I made this nice game - try it!" User: "Cool! What do I need to run it?" Dev: "Just install a browser version released after 2015!" User: "Nice, my browser is already updated"
In the vast majority of the cases. And for those who haven't yet updated, either they're not your target audience (they have some ancient IE6 they use mandated by their company policy) or they can just go ahead an update their browser.
They should update their browsers regularly anyway, and getting the latest Chrome/FF/etc is not like being forced to installed some unknown plugin from some unknown developer just to run a single app.
Is it really just difficult to understand?
>PS: the whole idea of API-mega-mutant browsers sounds like shipping Node.js/Angular with all the packages instead of just letting them be installed in a flexible way via NPM/Bower/whatever.
Also called as "batteries included". But unlike in your example, those packages are one for each kind (e.g not 100 competing templating engines, or 100 MVC frameworks, like in NPM).
No, you don't need a specific browser. You can polyfill it. (The reference implementation is a polyfill.)
If you want the performance benefits, you'll need to update your browser. But that's true for JavaScript in general. IE8's engine is very slow, for example. Only the most recent browsers can offer the best performance.
The web page shows a Mandelbrot which is a terrible example because they are best done with shaders. In face there is no real use case for this. There will always be specific things that cannot be done and no language can address them all. How many apps need SIMD based physics? In fact I am not even sure what kind of physics needs SIMD.
So, before we had shaders, we used hand-written SIMD to perform those calculations as quickly as possible. Even with shaders available, there are broad categories of numerical work where you can't justify the expense of putting things out on a GPU but you still want them to be done quickly.
Good examples of this are basic vector arithmetic, which is useful for graphics and scene management and physics, or certain types of hashing, or anything else.
In some magical future land where sufficiently advanced compilers drop out of trees, maybe you have a point--but this is the web, son, and we're as quick and dirty as you can get.
Also, a machine can't just jumble the operations around because that would change the result.
>> 0.1 + 0.2 + 0.3
0.6000000000000001
>> 0.1 + (0.2 + 0.3)
0.6
So, this is apparently something you have to do yourself. A compiler won't know what kind of data will be fed to that function. It can't make an informed decision.> In fact I am not even sure what kind of physics needs SIMD.
It generally just means that you use some physics library which makes use of SIMD. Without having to do anything special, your game will run drastically better an use less energy to boot.
That's the primary use case; using libraries which use SIMD. Most people won't bother doing that by themself.
I am not saying this won't make things faster. There are many things you can do to make things faster but this really is such a niche feature (for web based apps).
That is a much harder problem, especially since you brought up performance. Getting the interaction between dynamically typed and statically typed code to work in a way that's both easy and natural to use and allows compilers to get significant benefits from the statically typed code is an unsolved research problem.
Yes, it is. However, it's relatively easy to implement and the performance improvements and battery savings are fairly huge. All things considered, it's a pretty good deal.