HolyJit: A New Hope
blog.mozilla.org
blog.mozilla.org
This is a pretty well explored technique, it's rarely done these days because historically, people could not get good enough performance.
(I am not trying to knock them, just put it in context)
For more context on how you'd do something like this, read this:
https://en.wikipedia.org/wiki/Partial_evaluation#Futamura_pr... and http://blog.sigfpe.com/2009/05/three-projections-of-doctor-f...
http://tratt.net/laurie/blog/entries/fast_enough_vms_in_fast...
These issues, as well as developer velocity in translating VM features into optimized VM features, figure more prominently in our problem set than I would expect it historically did.
We sat down with nbp and went through his proposal in some detail yesterday. I'm reasonably sold on the theoretical soundness of the idea (actually I was excited about it from the first time he proposed it - hand optimizing every new JS feature is soul-sucking, and Graal already demonstrated the feasibility of variant of the concept).
As with any far-reaching idea, though, there are risks associated with it. And we definitely can't rewrite Spidermonkey from scratch.
Personally, I think prototyping the tech on top of a small toy language (objects, properties, proto chains, primitive types, functions), proving out the toy implementation, and then examining how we can incrementalize our transition to HolyJIT is the way to go.
I think a prerequisite is a good codegen backend that we can target. Cretonne is a good candidate, but we need features that aren't on its roadmap - primarily support for on-stack-invalidation.
There's a bit of a road ahead of us on this concept. I think we have a good rough idea of a viable path from here to there, but we are yet in early stages with this.
Since when? Since the 70's? Sure.
But since the 90's, not really in terms of techniques, only in terms of engineering and feasibility of advanced techniques. It's not that there is no research, mind you, but it's definitely more engineering than research. That also doesn't make it any less cool, exciting, etc.
"another difference between the days of old and new is that this is being targeted at runtime code generation, which naturally trades off throughput of generated code for speed-of-generation."
This actually was true then too, FWIW.
"These issues, as well as developer velocity in translating VM features into optimized VM features, figure more prominently in our problem set than I would expect it historically did."
While i'm not sure how much it matters, i guess i'd just point out these are not different concerns than history had
:)
I certainly hope you succeed, FWIW.
Ah, you were referring to a era before my time, and it seems I made some false assumptions about motivations back then. Thanks for the correction.
> But since the 90's, not really in terms of techniques, only in terms of engineering and feasibility of advanced techniques. It's not that there is no research, mind you, but it's definitely more engineering than research. That also doesn't make it any less cool, exciting, etc.
Are you referring to meta-compilation techniques here, or the techniques for runtime type modeling developed to drive type-specialization of dynamic code?
If you are referring to the latter, I agree completely. If the former, I'd argue that the runtime type modeling work brings something new to the table which changes the dynamic. But in general I agree with your point - the main difference between now and then is the sheer level of engineering effort by multiple parties, cross-pollination of ideas, and other prosaic matters.
In terms of research, my exposure has been to two main pedigrees of thought in runtime type modeling serving to drive type-specialization of dynamic code: the Self work by Ungar and friends, and Type-Inference by Hackett and Guo (both of whom I have the pleasure of working closely with).
> While i'm not sure how much it matters, i guess i'd just point out these are not different concerns than history had
It always helps to understand the motivations and efforts of what came before, so thanks for the clarification. Any more insight or information or references would be welcome.
> I certainly hope you succeed, FWIW.
This is nbp's baby, but yeah, I hope it succeeds as well. I will never get bored of working in this space :)
Graal has shown that you can use interpreter specialization to create a Javascript engine that's essentially as fast as V8. It relies on a lot of clever tricks to do that, like partial escape analysis. Graal has generated quite a few published research papers. It seems like both engineering and research, to me.
I tend to draw it at "research produces new things that were not previously known, engineering may produce new insights or improvements of things that were already known".
(but again, I admit this is not a very very bright line)
I consider graal to be good engineering. It is a new arrangement and engineering of existing techniques. That will in fact, often produce new papers.
For example, I built the first well-engineered value based partial redundancy elimination in GCC. before that, there were zero production implementations, and it was considered "too slow to be productionizable" until i took a whack at it. I helped with some papers on it. It's not research, just good engineering. The theory was known, etc. I just made it practical. That wasn't research.
Another example: LLVM now has the first shipping implementation ever of an efficient incremental dominator tree updating scheme (that i'm aware of. GCC has a scheme to do it for some things, but it's not efficient). Again, previously not efficient. Theory has been published. Again, making it work well is just good engineering.
Another example: LLVM's phi placement algorithm is a linear time algorithm based on sreedhar and gao's work. If you read further research papers, they actually pretty much crap on this algorithm as very inefficient.
It turns out they were just bad at implementing it effectively, and LLVM's version is way faster than anything else out there. Is it research because our results are orders of magnitude better than anything else out there? No. It may be cool, it may be amazing, etc, but it's still engineering.
Remember that conferences like PLDI and CGO accept papers not just on research, but on implementation engineering.
All that said, i also don't consider trying to differentiate heavily between research and engineering to be that horribly interesting (though i know some prize one or the other).
This operates at the MIR layer. You can sort of think of MIR as "core Rust", in that it's the final, desugared form of everything.
This is why the non-lexical lifetime stuff has taken a while; the precursor to that is "port the borrow checker to MIR". MIR/HIR are also fairly new; using MIR is only a year old.
Though I don't want anyone to get the impression that MIR is a source-compatible subset of Rust; it's a pretty different thing in its own right. (Worth clarifying because one could imagine a "fully-desugared" maximally-explicit subset of Rust, where e.g. all method calls are maximally disambiguated via UFCS, all types are explicitly annotated, no lifetimes are elided, all macros are expanded, etc.)
First step that was necessary would be converting the Rust to its lowest-level form. That sounds like the fully-desugared" form you describe.
But you still need to resolve types to do dispatch, so that's a nontrivial amount of work.
mrustc can currently compile rustc, but the produced rustc doesn't pass the entire rustc testsuite (yet).
A colleague of mine was considering writing a Rust-to-C++ compiler that was similar, but offloading most of the dispatch/resolution work onto C++. This is actually possible, you can turn method dispatch and autoderef into template resolution. You can do stuff to fake type inference too if you know the program is correct already. This is much harder however and I'm not yet sure if it's 100% possible without doing some typechecking in the Rust-to-C++ compiler itself.
You could however take rustc --unpretty=typed (or whatever that option is these days) output and transform that really easily.
This requires you to be able to independently verify that the two ASTs are equal if you resugar, because you can't trust rustc's output for this. In fact, the poc trusting trust attack I wrote[1] would still go under the radar here.
Verifying ASTs as semantically equal is much less work. But there's a loophole here, it is possible to add type annotations to type inference'd Rust to get different behavior. This is because (among other things) integers are inferred more loosely (an uninferable integer type is defaulted as u32). It's possible that a trusting trust attack would be able to propagate itself merely by flipping around the results of inference.
Even if we looked out for that, you still have the problem of dispatch, where a backdoored rustc could change the method being dispatched to be a different trait. Now this isn't something that would work if you assume that the original code was code which compiled fine with rustc (because rustc complains when there's unqualified ambiguity). But we can't actually trust the original compiler to have handled this correctly either!
In both cases there would be need to be traces of weird code in rustc for this to work, but it might be possible to hide this.
(This is also somewhat a problem for mrustc, but to much a lesser degree)
[1]: http://manishearth.github.io/blog/2016/12/02/reflections-on-...
PyPy and Graal show good performance. The can be considered special cases of partial evaluation, where the first Futamura project works. They are not generic enough for the second and third projections, though. They even require some help (e.g. annotations) for the first.
Why is it stupid?
Eventually, one day they will find out that their blocking is too strict and that they should restrict the blocking to sites which actively try to attack the users computer...
So it is more of an execution problem, but as it fails quite often you could call it stupid to invest into such a feature.
2) Since everyone has smartphones they will just access the same sites with their personal device which is more time-consuming.
Added: if you think I'm joking or being unfair, just look at the compensation tables for just about any government outfit. They top out around a salary that is considered average for software folks in some places.
If he's not paid well then I hope he realizes that by merely being aware of HN and GitHub puts him in the top 10% of developers and he can do a lot better than working at a place that restricts his ability to educate himself.
It seems odd that github would be blocked especially given HN isn't unless your company has some sort of pathological fear of accidental IP dilution, so perhaps it is a mistake that will be quickly corrected once pointed out?
I guess contractors are all over the map in how they try to make sure to comply with the "rules", though.
I didn't read it as a request for assistance, more a complaint that we weren't doing enough to assist by providing a link that would work in the specific circumstance that the poster finds themselves in (that we could not have known about ahead of time even if it was something we should be responsible for fixing or working around).
And I did offer assistance by suggesting the only practical way forward (unless you count HN banning github links because some of its readers can't access them as a practical way forward!): discussing the matter with the IT department. Especially as it _could_ be a mistake (externally sourced block lists being overly aggressive unbeknownst to them?) rather than a deliberate action. And if it is a deliberate action the poster may need to investigate what policy the block is part of to make sure they are not accidentally breaching it by other actions.
> no need to be so rude.
I used the exact same tone in my reply as I was replying to. A little passive-aggressive maybe, but if I was rude then so was what I replied to. I know two wrongs don't make a right, but then again neither does the first one on its own so I've not made the situation any worse.
1) holyjit aims to be easy:
> As a user, this implies that to inline a function in JIT compiled code, one just need to annotate it with the jit! macro.
jit!{
fn eval(script: &Script, args: &[Value]) -> Result<Value, Error>
= eval_impl
in script.as_ref()
}
fn eval_impl(script: &Script, args: &[Value]) -> Result<Value, Error> {
// ...
// ... A few hundred lines or ordinary Rust code later ...
// ...
}
fn main() {
let script = ...;
let args = ...;
// Call it as any ordinary function.
let res = eval(&script, &args);
println!("Result: {}", res);
}
> Thus, you basically have to write an interpreter, and annotate it properly to teach the JIT compiler what can be optimized by the compiler.> No assembly knowledge is required to start instrumenting your code to make it available to the JIT compiler set of known functions.
2) holyjit aims to be safe:
> Security issues from JIT compilers are coming from: > * Duplication of the runtime into a set of MacroAssembler functions. > * Correctness of the compiler optimization.
> As HolyJiy extends the Rust compiler to extract the effective knowledge of the compiler, there is no more risk of having correctness issues caused by the duplication of code.
> Moreover, the code which is given to the JIT compiler is as safe as the code users wrote in the Rust language.
> As HolyJit aims at being a JIT library which can easily be embedded into other projects, correctness of the compiler optimizations should be caught by the community of users and fuzzers. Thus leaving less bugs for you to find out.
3) holyjit aims to be fast
> Fast is a tricky question when dealing with a JIT compiler, as the cost of the compilation is part of the equation.
> HolyJit aims at reducing the start-up time, based on annotation made out of macros, to guide the early tiers of the compilers for unrolling loops and generating inline caches.
> For final compilation tiers, it uses special types/traits to wrap the data in order to instrument and monitor the values which are being used, such that guard can later be converted into constraints.
> Moreover, the code which is given to the JIT compiler is as safe as the code users wrote in the Rust language.
> As HolyJit aims at being a JIT library which can easily be embedded into other projects, correctness of the compiler optimizations should be caught by the community of users and fuzzers. Thus leaving less bugs for you to find out.
It's funny that the current tiny non-informative blog post starts with a "tl;dr". Maybe it should be "ts; du" instead!
It’s been around for a while but was most recently popularised by Dan Brown in The Da Vinci Code.
He has a small following on Reddit and users post updates of his whereabouts from time to time. He still sporadically streams live video to the TempleOS site. Last broadcast was from an internet cafe.
You say that as if the schizophrenic person, someone who lives in an alternate reality due to psychosis, has any kind of agency over taking his medication. You can only make that choice when you're sane, and even then the drugs aren't perfect and people routinely decide to come off them for some psychotic reason - yes, you can become psychotic while on anti-psychotics - and failing that they come off them because they're so ashamed of being mentally ill from all those condescending and "well-meaning" (read: superior) people telling them what to do and how defective they are all the time that they want to prove they can handle it. Plus there's the issue that the alternate reality is way more interesting than this one.
Your suggestion that it's somehow Terry's fault is about as helpful as telling a homeless person - who is much more likely to have a psychotic illness by the way - to "just get a job".
> you'll show them
> go off the meds
> you can't handle it
> man fuck those people tho, it's all their fault
> and failing that they come off them because they're so ashamed of being mentally ill from all those condescending and "well-meaning" (read: superior) people telling them what to do and how defective they are all the time that they want to prove they can handle it.
Which is a really bad reason.
Sometimes it's for a non-psychotic (but maybe ill-advised) reason like "they made me obese and diabetic."
Edit: The comments were posted 6 hours ago, check his user page https://news.ycombinator.com/threads?id=TempleOS
It might be that, but not "obviously". In fact it could also be totally unrelated.
One can imagine devs naming something "Holy" without wanting to reference TempleOS, the Graal VM etc.
Religious inspired references are perfectly common in themselves.
"The name is a reference to another project named GraalVM, except that the goal of this project is not to make a VM, but only to make a JIT as a library."
The main point was that if they came up with this name without the intention of punning holy shit, I have low confidence in their abilities to succeed in general and with Firefox in particular. But also don’t take this too seriously.
Also, this[1] blog series is a must-read for interested beginners.
0 - https://github.com/stoklund/cretonne/ 1 - https://eli.thegreenplace.net/2017/adventures-in-jit-compila...
For the moment it uses dynasm, just to get the prototype working, but I expect to change that in the upcoming months.
(I am the author of HolyJit)
Jokes aside, I love this kind of work from the Mozilla team. This and the bits going into Quantum are really amazing pieces of software engineering in my opinion.
AFAIK Mozilla does not drive the evolution of ECMAScript (not on their own anyway), so it would be a way to provide new features faster, providing better feedback during phase 3 (and possibly even phase 2) and providing time to implement features they could not so far (e.g. ES6 TCO)
So, I agree it's overwhelming; books about JavaScript from two years ago are already out of date on a lot of fronts. But, it's also resulted in a really powerful and concise language.
It's a feedback loop, I think. Almost nobody uses the web for serious audio because the web sucks for serious audio, and thus almost nobody who uses the web for serious audio is working on the standards for web audio. I admire anyone who can make the web platform work at all for anything audio related.
Is that just hyperbole, or has there been a recent (within the past 6 months) development that has made Linux audio better?
Is it basically just a Rust rewrite which also tries to reduce the complexity of their current just-in-time compiler?
Edit: By calling it "just" a Rust rewrite, I'm not implying that's a simple undertaking, even moreso considering the complexity of modern JS engines.
I wonder why they don't mention it.
also would that mean that the asm snippets could change as llvm changes and possibly cause security bugs if not carefully hand-audited/tweaked anyhow?
Basically, it seems (and I haven't fully digested the code yet) that it uses https://crates.io/crates/dynasmrt for codegen.
My understanding is it's rather more similar to RPython: the developer does not write the JIT, the developer writes the interpreter and the meta-jit generates a JIT from that. The developer can further add various annotations to guide JIT generation for improved performances.
See Laurence Tratt's Fast Enough VMs in Fast Enough Time on rewriting their Converge VM in RPython: http://tratt.net/laurie/blog/entries/fast_enough_vms_in_fast...
Seriously. Look at the example for brainfuck [0], it's less than 70 lines of completely normal interpreter loop, including a one-line macro on the top. What the heck.
[0]: https://github.com/nbp/holyjit/blob/master/examples/brainfuc...
I think a simple next step is to take a small toy language (there's a small rust scheme implementation written by another team member that might serve as a good candidate, if extended with js-style prototype-based objects), and prove this out for an actual type-driven specialization.
But yeah, the possibilities are certainly exciting.
$ git clone https://github.com/v8/v8.git $ cd v8 $ find . -name '*.S'
There's still some macro assembler in the built-in's but it's emitted rather than being assembler.
The MacroAssembler, is basically what is used to produce assembly code in both JavaScript engines.
Almost reminds me of early on in Servo's history: https://news.ycombinator.com/item?id=6268521
(Although I may just be projecting what I want to hear :) it's exciting hearing about new Rust projects, especially new stuff going into Firefox.)
HolyJit:Rust::Rpython (toolkit):Rpython (language)
Rpython being the language and toolkit that is used to create pypy.
Why would you want one?
Can it help you write in a jot?
Can it help you write in a dot?
Why do I feel like I’m trapped in The Land of Doctor Seuss?
I haven't seen any of their posts for a while. I hope they are okay.
I should set a blog up before that.
How is voat, like is the programming/tech comments worthy?
No, no it is not.
I would suggest changing the name from HolyJit to anything else.
I don't have religion, but given an essentially infinite number of alternatives to this not-very-funny one, why do this?
People also have a right to be idiots in other ways, such as poking hornets' nests because it's funny.
Maybe that's why as a start-up, though I have very much done the 'disruptive' thing in the past when needed, I didn't do it to be 'disruptive' for the sake of it, and I always try to find a gentle path to what I want to achieve...
I've recently been reading Ogilvy on Advertising[1] by David Ogilvy, one of the most successful 20th-century figures in the industry, and while it's a bit dated (it was written 30 years ago) and, obviously, the subject matter is advertising, it's filled with excellent advice that's just as useful in all professional and personal contexts.
One aside that jumped out at me was "While we are on the subject of taste, I deplore the current fashion of using clergymen, monks and angels as comic figures in advertising. It may amuse you, but it shocks a lot of people."
I had honestly never really thought of it that way, but it's true. Religious people (that is, most people) find it hurtful and disturbing when you mock their religious faith. That might not be your intent, but it's kind of like 10 or 20 years ago when people used to explain that "by f-- I don't mean gay, just stupid." Just because you don't think or don't know what you're saying is hurtful does not mean that others aren't hurt by it.[2]
If Mozilla had used terminology in their software that was offensive to women, or gay people, or non-Western religions, I'm sure they would alter it, and rightly so. If they called a copy-on-write library HolyCOW, and Hindus said they were offended by this name because it mocks their religion, I'm sure Mozilla would, rightly, change it. And they probably know enough not to use such a name in the first place.
Of course, "holy" is not a specifically-Christian term, but a concept shared by all religions, Western and non-Western alike. A devout Catholic, Muslim, Buddhist or Jew is equally likely to feel hurt and excluded when they see you comparing their beliefs to shit.[3]
As for the other half of the name, the profanity ship has sailed in the broader culture and especially hacker culture, but it's also true that many people don't feel the same way, and nearly all of these people are deeply religious. I think it's reasonable to expect that religious people in the open-source community ought to accept that you or I will sometimes say "shit" for humor or emphasis, but it's also reasonable for them to expect we'll meet them halfway by not making them the butt of our jokes.
There are plenty of equally-funny "JIT" puns that don't come at anyone's expense; "GoodJIT," "JITHappens," "JIT'sTheBomb," and so on. I recommend using one of those.
I'm not a prude. I wouldn't scold you for saying "holy shit" (or "holy cow") in a conversation with me. I haven't scrubbed those phrases from my own casual vocabulary either. But, if someone told me I'd offended them, I would apologize and probably feel bad the rest of the day, just as I imagine almost any of us would. When you're participating in the open-source community, your audience is mainly strangers, with many different beliefs and backgrounds, whose first impression of you[4] is formed by what you've written on the internet. So it doesn't hurt to be a little more careful.
[1] https://smile.amazon.com/gp/product/039472903X/
[2] Hearing that speech a few hundred times, and maybe even giving it a couple, certainly didn't make high school easy for closeted me.
[3] Not that we shouldn't try our best to avoid needlessly causing offense whether it's to one group or many.
[4] And not only you, but the organizations, projects and communities you are involved with.
I hear what you say about my use of "idiots", but someone stirring the sh*t (yes, I can use (light-to-full) industrial language too, in its place) without any need, casually and semi-deliberately offending (or worse) many many others, is I think behaving idiotically.
BTW on your point [2]: I am enraged when I observe people slandering two groups at once, casually: a direct target (A), by noting them as obviously as bad/gross/etc as assumed obvious horrible out-group (B). Gahhh!
Holy jit, that JS is running quickly. (Is it fast in that case? Quickly doesn't feel right...)
Let me repeat: God approves of JIT compilation.
You might want to submit a bug with your specific hardware to see if anyone can identify the cause, because that's not normal.
> With the advent of WebAssembly appearing in browsers, the virtual machine that we talked about earlier will now load and run two types of code — JavaScript AND WebAssembly.
I bet all those plugins will be back.
Its just that implementing the DOM requires a lot of other (complex) things 1. Stable JS object ABI 2. Integration with the GC garbage collector 3. Better integration with modules
Personally, I'm hoping that we actually get a lower-level subset of the DOM apis that doesn't rely on as many OO features so we can bind to it easier, and avoid more DOM manipulation overhead (though I have no idea what this would look like).
And just because it isn't there today, it doesn't mean it won't be there tomorrow.
Will there be a day where we can make desktop class apps that can run in the browser and not have to wade through the insanity of what is out there now...Should I use ReactJS, VueJs, Flow, Svelte, AngularJS, EmberJs, nextJS, on and on and on.
Pick an approach and standardize? Or just let people ship large WebAssembly binarys with their own runtimes and UI frameworks.
If you aren't making games or crunching big numbers, wasm isn't for you yet.
Even so, let's assume wasm added all those features today. History shows it would still be a decade before you could ship to all your users. That may not matter for fancy startup X, but it certainly matters for the biggest and most important businesses.
I love the potential of wasm, but I think it's way too soon to be preaching about it being the end-all be-all of the web.
If WebAssembly is good enough as C and C++ target, it is good enough as any of those processors.
As for the potencial of WebAssembly, there are already ongoing efforts to port .NET and Java runtimes to it, and I am looking forward to Adobe porting Flash to it as well.
So it will come, WebAssembly + Canvas + WebGL is already quite usable.
Typical client programs intended to run on nascent versions of those architectures were not intended to be delivered and installed very often. In contrast, web pages might get changes deployed to production multiple times a day. Needing to deliver a compiled runtime solely in order to run your client side code is going to be a nonstarter for the overwhelming majority of developers.
Now, at best, we can hope that many developers will collaborate to make caching easier by agreeing to only use one specific version of each runtime, delivered from a single well-known source (though experience with e.g. jQuery means that we shouldn't hold our breath). In the meantime, extending WASM to obviate the need for delivered runtimes will even the field between Javascript and every other managed language.
A GC runtime for a language with like Oberon semantics is just a few hundred KB, way less than sonething like minified jQuery.
Using LLVM to target WebAssembly is not a requirement.
Also the point isn't implementing the best GC algorithm, rather a good enough one.
Daniel 5
Does that have some significance to Rust or Mozilla? Or is this a case of copy pasta?
[1]: https://github.com/nbp/holyjit/blob/master/Cargo.toml
[2]: Permanent link to line: https://github.com/nbp/holyjit/blob/1f20eb41de2dae14179815c7...
The HolyJit repo has a brainfuck jit example in the repository: https://github.com/nbp/holyjit/blob/master/examples/brainfuc...
(They actually don’t even need that line, as Cargo already infers this via convention)
[0]: https://en.wikipedia.org/wiki/Brainfuck
[1]: https://github.com/nbp/holyjit/blob/master/examples/brainfuc...
It is a reference to the best programming language ever created.