Build me a JIT as fast as you can… [Squeak COG VM]
mirandabanda.org
mirandabanda.org
The actual language of Smalltalk is pretty nice and it is fun to read through posts like this because the simplicity of the language allows for easy learning of different algorithms. The underlying algorithms are pretty transparent.
It also makes it relatively easy to read through code to figure out what's going on and you learn a lot about implementation details by doing this.
Edit: forgot to mention, his Cog VM is pretty nice! It definitely helped improve performance of the system as a whole from slow & pokey to feeling quick and more modern.
Could you expand on this? I've been looking at Pharo w/Cog + Seaside for production software and I would love to hear about any pitfalls.
When you upgrade to a new version of Pharo, you're upgrading your IDE with it. When there is a bug in your system its easy to fix but its pretty tedious to make sure you track every little patch and make sure they're re-applied the next time you build a fresh image. The style of putting your overrides into a method category named the same as your package doesn't sit right with me. There are a lot of conventions like that that are hard to discover that you're expected to "just know". You're expected to peruse mailing lists frequently.
For example I found it really easy to hack in emacs-like key bindings for cursor motion... but I never load those changes into new images just because my code is slightly different from new code and I don't feel like tweaking it every time there are image updates... I've just given in and use the images as-is in the interest of moving forward.
If you run into a memory leak or slip up with some bad code in development, all of those leaked objects hang out in your image until you re-build a new image, or can manage to destroy them manually (I've had mixed results with the latter).
I repeatedly have been bitten by doing something stupid in development and have my image balloon to hundreds of MBs. You can't just Ctrl-C out of that situation... you better hope you have a backup image hanging around and that your most recent changes were committed to Monticello, or else you'll have to go recover them all manually and commit, then load them into a fresh image.
A lot of the libraries you'll find are half-baked. It is a lot easier to whip up something in Smalltalk than other languages (like say, Java), however the community isn't large enough to get enough people doing enough things for the platform to get some good quality work going outside of a few main projects (like Seaside).
Smalltalk's image-based nature does not play well with the host OS. There are some great packages for interfacing to the OS, but you'll be spending some time learning the unique ways they 'smalltalk-ize' typical OS interface areas differently than a language such as Ruby.
Now I still think Smalltalk is awesome. But I'm a lone developer at a tiny company and its akin to using a Formula-1 car to commute to work. One would be more comfortable in a boring Honda Civic. I just don't have enough resources to use it in production is all. F1 cars require large teams, but they're awesome technology!
Lately I've gone back to using Ruby, Rails, and a little Node.js for this company as those are more pedestrian technologies right now. I know one would object to Node.js being "pedestrian", but it really is pretty simple to wrap your mind around, its just Javascript and has decent documentation and isn't too much code to look over.
Smalltalk on the other hand has code for everything, so if you're averse to reinventing the wheel you'll find yourself reading a lot of (good quality) code to figure out what is where and what you should use that is already there. There also isn't a great way to get an idea of what is in libraries you might find on SqueakSource without reading through all of the code, most of it is poorly documented.
In the end, Smalltalk is awesome and frustrating at the same time. Image-based development is both a blessing and a curse. Some conventions of the culture I think are bad practices (e.g. "the code is the documentation, so we don't need to write documentation").
App server farms running hundreds of processes with instrumented JIT VMs could gather a lot of data. This would allow programmers to determine the "steady-state" type behavior of an app server after the "warm-up" period with a high degree of certainty. Would it be possible to save such data as an aggregation of runtime-profiles that could be used to 1) speed up execution and 2) back-port static type annotations to the source code? I could also imagine a special command to tell the VM that "warm-up" time is over now, take type annotations as mandatory, and start aggressive compiling now!
I am quietly (and very slowly, its very much a side project) working on an experimental language, which, while statically typed (using type inference), allows for many dynamic operations as you would expect from a dynamic language. What I am trying to do is detect as much as possible during compilation and storing the rest as metadata. The metadata also contains a lot of other information, such as register usage. Then the runtime is a hybrid between AOT and JIT compilation, where most of the program is native AOT compiled, but that the runtime can choose to recompile segments of code using the JIT, eg, to optimize the common path once it has been detected (inlining, avoiding branch misprediction, optimizing for cache usage and avoiding register spilling etc), or to resolve dynamic types that could not be determined at compile time.
Every now and again, discussions of AOT vs JIT arise and people on both sides have good arguments as to why their compilation strategy of choice can generate the most efficient code. I want to explore mixing both strategies in a single language/runtime in an attempt to get the best of both worlds. I have searched online, but have not yet found anything similar (but if someone knows of a language or VM which mixes native ahead of time compilation with JIT compilation in the same executable, please let me know!)
So far, I have most of the type inference engine and abstract syntax tree done and some prototype code in assembly and C which will eventually become the code generator, runtime library and JIT. I am also still working out semantics and syntax of the language itself (which I need to finish before writing a parser, finishing the AST and writing the code generator and JIT..) A lot still needs to be done and I'm not finding as much time for it as I'd like (too many other projects :P), but I'll get there eventually. I'll probably blog about it when its closer to working.
"back-port static type annotations to the source code"
I love this idea. Profile-guided type resolution or something like that.
In my case, I try to have the compiler deduce the information to make it fast, where it can, but fall back to JIT when it can't, or where running the application is the best way of figuring out what needs to be optimized. The missing link between what I'm doing and what you are suggesting is that the runtime should be able to dump metadata for the compiler to use next time around. I'm going to have to think about this a bit, I think there may be some great applications of this idea, eg, turning assertion failures into unit test stubs. It would be great if the compiler and runtime worked together to help you program 1) more efficient programs and, probably more importantly, 2) correcter, less error prone programs. I love the concept of the compiler "hardening" code that was detected to produce errors in a previous run, especially if the JIT compiler can remove checks that are not needed.
You could write a parse transform that wraps all exported funs in a module in some kind of type analyzing shim function, which would eventually output type specs as an .hrl you could -include().
Type specs will eventually be used by Dialyzer and other Erlang tools to issue warnings and optimize code, so the effort wouldn't be in vain.
Seems like a quick way to figure out the feasibility of your excellent idea.
For example, production code usually turns off a lot of runtime assertions, so this might only be applied to development builds. Having said that, a JIT compiler blurs the line between debug/development and production builds, since the JIT compiler could dynamically insert or remove debug, profiling or tracing code. I like the idea of being able to connect testing tools to a running production system (eg, when an error or bug is detected) and the JIT compiler would insert the required code to make this work for the duration of the test session. The debug information that was generated could then be applied to the next build of the development version (which would include new type annotations, unit test stubs, profile information for the code optimizer and whatever else).
Cool!
the JIT compiler could dynamically insert or remove debug, profiling or tracing code. I like the idea of being able to connect testing tools to a running production system (eg, when an error or bug is detected) and the JIT compiler would insert the required code to make this work for the duration of the test session.
That's going much farther than I'd envisioned it! The mind boggles.
We'll see how my own experiments work out. If anyone would like to chat about the idea, feel free to send an email to my username at gmail.
It should. I recently ran across a page looking to collect data reuse profiles (via user submission) for use in compiler research. Unfortunately, it doesn't appear to be active/working.
The problem with PGO is that you need a representative test which you can run automatically. The benefit of this idea is that the information could be gathered on any runs - not just ones specifically for PGO. Otherwise, yes, PGO does this. Of course, what I want (as explained in my other comments) is more than what PGO offers.
It's also yet another bit of FUD by the Pharo folks. It labels the interpreted VM "Squeak" and the JIT-compiling VM "Pharo-COG". In fact, both Squeak and Pharo will run on either VM, and in a fair comparison, their performance is very similar.