Pixie – A small, fast, native Lisp
pixielang.org
pixielang.org
It should be mentioned that I put about a year of work into this language, and then moved on about a year or so ago. One of the biggest reasons for my doing so is that I accomplished what I was looking for: a fast lisp that favored immutability and was built on the RPython toolchain (same as PyPy). But in the end the lack of supporting libraries and ecosystem became a battle I no longer wanted to fight.
Another goal I had was to see how far I could push immutability into the JIT. I learned a lot along the way, but it turns out that the RPython JITs aren't really that happy with VMs that are 99.99% pure. At one point I had a almost 100% immutable VM running for Pixie...as in each instruction executed created a new instance of the VM. It worked, but the JIT generated by RPython wasn't exactly happy with that execution model. There was so much noise in the maintenance of the immutable structures that the JIT couldn't figure out how to remove them all, and even when it could the JIT pauses were too high.
So anyways, after pouring 4 hours a day of my spare time into Pixie for a full year, I needed to move on.
Some other developers have commit rights and have pushed it along a bit, but I think it's somewhat a language looking for usecase.
And these days ClojureScript on Node.js could probably be made to handle most peoples needs.
I use Gambit-C scheme for most of my scheme needs because of the speed, but it seriously lacks in libraries, that Racket has in abundance.
As for libraries, the only major feature that Pycket does not support is Racket's FFI, most built in functions can be implemented pretty easily if missing.
The only reason why I have not tried it out yet is
* No mention of installing it with `nix-env -i pixie`, `apt-get install pixie` or `brew install pixie`. Sorry, but I'm that lazy :(. If you make it Linux-only, a mac user can do `docker run -ti pixie -v /src:.` as well.
* No mention of how to connect from your editor. A small Emacs mode which let's you C-x C-e is all that would have made me use it.
It's lazy and dumb, but that is the truth. So for me the only reason for not being a Pixie developer full time are those 2 UX things.
Thanks for your work and keep it up :-).
Oh how the definition of "small" has changed. I actually would like to know how they managed to make something like this so big.
To compare, LuaJIT is about 400 KB, and that includes the Lua standard library, a JIT almost certainly more advanced than Pixie's current one, an incremental GC, and a C FFI.
Neither compilers (well, except C++ ones), nor stuff you usually find in standard libraries, nor a GC should require much code to implement, relatively speaking (e.g. compared to a WYSIWYG word processor). These things are usually small. The compilers for almost every language were < 1 MB in size for the longest time.
I am not saying that Pixie being 10 MB in size is a problem. We have a lot more bandwidth and disk space nowadays, 10 MB is nothing. My point is that a "JIT, GC, and stdlib" package weighing this much cannot claim to be "small" for what it does.
Back in the day, Smalltalk was criticized as bloated because one ended up with stripped binary+image of about 2MB. Someone around that time got SqueakVM+image down to just under 400k. There were specialized Smalltalks in company R&D that got their image down to 45k.
My involvement with Smalltalk started in the 90's. Then it still had high penetration in the Fortune 500, and it was a niche language for complex financial programs. It was also used in Energy.
[1] http://users.cms.caltech.edu/~cs140/140a/Oberon/system_faq.h...
Unfortunately, for an arbitrary CL program it's impossible to tell for sure how much of the CL compiler and standard library it will need at run time, so SBCL takes the easy route and just includes everything. Some of the commercial Lisp compilers are a lot smarter at stripping things out and and can create significantly smaller executables.
I usually avoid the issue altogether by not building binaries and running most things from the REPL or using a "#!/usr/bin/lisp --script" shebang line.
in this case
(funcall #'+ 1 2 3)
or the more scheme-ish (apply '+ '(1 2 3))
i think the symbol plus gets evaluated to #<function+> and then is called with 1 2 3 - CL will do apply, of course, not sure about scheme and funcall. I'm totally willing to accept my mental model is wrong, i don't have an interpreter handy.So if that's true, you run into this problem:
(funcall (string->symbol "+") 1 2 3)
(apply (string->symbol "+") '(1 2 3))
maybe string->symbol does not introduce in to the function namespace, but i bet there's a way or a flag to do it.so anyway, at the end, you run into this problem:
(funcall (string->symbol "any-random-function") 1 2 3)
(apply (string->symbol "any-random-function") '(1 2 3))
this makes tree shaking hard.you could, of course, look for the handful of ways to bind symbols to functions, maybe do constant folding and include a minimal set if the compiler can prove what's actually needed.
but in general i think tree shaking is hard when you can load things like that dynamically.
Treeshakers are generally used for application delivery in Lisp. When you do that, then one usually limits the amount of use of runtime dynamics. Typically you can tell the treeshaker what to remove or what to keep: the compiler, the symbol table, debugging info, etc etc.
Allegro CL and LispWorks have extensive facilities for this.
http://franz.com/support/documentation/10.0/doc/delivery.htm
http://www.lispworks.com/documentation/lw70/DV/html/delivery...
https://www.snellman.net/blog/archive/2005-07-06.html
There is also the possibility to compress the image. See :compression option when saving an image:
http://www.sbcl.org/manual/#index-g_t_0040sbext_007bsave_002...
Commercial Lisp implementations like LispWorks and Allegro CL have a treeshaker for delivery, but that's not too surprising, since these products are for developers of applications and the users pay for such features.
Other examples of CL compilers not saving an image are ECL and WCL.
The ECL manual had the terminology a bit wrong, but it actually does not dump an image.
https://common-lisp.net/project/ecl/static/manual/ch34.html#...
WCL:
https://github.com/wadehennessey/wcl
> For example, the executable for a Lisp version of the canonical ``Hello World!'' program requires only 20k bytes on 32 bit x86 Linux.
I wrote a size profiling tool that can give much more precise measurements (like size(1) on steroids, see: https://github.com/google/bloaty). Here is output for LuaJIT:
$ bloaty src/luajit -n 5
VM SIZE FILE SIZE
-------------- --------------
74.3% 323Ki .text 323Ki 73.8%
12.5% 54.5Ki .eh_frame 54.5Ki 12.4%
7.6% 33.2Ki .rodata 33.2Ki 7.6%
2.2% 9.72Ki [Other] 12.9Ki 2.9%
2.1% 9.03Ki .eh_frame_hdr 9.03Ki 2.1%
1.2% 5.41Ki .dynsym 5.41Ki 1.2%
100.0% 435Ki TOTAL 438Ki 100.0%
And for Pixie: $ bloaty pixie/pixie-vm -n 5
VM SIZE FILE SIZE
-------------- --------------
57.5% 4.39Mi .text 4.39Mi 44.7%
33.7% 2.58Mi .data 2.58Mi 26.3%
0.0% 0 .symtab 1.31Mi 13.4%
0.0% 0 .strtab 978Ki 9.7%
8.8% 688Ki [Other] 595Ki 5.9%
0.0% 8 [None] 0 0.0%
100.0% 7.64Mi TOTAL 9.82Mi 100.0%
In this case, neither binary had debug info. Pixie does appear to have a symbol table though, which LuaJIT has mostly stripped.In general, I think "VM size" is the best general number to cite when talking about binary size, since it avoids penalizing binaries for keeping around debug info or symbol tables. Symbol tables and debug info are useful; we don't want people to feel pressured to strip them just to avoid looking bad in conversations about binary size.
> This executable is a full blown native interpreter with a JIT, GC, etc.
> A small, fast, native lisp
I have issue with the use of the word 'native' here. For me 'native' mostly means AOT compiling.
Edit: Agreed that I don't like the author's use of native either.
Also has anyone tried to create a 'readable' [1] flavour of Clojure yet? If not, why is that, do lispers consider all the parens to really not be a barrier?
When writing lisp your editor should help you at least a bit to make things enjoyable. They can help a lot, but doesn't have to help much; I found for myself that if the editor just highlights matching parens, I'll be very effective.
I have tried a structure editor for Haskell but I found it pretty counterintuitive. I wonder if structure editing just assumes a language with very little syntax.
Nope, it has it's own notation which I find very close to the Smalltalk one, for example:
red>> a: [1]
== [1]
red>> pick head append a 1 + 1 2
== 2
red>> a
== [1 2]
So the second expression is good example of how things work, this is how it looks if I put parens (which are unnecessary in this case): pick (head (append a (1 + 1))) 2
Here interpreter/compiler knows how many arguments each `word` (you can think `function`) needs, so `append` needs 2 - series and item to append, `head` only one - series, `pick` needs series and index to get element. But what about `+`? red>> type? :+
== op!
Operators are infix things of two arguments and they get priority over function calls, that's why `append` is called with `a` and result of `1 + 1` and not with `a` and `1`.It looks a bit tricky in the beginning, I understand, but it leads to very compact and easy to read code.
This is brief explanation which should help you to start reading and writing Red/Rebol code :)
Their terminology is totally consistent with how the field uses these terms.
This technology operates at multiple levels of meta-implementation, so it is easy to get confused what is tracing what, and what is implemented using what at a given time.
But there are also interpreted interpreters aren't there? Which aren't real executables, and aren't native. It would be possible to write an interpreted interpreter for Perl, maybe in a language like Ruby. Jython is a real example of an interpreted interpreter (if you ignore that the JVM has a JIT, but it isn't AOT, which you've said that you think is an important criteria).
And so it isn't useless information.
The distinction is particularly relevant here because the RPython technology they are using to build their interpreter means they can either interpret their interpreter using the Python interpreter, or they can make a native interpreter by compiling their interpreter to native AOT. That's probably why they used that particular wording.
So, again, not only is their terminology consistent with the rest of the industry, they are also making a specific and interesting point here, and it isn't useless information.
Sorry, no, this is nonsense. Scala source code is compiled to Java bytecode, which is interpreted and JITed by the JVM. Pixie source code is compiled to the Pixie bytecode which is interpreted and JITed by the Pixie VM. chrisseaton claims that "native" refers to the interpreter, but this is rubbish ... that's not what the industry means by "native".
I mean, not in use.... What would be the point? Every other interpreter worth using is AOT compiled, it's hardly something to boast about.
One exception might be JRuby, I suppose. Kind of a strange comparison to go out of your way to disabuse.
Pixie programs are translated into Pixie bytecode, which runs on the Pixie VM. Bytecode is, by definition, not "native".
IIRC with RPython/pypy the two techniques are interleaved, but most users I've seen opt for some JIT optimization.
You have said that twice. It was wrong both times.
Sure, you can for example run a Prolog interpreter on top of a Lisp interpreter. It's just not fast, but may have other qualities.
See http://stackoverflow.com/questions/3107299/how-do-clojure-pr...
Sorry, but that's nonsense. "native lisp" means that the lisp source is compiled down to machine code. To claim that an interpreter is "native" is to misuse the terminology.
No, that would be a lisp that produces native code, not a native lisp.
To contrast, there are non native lisps written in Python, in Javascript , etc., including atop of other Lisps.
A native compiler for Clojure would be an interesting project, but completely new languages competing on speed or size are going to have a really tough time beating the existing Common Lisp implementations, not to mention the CL library ecosystem.
Not to say the CL library world is very big compared to Python or Javascript, but it has most of the important bases covered, and it's certainly bigger than a brand new language like Pixie.
My largely uninformed armchair opinion as to why, is that the author is very performance-driven, and in the end it's very difficult to beat the JVM performance-wise. Lesson: if you want high-perf Clojure, you already have it on the JVM.
Personally, I think there's room for a simple small native Clojure implementation where performance is not top-priority. Small footprint, quick startup, access to native C libs. Still holding out hope for that one.
But, that said, that is basically what it is: https://github.com/clojit
I would love to write a JIT for it, but currently its a half done interpreter. You compile Clojure to Bytecode (of our own design, see clojit-doc) and a C interpreter (we started in Rust but the pre 1.0 chances and some other stuff killed that).
All of this is inspired by LuaJit, specifically my goal is to work on a tracing compiler. It just takes to much time to get there.
I've been tinkering with some ideas for another language in that space (the Crafting Interpreters book is looking to be pretty helpful in getting me there: http://www.craftinginterpreters.com/), but it'd be really nice if there were enough options for me to not feel the need to create my own.
Especially in Python, which is what Pixie is written in.
I don't think a performance-driven developer would have picked Python in the first place.
> So this is written in Python? > It's actually written in RPython, the same language PyPy is written in. make build_with_jit will compile Pixie using the PyPy toolchain. After some time, it will produce an executable called pixie-vm. This executable is a full blown native interpreter with a JIT, GC, etc. So yes, the guts are written in RPython, just like the guts of most lisp interpreters are written in C. At runtime the only thing that is interpreted is the Pixie bytecode, that is until the JIT kicks in...
From the docs: "Ferret is a free software Clojure implementation for real time embedded control systems"
I kinda wish the likes of NPM would output cool skeletons and heavy-metal vampire chicks while I wait ten minutes for it to install stuff.
(To me, it is a context thing, I think - in lagnuages like C/C#, there parens, brackets, parentheses and even angle brackets flying around, and it is not a problem at all.)
I recently had to write a project in Ruby and all this random syntax is driving me mad.
Having maps with `{}` syntax is also nice if you ask me. Makes reading a little bit faster.
What difference does it make when representing code?
Vectors also differ from lists in that they are ordered and indexed, so in macros and special forms they tend to be used to represent positional bindings.
(let [a [x y]] ...)
(let [x y] ...)
(fn [x y] ...)
As for parenthesis, you have special cases too: (defn foo ([] (foo 0)) ([n] ...))
And sometimes, people aren't sure about what to use: https://github.com/bbatsov/clojure-style-guide/issues/64I was thinking about it over the day, and realized maybe having special syntax actually is confusing - maybe "normal" lisp with only parens is better. Thanks for the eye-opener.
The idea behind the Lisp reader approach is that custom syntax is used for reading/printing objects of different types, not to demarcate syntactic elements. For example, double quotes are for strings, #P"" will read pathnames, and so on. In Common Lisp, some characters like [ and { are reserved for the user, meaning that no conforming implementation defines a custom syntax based on those characters. And you can define [a b c] to mean (vector a b c), which will build a vector, when executed, to hold the current values of a, b and c. The existing #(a b c) vector syntax is a literal vector that contains symbols.
From this point of view, for Clojure, I don't think it is a bad idea to have a short syntax for vectors, set and map literals. But then, those objects are used when representing code, not because it has a real added value but because the visible syntax is a little bit nicer (maybe, maybe not). That IMO complexifies tools that work with code and does not really fulfill a practical purpose. If that was for practical reasons, I think bindings would had been better defined as maps: as far as I know, they seem to fit more naturally than vectors for this task. For example:
(let {a 20 b 10} ...)
Somehow when compiling or interpreting the code, you could do "(get symbol env)" where "env" is the surrounding lexical environment associated with the let, which would be computed partly from a parent lexical environment and {a 20 b 10}. The environment and the map could be of the same type and be easy to combine.
But there is no such consideration, and it is actually a good thing that there is no link between how the code is represented and how it is interpreted (also, there can be different interpretations).Basically, I find Clojure a little bit confused about its use of external data representation for code. I would have preferred a simpler syntax for code and keeping those notations for data, where they are self-describing without additional context.
That applies regardless of how they are spelled, be it [...] or #(...).
(let ([a 5] [b 6]) (display a+b))
I find this easier to parse mentally, but I can see why one wouldn't do it, especially for cond clauses where you might end up having to actually browse parens to add or remove before or after the ].
I use paredit now, so there is really no need for me to do think about parens much at all, but old habits die hard.
YES! :D
In Clojure it is impossible to surgically modify a data structure. That is, you can't do something like:
(SETF (CAR (CDR x)) 'foo)
which would alter a data structure.
You can modify a data structure, but it returns a new data structure, yet the old one remains if it is not GC'able.
All of the common data structures have this property. Sequences (eg, lists), arrays, maps and sets. If you change the 50 thousandth element of an array, this returns a new array. The old array is unaffected. Yet it gives the performance you expect of an array. (Meaning no apparent cost of copying.)
If you're writing a search procedure, it is trivial to transform one chessboard into a different chessboard, but without concern about the cost of copying (close to zero), or having altered the original value (you haven't). Other variables that have a pointer to that first chessboard don't see any changes.
There are other things such as a great story about concurrency.
Hope that helps.
For the few times when you want to squeeze the last drop of performance, you can use transients[0].
You can use map, which if I remember my CL, is like MAPCAR. But instead of map, you can use pmap which will do the processing on all of your cpu cores.
Besides, the semi-lazy approach of Clojure's pmap (allocate futures consecutively) doesn't seem to convince everyone:
https://stackoverflow.com/questions/2103599/better-alternati...
https://www.reddit.com/r/Clojure/comments/20lxmv/just_what_i...
On the one hand, when you want to parallelize, you are delegating tasks to workers and you want them to do what they need, independently of you. But on the other hand, laziness introduce a dependency from you, because workers cannot produce a result before they are sure you really need it. This basically slows down parallelization and that's why I am not sure Clojure's pmap is a good general solution.
Note that Clojure Reducers, mentioned here in Zombie Metaphysics (ppmap: http://www.braveclojure.com/zombie-metaphysics) and documented at https://clojure.org/reference/reducers, use strict sequences.
Laziness is great when you need it, but when you don't you shouldn't have to pay for it.
http://clojure.com/blog/2012/05/08/reducers-a-library-and-mo...
I found it useful to override the print and prompt method so that they are silent, requiring the script to explicitly print to stdout/stderr if desired. I reuse the socket-repl reader to enable exiting via :repl/quit which is echoed in the bash script after the script is loaded.
The obvious downside is you have to start at "server" before any of this works but I think the main interface limitation was the inability to interact with Clojure from the command-line and with other command-line tools.
(Not voluntelling halgari to do things! I would just like tn understand.)
I tend to agree... but doesn't 'all the way Clojure' imply at least some exposure of an underlying runtime, be it either Java (Clojure) or JavaScript (ClojureScript)?
I'm admittedly a bit behind the curve wrt the latest developments in the Clojure along these lines.
I'd settle for a gdb backend, really. But printf-debugging is unacceptable.
Also, how fast is this? How does it compare to other lisps(or schemes) in terms of speed? Can someone port the benchmarks to https://benchmarksgame.alioth.debian.org.
I really wish Pixie just had a REPL and compiled to native code via LLVM so I could give coworkers a .exe.
Edit:
Just looked at the github page. It doesn't look very active to me although I wouldn't say dead. The only small lisps I know of in development are Picolisp and Newlisp.
They say WSL apps can not interact directly with Windows apps. So yes and no to your question. I would say it would be similar to running in a virtual machine but more integrated.
Startup is considerably longer than on other platforms (30 sec or so), but I haven't had any issues with runtime performance, and having a Lisp running on what is essentially an embedded platform is very useful.
(It'd be easy to make the argument that my agent code doesn't do much, but it doesn't diminish the fact that Clojure is able to usefully work on an RPi. :-))
If you could magically port real Python or ruby or JavaScript code to pixie, but keep the same algorithms and architecture, I doubt it would change much.
Slowness is more influenced by things like data structure layout, allocation patterns, serial vs parallel I/O, context switches, and just plain not understanding your code once it reaches a certain size.
There seems to be a fetish for JIT compilation in a lot of new language designs and it confuses me. Julia is probably the language making the best and most appropriate use of it. It actually has good data structures and types which complement it.
If you want to make systems fast, you work on the bottleneck. I'm saying that people think too often that summing integers is the bottleneck, when it plainly isn't.
Although I have worked in the domain of numerical applications, thus my nod to Julia. Anybody who works in that domain isn't going to be using something like Pixie; it's too impoverished in terms of types and data representation.