Unexecute
emacshorrors.com
emacshorrors.com
At the time, building these systems took several minutes, so it really wasn't feasible to expect users to just load everything they needed on startup. It is highly non-portable, of course, and has caused headaches for Lisp builders ever since. Amortizing startup time over a larger amount of work is still the only portable solution I know, along with keeping initialized application state in databases rather than in-memory data structures.
If you want to be fancy, you can make the on-disk VM-image format a database (SQLite, LevelDB, whatever) so as to avoid writing it all out every time. Then it becomes cheap enough to write out a differential state that you can make the runtime do it automatically at intervals, after certain operations, manually with a sync(1)-equivalent call, etc.
Although, I wonder if I could hack together a lisp on top of the cog vm. That would be cool.
Anyways, what kind of awesome job do you have where you get to write st for a living? :)
I took this approach in a game engine I developed at one point. Common Lisp has a very general meta-object protocol that allows you do things like this transparently (see e.g. [1]). I believe I used Berkeley DB as the backing store, which supports in-memory caching of objects. With this approach, I didn't need an explicit save-file format, everything was just "there" on disk, automatically. As far as that was concerned, it was pretty cool.
Unfortunately, it was not fast. At one point, I did an experiment where I ripped out the DB and replaced it with a in-memory hash-map implementation. This was about 10x faster, (despite the supposed in-memory caching at the DB layer). I got an additional similar speedup when I ripped out the meta-classes for the objects.
Turns out, these abstractions are expensive. Writing nice sequential code on compact in-memory data structures has substantial benefits (if you want performance).
EDIT: "There are currently three different data stores that support the Elephant API: Berkeley DB, Postgresql via the postmodern library, and any database supported by the CLSQL library including SQLite3." — so, write-through, then.
I was working in a project in which a "make" would load a huge tree full of rules scattered in sub-makefiles, and take a full 30 seconds to evaluate before kicking off the first incremental build, really putting a damper on the edit-compile-test incremental cycle.
I got sick of this and so took the unexecute code from GNU Emacs into GNU Make, and added an option to do "make --dump" to dump an image after loading the rules. The restarted make image would kick off a build recipe almost instantly.
I imagine if the author of this article looked at a Unix kernel they would start arguing that we should remove context switches because they don't understand them.
Easier IMHO is to keep maintaining a proper malloc implementation. glibc were not able to update to ptmalloc3 anyway, and now they want to destroy ptmalloc2 even more. I wouldn't trust them.
"My .emacs is older... than most engineers..."
:-)
Does this quote have a source?
Until proven otherwise. >.<
To keep pushing a project from the 70's and keep it current is quite a feat.
Rebuilding from scratch is its own challenge, but to keep maintaing something at this scale and have such quotes be realized by hundreds of people is awe-inspiring.
This were the days when fvwm was considered new.
So even today, if my IDEs aren't around, I get to use my old friend Emacs.
I have run back to ensime and sbt-mode. I have an old Thinkpad but with 8GB memory. The class flipped between IntelliJ or Eclipse depending on platform. I find them exhausting. They consume a ton of resources (as IDEs go, Emacs and piped in tools can be 8 megs and swap all damn day and still cut through them in performance and speed).
There is just too much going, despite having nice auto-complete and error check defaults, they struggled to pick up Scala compiler bin paths, even at defaults, and there are just so many damn menus. What did emacs, and later helm perfect? M-x and search all the damn options. Even every OS on the desktop imitates this stuff now! I know Intellij, to a lesser extent Eclipse can do that, but they are so damn busy with GUI tiles and buttons and UI choices I cannot be bothered to wade through the chunkiness of it all, even if it means I will have the potential for superior coding. I feel it is the lazy way out.
I am slowly eeking my way to intermediate Emacs use, and I keep forcing myself because all the GUIs and lack of consistency cannot make up for the minimalist, trim your own topiary of combined tool bliss that is Emacs for me!
As for Scala I don't know how good the support really is, as I never managed to use anything besides Java for production code on the JVM.
I'm an idiot.
Double shift (to lookup every project symbol) and meta-shift-a (to lookup every IDE function, like M-x) gets me almost everywhere, aside from navigating along definitions/implementations/usage sites which is the reason to use an IDE. Toggle distraction free mode (via meta-shift-a) to get rid of everything beside the open file.
I think the typical way to set it up is to manage the Scala runtime itself (maybe via SBT? I don't even know off-hand), so having it pick up an existing installation is kind of non-standard. Can't remember when I've ever manually downloaded Scala.
Emacs can do those things, I just have to selectively decide what and how and when. I also do not like dependence on proprietary software. Yes, I know there is a free copy. But when they try cool new UI or feature X and I must suffer through it, I yearn for the consistency of Emacs for three decades for a reason. I have been doing Windows sysadmin for approaching a decade. W10 imposing changes I hate on me was the last straw: at home, open source or nothing. Haha.
Yes, re SBT. I tried both methods (system-level install of SBT and Scala with pacman and ~/.sbt or whatever install), it went so-so.
I've never loved IDEs, because they increase the amount of time between having an idea and starting to code it, whereas with emacs, I can just start writing code.
The only IDE that is any good is the Smalltalk IDE, because it's not so much an IDE in the java sense as it is a realtime window into the soul of your application and environment. I mean, if you thought the modern LISP's realtime interaction was good...
But yeah, even then, I start to miss emacs. And I'm not even a serious emacs user. Which is to say, my config can probably fit on only a few pages, and I don't know the 100+ set of basic keyboard shortcuts yet.
Regarding Lisp many in the FOSS camp that never experienced commercial Common Lisp IDEs should give a try, the REPL is only the tip of the iceberg. Maybe Racket is the closest one can get without paying.
Back in the mid-90's I couldn't even get Emacs to do what Borland got me with their tools or even what I later learned Xerox environments were capable of.
Energize C++ with their custom Emacs was probably the closest thing that one could get, if the company could afford it.
You've said. However, I am unsure as to whether any Lisp implementation outside of the lisp machine truly allowed for programming inside a live environment, such that there is no distinction between the live environment and the code on disk, where the image IS your environment.
Although being able to serialize your live environment to disk goes a long way. As a schemer, I'm still drooling over that particular feature.
http://basalgangster.macgui.com/RetroMacComputing/The_Long_V...
Allegro Common Lisp now has an express edition:
http://franz.com/products/allegrocl/acl_ide.lhtml
http://franz.com/downloads.lhtml
Also for a bit of time travel, there are quite a few documents from Interlisp-D at Xerox PARC available:
https://archive.org/details/bitsavers_xerox?and[]=subject%3A...
On Xerox, images are called symbolic files.
The Xerox Lisp Machine (different hardware, different Lisp, different OS) used a different approach will full managed source code in the image, with managed files as kind of a way to persist sources.
See: http://lispm.de/genera-concepts
LispWorks or Allegro CL are in that direction.
It's not as radical as the InterLisp-D system, which is in many ways similar to Smalltalk - when in fact many ideas in Smalltalk are coming from Interlisp and its earlier BBN Lisp, including part of the runtime technology, which enables dumping/restarting images.
http://flukus.github.io/2014/08/19/2014_08_19_Android-Take-b...
There's a different workflow. I always feel like I'm fighting the IDE, the instant I do anything that is even slightly off the beaten path. And they're complicated: everything about the IDEs that I've used seems to be an overcomplicated mess of settings, options, and menus, whereas in emacs, I could just put a single config option into my init.el to do what I want instead of jumping through all those hoops just to get to the setting. This is a problem compounded by the fact that Java, the only language I've used in an IDE, has a complicated development environment and multiple complex build systems, meaning you HAVE to fiddle with things to a fairly large extent.
To further things, it's rare for an IDE to be as programmable and configurable as emacs, or even close. It's considered above and beyond the ordinary to write extensions for most IDEs: the job of an extension is to add some big feature, like support for a new programming language. My first emacs "extension" was a simple function that wrapped around some actions I found myself constantly using in sequence, so I could type one command sequence instead of six. It wasn't something I wanted to dedicate a macro to, I just wanted a command to exist that would do what I wanted.
And then, 5 minutes later, I had one.
It's rather a stable and very useful feature, which just recently got under attack, because glibc doesn't want to maintain malloc_get_state() / malloc_set_state() anymore. XEmacs has a portable dumper pdump, which is a hack compared to emacs unexec.
See https://lwn.net/Articles/673724/ and esp. https://lwn.net/Articles/673815/
I recently re-added unexec support (i.e. native compilation) to perl5 in my cperl fork, but I haven't got it stable yet. Super trivial on solaris, but not so easy on elf, darwin and windows with its various compilers and the different way to treat their segments. But it's still the easiest way to do it, compared to pdump or a seperate compiler or criu, which is still not in the kernel and not in debian. They are saying it's unstable for 2 years, where it's stable for 1 year already.
self-dump via crui is besides unexec the most stable variant, but it needs either a service or root perms, first of all a package, and then it's not so attractive because it produces many files instead of just one binary.
If glibc removes malloc_get_state() even if darwin still has a similar API, I'll happily build with a static ptmalloc3, which is the better variant of the glibc ptmalloc2 anyway, and they never where able to update this. (much faster, but needs a bit more memory for housekeeping).
https://github.com/perl11/cperl/issues/176
https://github.com/perl11/cperl/commits/feature/gh176-unexec
Being used to this layer of abstraction being hidden in ld(1) doesn't mean that reimplementation of it is wrong, just perhaps an ill-advised maintenance burden.
An appeal to ASLR is a bit fallacious - that technology developed for C's deficiencies, including ISAs tailored for it. There are likely better ways than using an untyped language that begs attackers to forge object handles, and then kludging around that by making attackers guess.
https://www.freebsd.org/news/status/report-2016-04-2016-06.h...
https://news.ycombinator.com/item?id=12178766
I love Lisps, but to an amateur with rudimentary infosec coursework this does scream scary.
I LOVE THIS SITE. Hello early weekend entertainment reading, emacshorrors.com ...
ASLR is also a weak defense by itself as I noted in that FreeBSD thread. Any ASLR-enabled application is, essentially, about 1 infoleak away from being no different than a non-ASLR application. If you want to stop exploits, invest in real mitigation tech, not cheap defenses, and suddenly you won't need to worry about this so much. (Feel free to add randomization on top of working defenses as an extra layer -- just not by itself.)
People always hem and haw over how this makes ASLR not work for XYZ, and thus is a 'security nightmare'. But then, very strangely, we still find that it's very possible to write all kinds of exploits that bypass usable ASLR anyway in a variety of applications, with only a single infoleak, coupled with a vulnerability, at only marginally higher work effort. ASLR does not actually eliminate a class of vulnerabilities, it only adds an extra step in the process of exploitation. It can only truly prevent a narrow class of exploits, under very specific constraints.
So, it does seem like there's a real security nightmare happening, but it's almost certainly not because random XYZ thing lacked ASLR at compile time. It's because we invest our 'faith' in stop-gap, mostly futile defense mechanisms that are obsoleted without much extra effort, normally.
So I wouldn't call ASLR a weak defense - it closes off a lot of exploitation avenues by itself, and it can make exploiting interactive situations quite a bit harder. Finding that second infoleak bug isn't always quite so trivial.
The Windows implementation is actually better in this regard since executables are randomized as well as libraries. However the randomization is the same for all processes and only changed on boot (because libraries on Windows usually uses relocations rather than PIC so the pages wouldn't be shareable if they were randomized per process), so an infoleak in one process can be used to attack another.
I have played with CCL and SBCL on and off for a while. Ironically, the AUR package I used Clozure Lisp with does not even exist anymore on Arch AUR repos.
https://aur.archlinux.org/packages/ccl-bin/?comments=all (That won't work)
I get the trade off, but if you read SBCL dev blogs, like PVK, you will recognize really competent programmers and elegant solutions.
The problem is these gods (I revere them) cannot possibly encode all the hacks and knowledge. They have to code, and the nuances of such "hacks" are internalized by them and forgotten. Ironically, a lot of shit talk on Lisp transpires here, and it is one of the most sophisticated developer-productive tools I have seen, with many laudable features other toolchains brag about. Lisps had them for years. This is not to be an elitist jackass; the cultural shift to obsolence (sadly, my view) is why only diehards know and others avoid implementation details and functionality: we don't use it, so we don't care.
Either way, dead code is scary, especially in such scenarios with binary image dumping and bit fiddling hacks. I was ironically aware of such issues with SBCL when I stumbled on Gentoo Hardened list info when people explained how Lisps have trouble in such constrained environments.
https://archives.gentoo.org/gentoo-hardened/message/d2fb14f1...
I worry later on such methods leave blobs of weak executable routines and other items of interest to get a foothold once the unsung heroes of now move on.
On the other hand, it is painfully obvious to me higher level languages, with or without memory control, allow more flexbility, depth, and alternatives to avoid buffer overflows and other counterpoints made in other threads here. The devil is in details that fly over my head at the speed of light.
Life is a tradeoff. The fact that the world uses FreeBSD (WhatsApp) and ASLR is not fully implemented means, yeah, it is complicated.
It is scary to me, a guy who shits himself at macros and could not code a rudimentary one-pass compiler if you held a gun to my head.
https://en.wikipedia.org/wiki/Action_at_a_distance_%28comput...
"In computer science, action at a distance is an anti-pattern (a recognized common error) in which behavior in one part of a program varies wildly based on difficult or impossible to identify operations in another part of the program."
Why are people so afraid of C buffers? They enable great performance, and they are no more and no less secure than bounds-checked buffers (as long as the array indices are properly contained), and well, this is what makes computers and programming run fast rather than being forced into slow interpreted execution.
In my reading, the author doesn't even complain that it's a security issue, just that it will likely fail on systems with certain security measures. And the author never says it's bad simply because it's self-modifying; rather, that the self-modification is a "platform-specific reimplementation of a linker".
A maintenance burden that is over 20 years old still working fine.
unexec is hairy, but it's mature, greybeard hair. It is no worse than any JIT, for starters.
But now they want to deprecate malloc_{g,s}et_state(), without a plan to improve ptmalloc2? They already failed.
Deprecating an API for no good reason is failure, not an improvement. It not only breaks emacs, it breaks other software also. unexec is used in perl5 also, btw. just not in the official perl5 packages.
Because it's several orders of magnitude more incomprehensible than code that fails a cyclomatic complexity check, is why.
When you have table lookup, and a table is dynamically modified in such a way that this is influenced by the table lookup, that is as "exciting" as self-modifying code.
https://msdn.microsoft.com/en-us/library/office/gg615596(v=o...
>Applies to: Office 2007 | Office 2010 | Open XML | Visual Studio Tools for Microsoft Office | Word | Word 2007 | Word 2010
I feel as though the gp comment is referring to far older versions, although without clarification, it's hard to be sure.
FWIW, old Office documents were actually CFBF (Compound File Binary Format) files - think of it as FAT-in-a-file, allowing for multiple independent streams inside, with transactions. This was very commonly used on Windows in the OLE/COM era, because it was the underlying format for OLE Structured Storage. It's what allowed a Word document to embed another arbitrary document in an extensible way. The underlying data in the streams within CFBF was a loose object graph dump.
It all makes a lot of sense when you have your OLE glasses firmly on - it's basically a natural design that follows if your world consists of OLE objects and interactions between them. Look up IStorage and IStream to see what I mean.
The side effect of all this, however, is that the data inside an old Office file is not laid out in a logical way - streams consist of non-sequential interleaved blocks in a seemingly random order (depending on what was written when), some blocks may contain garbage data, and so on. So it's very difficult to reverse engineer, which is why it took so long back in the day, and the results were often unreliable.
That's actually the "new" binary formats. The usage of CFBF seems to have been introduced in Office 4.2 (at least Excel 5.0 is the first Excel version to use them, it's hard to find information about the old Word document file formats).
> The side effect of all this, however, is that the data inside an old Office file is not laid out in a logical way - streams consist of non-sequential interleaved blocks in a seemingly random order (depending on what was written when), some blocks may contain garbage data, and so on. So it's very difficult to reverse engineer, which is why it took so long back in the day, and the results were often unreliable.
I don't believe the OLE compound file format has ever been much of an effort to reverse engineer. But the CFBF based Office documents are also basically just blobs of the older binary formats saved in a more structured way. The issues with Office documents have always been a question about their sheer complexity combined with their tight coupling to the internals of the Office programs. This still shines through in the OOXML formats which contains lots of stuff like "position something the way it was done in Word 5.0".
The very first Lisp implementation did that already. It could dump and read memory images to/from tape. From that on, most Lisp implementations, and not just Common Lisp, are doing it. Some have extensive capabilities in this area (like tree-shaking or generating shared libraries which can be included in programs).
Guile Emacs is dead, isn't it? No activity in over a year: http://git.hcoop.net/?p=bpt/emacs.git
> 28/02/2015
Emacs-guile is working today. That's a first after many many years of talk and proof-of-concept stage efforts.
I hope the development will pick up speed once guile 2.2 and emacs 25.1 are out. Both projects underwent some big changes/improvements lately.
One sign of live is that bpt's improvements to guile elisp were rebased on a recentish guile-master. This branch lives in the main repository now.
http://git.savannah.gnu.org/cgit/guile.git/log/?h=wip-elisp
Also, I have seen some recent commits from emacs developers to guile. So the two projects are talking to each other.
Now, imho, the success if emacs-guile strongly depends on whether they manage to share the load across several developers. A more open development culture surely would help.
The most serious problem with it I had is that there's something wrong with the makefile dependency checking. For certain types of change you have to do a full rebuild. But I'm pretty certain autotools is scarier than any memory dump, so I just put up with this.
I never felt confident to implement the feature directly :\
Your solution needs several seconds, unexec just a few milliseconds.
Oh wait does this all get serialized to disk every time you exit emacs? Is that what I'm missing? If so, why not run the linker at emacs exit time, to optimize the loading sequence?
It's never serialized. It's just dumped. Like a core file, with just proper headers, sections and segments, so that it can be executed. A proper COFF/ELF binary. A core file has all the segments but misses the headers.