Warp – self-contained, single binary applications
github.com
github.com
This is not to take anything away from this project, just making an observation. The programming field often feels like it's in a giant loop constantly re-discovering what went before.
It might feel that way, but this is driven by economics. Basically disk, memory, and bandwidth is far cheaper than it used to be so who cares if you waste a GB or so copying the same libraries all over if you don't have to solve for dependency hell?
Not in the cloud.
2) Application space on the server is vanishingly tiny compared to the data you're processing, so it is a far more generous trade-off.
Just an example: https://blog.acolyer.org/2017/04/03/a-study-of-security-vuln...
If there's a bug in Go's TLS implementation, you need to identify, recompile, and redistribute every Go application you use. Whereas if it were using the system's OpenSSL library everything would magically be fixed by a simple `apt-get upgrade` or `yum update`.
I don't buy this argument; it's every bit as easy to `pip install` something as it is to `go get` something.
> If there's a bug in Go's TLS implementation, you need to identify, recompile, and redistribute every Go application you use. Whereas if it were using the system's OpenSSL library everything would magically be fixed by a simple `apt-get upgrade` or `yum update`.
I think the fix is fairly trivial; package maintainers just need to introduce vuln scanning into their automation. As soon as a security issue pops up for one of their dependencies, their automation automatically compiles a new version and publishes it to the package registry so the end user can pull it down with their next `apt-get upgrade` or `yum update` or whathaveyou.
Not at all. Package maintainers and the Security Team in Debian do plenty of manual work to backport and test security fixes.
> As soon as a security issue pops up for one of their dependencies, their automation automatically compiles a new version
That pulls down new releases for the dependency instead of backporting a security fix. Now you have no guarantee that the new binary will behave like the old one (minus the vuln). On the contrary, it would be practically a new release.
A lot of companies have security policies allowing security updates on live production systems. A complete rebuild against new libraries is not that.
I wonder if that sounds arrogant to just me. As a user I DO CARE. Its frustrating to find applications shipping their own copy of libraries wasting precious disk space. When I was using a Mac, I had at least 4 applications shipping their own copy of Qt5.4 ... aaaarrgh
Back on the Linux desktop, package managers handle this beautifully. And for the cases where you need to skip the distro, Flatpak still lets you use system dependencies. Everyone wins.
And for whoever claims storage is cheap, please link me the magical place you're buying yours from. Thanks.
And storage is ridiculously cheap. You can get an 8TB HDD for $150[0]. Or a 1TB SSD for the same price[1].
[0]https://www.amazon.com/Seagate-Expansion-Desktop-External-ST...
[1]https://www.amazon.com/Blue-NAND-1TB-SSD-WDS100T2B0A/dp/B073...
In many package managers you can have multiple versions of the same application without a big deal -- in most cases you don't want to SUPPORT multiple versions of the same application, so there's always a trade-off.
In my experience, only if the distro has specifically set things up to allow that (Python2/3), or you find a PPA maintained by someone. Otherwise, time to compile from source I guess.
I think part of this comes from various tradeoffs becoming more or less important at different times. Processing power goes up so you move to fat clients, then people move to battery powered devices and heavy lifting makes more sense on servers again.
But nowadays it has terrible security (and to some extent reliability) trade-offs - if everything is a static binary and there's a security issue in a library, you need to upgrade every single binary depending on it, which is much less reliable and more prone to error/oversight.
As security and reliability have increased in importance, so the "cost" of the static binary route has actually gone up significantly in my view compared to 20-30 years ago. So it's a bit odd to me to see this model (and closely related ideas like flatpak etc) gaining currency. There must be better ways to manage the "DLL hell" issues than that.
The concern has more relevance to client applications, but then the cost becomes related to update strategy, and if every app is tied to an update manager process(whether it's an app store type of thing or a custom one) then the thinking becomes similar to the cloud server setup: Why bother, I'm going to blow it all away in the next update.
When the dependency management is exposed to users, there's a much bigger surface area for troubleshooting and support. It makes sense that devs favor static if they expect the whole environment to be disposable. It isn't like it was in the 90's, where patching was occasional but major new versions were essentially new products.
Also, if you're talking about native, highly-distributed desktop/server applications, then the problem with using shared libraries is even worse. The only exception is system-level libraries that are updated automatically by the OS/distribution.
Finally, on Windows, with static binaries you can be sure that the binary is a particular version/build and is signed (this ensures that the binary version/build is valid). You don't get much better peace of mind than that. We've been distributing our installations and other binaries like that for over a decade, and it works great.
Even GNU/Linux made such a big deal of moving from a.out to ELF binary format, due to easier support for dynamic linking, which quite cumbersome with a.out.
Now almost 30 years later it is a big deal to support static linking, go figure.
* those who want applications that take the least space on the drive, and share most of their data / libraries across themselves - shared electron runtime, steam runtime shared with the os * those who want applications to conform exactly to what the developper used since the developer can guarantee a "working" configuration
Ideally it should be up to the user to choose what configuration he wants, but the problem is that it's up to the developer to make that choice. Thus, the eternal loop :
applications released as a blob -> this takes too much disk space, we should use shared libraries ! -> applications released as shared libraries -> DLL hell -> applications released as a blob -> ...const varname = staticRead"filename" embeds a given resource at compile time - to make the project completely selfcontained if need be, and just for the hell of it.
I know you can do this in several other languages, but the Nim way is - unsurprisingly - the most ready-made, obvious and transparent one I've come across.
The most popular one is Starkit/Starpack which uses the Metakit database (which is a heirarchial database) and appends that archive to a loader (either a script, as a stub, or to an executable which includes the stub with the ability to open itself). This database can be modified externally using the starkit tools (sdx), and when you pack all your resources into it, it appears from within the interpreter that all the files are just normal files even though only the database is being accessed.
It has been a often-requested feature for at least 2 years.
https://github.com/dotnet/corefx/issues/13329
(The last comment referring to warp is indicative)
The current .Net Core 2.x runtime has had this feature since it came out though (last year?).
(Hello world + runtime is like 100 MB and several hundred files)
https://github.com/dotnet/core/blob/master/samples/linker-in...
If you use some extra options to disable crossgen and link System.Private.CoreLib, you can get the size down to 10MB, at the expense of startup time:
It's still early days in terms of what's available in the runtime (e.g. the AWT subsystem is missing), but it's pretty damn impressive anyway.
We use Kubernetes to trigger rolling restarts with new images when we release a change to the base image, and so far it's been painless, but a lot of work went into it. We use Gitlab, but any CI/CD should allow you to do it.
See the linked code at the bottom of the page
We just use a requirements.txt [1] for each service and run pip with the -r flag in the dockerfile. No Virtualenv in the container.
Most of the time we run these containers with docker-compose locally, but sometimes we want to run the service outside the container. For that we create a virtualenv outside the folder where we keep the source for the service. The reason for that is that we don't want to accidental copy the virtualenv in to the container. It wouldn't do much there but it would increase the container size.
You can do this easily with just "python3 -m venv", but there are some tools that help with that as well. I personally use pyenv-virtualenv [2] which just keeps all virtualenvs in ~/.pyenv/versions/<env>, but there is also conda [3] and virtualenvwrapper [4] which also store the virtualenvs in a central directory.
I am not sure if there really is any more to it.
[1]: https://pip.pypa.io/en/stable/user_guide/#requirements-files
[2]: https://github.com/pyenv/pyenv-virtualenv
[3]: https://conda.io/docs/user-guide/getting-started.html#managi...
There are a few things of note:
1. The .dockerignore file can be used to prevent use of your `node_modules` folder during `docker build`, even if you have something like `COPY . .` in your Dockerfile.. This can let you create a new `node_modules` folder from your lockfile as part of the docker build process to create an image for testing/deployment.
2. You can maintain separate Dockerfiles and such for development and for release, so e.g. for development you might use volumes and not copy things in, but for release you wouldn't use volumes for source code.
I can't tell from what you said, but it seems like one of those two tips might be relevant.
It's okay to have a development setup which doesn't use containers and then use containers for deployment (so long as you have an integ or staging environment that uses containers) as well. It's quite reasonable to have a venv during development, but inside the docker image to not use a venv at all since things will already be reasonably isolated inside the container's fs.
We use venvs inside of the Docker container because we use pipenv and apparently pipenv's support for installing to the system is buggy and/or idiosyncratic.
The pipenv bit is probably more reasonable, and I'm afraid I haven't used it enough to be sure what rough edges are likely there.
You might wanna check it out.
Here is a small readme I put together...
https://github.com/devxpy/shiv/blob/8d8298d21380dcf0b1970856...
$ zip -r mymodule.zip mymodule/
$ echo '#!/usr/bin/env python3' > myapp
$ cat mymodule.zip >> myapp
$ chmod 755 myapp
CPython will automatically detect it's a zipfile, unzip it in-memory, and import the __main__.py in the root of the module and any nested modules.The reason why you have to handle shared libraries and any other non-.py data yourself is because Python doesn't know what to do with it. You can access it as binary data via pkgutil.get_data(), but linking shared libraries is system-dependent and Python doesn't load them itself. As the dynamic linker can't find a shared library in a zip file, the only thing you can do is extract it separately. This is documented in https://docs.python.org/3/library/zipimport.html
Any files may be present in the ZIP archive, but only files .py and .pyc are available for import. ZIP import of dynamic modules (.pyd, .so) is disallowed.
- https://github.com/gobuffalo/packr
I like that it allows you to `go run` without changing any code.
It was used in multiple programs in both Windows and Linux. On Windows there was a rush of individual updates as each application was fixed... on Linux, simply upgrading that library fixed all applications.
I'd rather be in the latter than the former situation when the next vulnerability happens.
On Windows, the only platform-supported global library format is COM, and that requires a very specific programming style--HRESULTs and all the rest. True, registered COM DLLs can be linked to without separate header/symbol files, and installing new programs/libraries never requires local compilation, but the extra overhead of COM itself, and the fact that it's a walled garden code-wise, makes it rather unpopular. It's also nigh on impossible to port COM code to other apps without resorting to something like WINE (which can't compete performance-wise when running highly parallel COM apps/services). Not to mention how ugly C++ COM source files are.
It would be so nice for us all if Windows shipped with compilers, straight-up C libraries with headers, a package manager, a better shell, ...
Every compiler that produces native binaries for Windows links with these DLLs (C, C++, Go, FreePascal/Delphi, etc.), either statically or dynamically, using C calling conventions.
COM is the Component Object Model, which is implemented in DLLs, but is something different:
https://en.wikipedia.org/wiki/Component_Object_Model
COM, or some form of it, is used heavily with .NET and the new UWP runtime, but those are newer runtimes that sort of sit on top of the Win32 API (to some extent, UWP is actually integrated into the OS itself, but, AFAIK, still uses the underlying Win32 APIs).
Rather than making all of those applications vulnerable at the same time, they slowly become vulnerable as the release binaries are linked against bugged code. If it's not linked at runtime, or recompiled, it'll be vulnerable forever.
I believe you're referring to an OpenSSH vulnerability?
...though, I did a search just now to find when it was from (I remember there being a major one around the time you're thinking of), and instead found posts from only two months ago about another one that has existed for at least 18 years.
I don't see how warp handles the most trivial case of a dynamically linked executable; I see no references to patchelf or other such tricks to ensure it looks within the application's cache directory for dynamically linked libraries, nor do I see any mount namespaces.
Does warp do that? Why would I use warp if it doesn't when there are more generic solutions that can handle such a case, namely binctr, docker, guix pack, nix....
binctr is Linux-only, and appears to need Docker as a runtime requirement? This has neither constraint, apparently?
binctr does not require any runtime component other than a modern linux kernel with unprivileged usernamespace support.
Targetting OSX / Windows does make warp distinct from what I mentioned, though I don't use those OSs so I admit I didn't think about that possible issue.
Anyway, I thought this was a solved problem on both Windows and OSX since they both have conventions for self-contained applications, while linux does not.
This is mistaken for Windows. Even if programs are run from Program Files, they will often source their DLLs from other locations.
tVQhhsFFlGGD3oWV4lEPST8I8FEPP54IM0q7daes4E1y3p2U2wlJRYmWmjPYfkhZ0PlT14Ls0j8fdDkoj33f2BlRJavLj3mWGibJsGt5uLAtrCDtvxikZ8UX2mQDCrgE
Anyone know what this magic is all about?Honestly not feeling warp-packer downloading executable blobs during runtime either.
This is a known and established technique. I don't have the link right now (mobile) but .NET Core does something similar for their native deployment, for example.
Is it possible to instead execute the compressed application code directly and present to the resulting process a virtualized file system that decompresses the dependencies on the fly too? Then you could run the single binary without giving it any filesystem access. One way to do it would be to have have the parent process provide some routine that the kernel can call for reading a particular "file". Or perhaps just have standard embedding and compression methods that the kernel supports, so you would just provide the offsets to the kernel.
I guess what I mean is: do any platforms provide the necessary syscalls to pull that off?
You've basically described a container where the filesystem is virtualized to a safe location on disk.
Disallowing reading/writing data entirely is pretty easy with seccomp or ptrace, though that's not required with namespacing.
https://twitter.com/jordwalke/status/1050277143085608960
Maybe it's just not worth it and a temp directory is not so bad. I just worry about little changes to the way `mktmp` works across OS updates etc. A totally self contained executable that leaves no trace, with virtual file system support is about as isolated and reliable as you could ever hope. I just don't think anyone has been determined enough to achieve it.
I don’t think you can have a truly robust solution without kernel support. Maybe the FUSE kernel-side code already supports, or could be modified to support, per-process ad-hoc virtual file systems? If done right, you might even get existing fs caching and other optimizations for (closed to) “free”.
I also can’t quite explain why I find this so desirable. It just feels intuitively cleaner and more robust to not have to spray files elsewhere into some cache.
[1] https://github.com/dokan-dev/dokan-dotnet/blob/master/README...
Isn't this a huge security hole? The user should not have write access to the application binaries.
I'm sure this still has some uses, but they are rather limited, and very dull.
comes with its own SAT solver for dependency resolution. zeroinstall is what Canonical was considering before it created Snap Packages.
Digikam moved to AppImage, which sounds cool but adds a further system one needs to update, and the automatic update never worked for me (I think their package didn't include the necessary metadata?).
Having all application updates available in one place, and official packages be recent, is the 'killer feature' for me. I used to be a Slackware user, but when I migrated to slapt-get I realised I was done with rolling my own packages and checkinstall-ing them and wanted system where things were more managed - not having half the apps installed ad-hoc.
When you have to install a new system to get just a few apps it feels like I might as well go back to ./configure;make;[check|make ]install.
Seems we need an xdg install record with a unified app that calls the sub-system (pip, 0install, AppImage, shell script, dpkg, whatever) to abstract all that away.
Im using a SQLite btree file as backend, working basically as a key-value store. On the header i define what target triples are supported (the ones the dev built the binaries), and on the runtime open, see if the current host is supported and unpack the bynary payload always checking the hash before running it. (If its on the filesystem already see of the hash matches with the one saved before on the DB as a record)
Giving its a db the owner of the public-key signature(the creator of the file), can modify and add others binaries latter.
Im not disclosing it now, because its just part of a bigger platform, and you dont distribute only the apps, but a whole "container"(not in a Docker sense, but as a multiplatform container more focused on final users instead of cloud backends).
I hope i can launch it soon here on HN.
A better analogy for this system would be jar files that allow more languages than just java.
I'm stoked for anything that eases the pain of desktop deployment that isn't Electron. Right now the cross-platform story is... complicated
Edit: just looked at the warp implementation, it's a self extractor and executor
I wonder what benefits this has over other solutions? Originally when it said multi-platform i thought it meant that one binary could run on multiple platforms like a fat binary. But i think all of the examples require you to have a specific target in mind.
There's something to be said for an entirely self-contained native binary, at least as an easy on-ramp for users.
What Warp does is act like a packer, creating a thin, native executable that compresses and embeds all the files required to run your app, then decompresses them to disk the first time the app is executed.
- Bundle all C/C++ native libraries, and link them all together with the java launcher - e.g you produce your own java.exe that has your C/C++ libraries linked in, then you slap the .jar file - for extra bonus you can slap at the back of your .exe and make it load from there
Voila - you have single executable (almost) with all you need, and you don't need to unpack .dll's anywhere
But... you must compile java.exe yourself...Is that 512 the lowest you can get it with command line options?
?
Jar files are not a self-contained, single binary (executable).
I can't
curl -o binary http://example.org/binary
./binaryIt's really not much better than running a fat jar that contains all of its dependencies.
curl -o binary.jar http://example.org/binary.jar
java -jar binary.jar
This is more like a self-extracting shell script, but more opaque.How is it self-contained if it's not self-contained?