HyperCore Linux: A tiny portable Linux designed for reproducibility
hyperos.io
hyperos.io
So, let's break down the fields of CS for which this should be applicable:
* Systems: This won't work except for the few systems projects that are entirely in-RAM AND will work on tinycore's kernel version
* ML: This, I can see, especially with the seeming focus on dataset management. Much ML is compute-bound and the overhead of using the FUSE FS's is hopefully negligible.
So, is this focused on ML and ML-using code and experiments? If so, I think that should be clarified. I think a lot of systems folk will be (rightly or wrongly) turned away from it due to the seeming overhead of the various hyper* extensions. Not to mention that they are all written in Node/JS (Again, rightly or wrongly, many systems folk will not want to run their stuff on platforms written in JS)
I like the direction this project can go, but there seems to be a lack of focus or direction in your mission right now.
> So, is this focused on ML and ML-using code and experiments?
Your completely missing the point. Please look into 'Computational Science' (or Scientific Computing, or Numerical Analysis), that applies to 80%+ of all disciplines that exist today (e.g., computational physics, comp. biology, comp. economics, comp. aspects of engineering disciplines, the list goes on).
However, this still isn't clear in their website. I will give them the benefit of the doubt since they are early in their project, but I think it would behoove them to nail down their mission sooner rather than later.
This is probably what I get being in the CS bubble. =)
But I agree, it would help if they expand on this from the pure CS point of view. Especially if they mention things like containers, CS people would be interested in finding out what they're up to.
Seriously though, this kind platform is a critical component in scientific reproducibility. The dream is that we can have code, data, and the results of the composition of the two in the same revision control system. A minimal layer to allow the execution of linux software would support the use of legacy code and binaries in this new platform. Javascript has its advantages, but it's a waste to build a data RCS and require all functions on the data to be written in it.
And to go a bit further, it's not just for science. For example, you could write a HN clone in dat. I could fork it and get both your code and all the posts.
Although I think he might have been pointing out the problem of reproducible builds[2].
[1]: https://nixos.org
Debian is updating a lot of their compile/build scripts to make things 100% byte-for-byte identical and you can see that while there isn't a lot of work involved it's still going to take a while to hand-update many thousands of pacakges:
If you compile a "hello world" type C program you should get the same binary when you compile it again provided your toolchain (C library, C compiler, linker etc) are all the same.
However certain C macros like __DATE__ make the binary change (in this case based on the time of compile). Additionally sometimes environment variables like your working directory and your username get into the binary.
Why is this bad?
If the build server for Debian gets hacked or if a developer's machine gets hacked (for some projects), the hackers can modify the binaries. If the program is not reproducible then there is no way to tell that something has gone amiss. If the program can be built reproducibly, someone else can build the code, produce the same binary, and validate it.
This is more scarier in the case of a "Ken Thompson" style hack, where the C compiler binary is modified so that it compiles normally but inserts backdoors in certain libraries, and also inserts its modifications whenever it is building another C compiler.
If the "Ken Thompson" style hack is ever pulled off on a linux distro, there would be no real way to tell without analysing the binaries.
Provided your initial C compiler is good. Having a chain of reproducible builds where each build produces the same binary would prevent against this. Currently we are just producing random binaries and relying on trust which is horrible.
You likely don't have the old toolchain installed, so you reinstall everything, pull a VM, whatever, and then rebuild the last official version.
I would be so much more comfortable if what you just rebuilt had the same md5 that the one deployed on the field. Because even before applying and testing the patch you are not absolutely sure you have rebuilt exactly the same application.
/usr/local/lib/node_modules/linux/cli.js:60
fs.accessSync(keyPath)
^
TypeError: Object #<Object> has no method 'accessSync'thats because xhyve fails with
vmx_init: processor not supported by Hypervisor.framework
Unable to create VM (-85377018)
might be better to output the error from xhyve right away- use an abstract concept as a metaphor/marketing for your product
than
- use product "L" that has been influencial and important historically, currently and likely well into the future as your Product's dependency.
- Then market your product as product "L" on one of the most important package managers for startups, enterprises, hobbyists.
They have but a single distrobution. They are not "Linux". They just are an example of it
Please be explicit and say that this is NOT a linux distribution, and in all seriousess I don't understand how this can be called linux, can someone explain me?
Interesting, using NPM. Now that I think about it, it's the only package manager that I know of that runs on and is commonly used on all 3 major OSes.
/s kids these days, Javascript this, node that...
Which is all of them, except for windows update.
Except for if you have Microsoft Visual Studio installed, in which case you probably have a make system, nmake.
Unfortunately, nmake isn't quite compatible with regular Makefiles.
Luckily, there's plenty of resources for how to set up Makefiles so they work for both nmake and cmake.
But since make is so simple, it's been compiled and available for windows for decades,so just shipping it with the makefile for windows is probably simplest.
That or docker toolbox.
The Dat team has a solid track record creating separate packages for binaries; they're responsible for maintaining Fuse, Electron and LevelDB packages on NPM. Optimizations are likely to follow.