Keep your files inside your VM
msimkunas.lt
msimkunas.lt
I've developed almost exclusively in VMs for over a decade. One reason I use VMs is to isolate the execution context from development and deployment. I used to passthrough from the host but poor performance and the lack of inotify are DX barriers. Passing through the FS to the host is a no-go because of subtle executable things like git-hooks that could enable sandbox escape.
The simplest and best approach I've found is to use a git remote on the host to push branches to/from the guest sandbox. I can commit on the sandbox fs and treat the host as an upstream remote. On the host I pull from the sandbox and push up to GitHub/etc. It's a bit more process but becomes second nature quickly and requires no extra tools. This also works well for remote servers.
Another approach I've used is lsyncd to sync files from the host to the guest (Mutagen is another cool syncing tool). In practice, though, I've found syncing to be a footgun too. It's too easy to edit a file on the host and blow away a change inside the guest with no undo. This is one reason I've found explicit git push / pull to be cleaner.
Additionaly, having a VSCode server inside the VM is also a pretty sweet solution if VSCode is an option.
I have started to use Syncthing to share the entire workspace between host and VM and it works great. It's near instant, and even works between Windows and Linux, and it's local sync.
Keeping the IDE inside the VM would feel more appropriate for your isolation, but would be a continual source of friction to keep IDE plugins/settings/tools in sync. Using a remote development workflow feels like it would eliminate many isolation guarantees.
What does your VM provisioning setup look like?
Re: provisioning, really basic automation around KVM. I've been meaning to experiment with Firecracker for more transient runtime environments.
Could you please expand on this? I’m having trouble understanding this.
Of course, there’s also the question of dependencies. If I’m working on a project that I trust that pulls in hundreds of NPM dependencies, I implicitly trust that the project deemed those dependencies safe to pull in. It would be impossible to always operate on the assumption that you’re potentially pulling in hundreds of malicious packages so the chain of trust has to take over at some point.
However, I imagine you were more or less referring to singular projects that you don’t fully trust and like to experiment with in which case I agree with you and would probably use a VM without any kind of file sharing tbh.
Ah, interesting indeed.
Unfortunately recent versions of Samba makes it hard (impossible?) to spawn remote processes; else I would be able to use eshell inside Emacs, which is pretty awesome over SSH.
Fwiw though, windows now supports openssh via a powershell command to install (e.g it’s not some third party binary). I’m not sure if it supports sftp, but it may be worth looking into.
I did try using OpenSSH, but there were issues with using Powershell through eshell over TRAMP. It expects a POSIX shell. I also tried using Cygwin to run OpenSSH, but that didn't work well for my needs.
[0] https://www.gnu.org/software/tramp/tramp-emacs.html#Connecti...
[1] https://www.gnu.org/software/tramp/tramp-emacs.html#Ad_002dh...
I also have some little quality-of-life scripts on the Linux side for things like editing shared files in the Mac editor, opening shared folders in Finder, etc. These make it a fairly seamless experience to go back and forth between the two OSes.
Not for me! I did that and one day ran a Git command on the Linux side of a folder shared from the host.
It crashed Git!
This turned out to be reproducible. It was quietly corrupting my local Git repo too.
Turns out VMware Fusion shared folder has extremely bad incorrect caching in its Linux client, with no option to turn the caching off.
So I switched to using SMB and Samba to share a Linux guest folder with the Mac host. That mostly worked but for 9 months I found a random file would occasionally just disappear, without me noticing for a few days.
I kept losing data, about once a month at random. I thought there must be a subtle bug in my application's careful handling of those files. The concern delayed production deployment for 9 months. But no, when running on a proper Linux server my application code had been fine all along
It turned out to be the Mac SMB client had a an awful bug that caused it to very occasionally delete random files in busy directories. The files had no relation to files being operated on, except for being in the same folder and recently used. It wasn't an operation race, it was blatantly sending the wrong file id for an operation on a different file by another process in parallel. Debugging that was intense. For months I had no idea, and thought it was my fault I lost data about once a month. Eventually I managed to trigger the fault more reliably with an artificial load test, and then it took a solid day of tracing everything you can imagine to confirm the cause was the Mac SMB client sending an egregiously wrong command on rare occasions.
Because the shared folder random file deletions were so vexing, rare and distressing, I investigated all sorts of causes, and in the process found another. A sequence where Emacs (GUI Emacs on the Mac host) asks to confirm editing a file "locked" by another Emacs (not really, just due to stale state), where if I gave the wrong answer sequence and aborted at the wrong monent, and editor backups were configured a particular way, the file ended up removed unexpectedly. I figured: Jackpot! Found the cause of disappearing data due to a race in Emacs in rare circumstances. I thought I must have triggered that sequence a few times. So I fixed it in my Emacs with great relief.
Than a month later another file was missing and I realised there was a second cause of disappearing files. That turned out to be the crazy bug in the Mac SMB client.
I have shared folder solutions that seems reliable enough now, both for sharing host folders with the guest and guest folders wih the host. It involves NFS and duct tape with Samba to translate permissions, because Macos doesn't support NFSv4. And CIFS (SMB version 1, obsolete) because later SMB versions had some problem. It's not great. The permissions don't map well and I had to change the config last time I upgraded Macos version.
All because VMware Fusion shared folder had such poor cache coherency.
> I also have some little quality-of-life scripts on the Linux side for things like editing shared files in the Mac editor, opening shared folders in Finder, etc. These make it a fairly seamless experience to go back and forth between the two OSes.
I have little scripts like yours too. They're great! Being able to edit a file in the Linux guest using the Mac GUI version of Emacs is really helpful, as is being able to open a PDF or DOC or XLS that's stored in the Linux guest. That sort of thing is the motivation for sharing my Linux home folder with the Macos host.
Oh, I also use Emacs TRAMP in a way that's so seamless I forget I'm using it, When editing Linux guest files at the path where the Linux home folder is mounted.on thr host, Emacs is configured to do so transparently over SSH to the guest instead. It works so well I forget I'm using TRAMP, and this avoids weird file permission issues and potential bugs like those I've seen before when writing to the shared folder directly. It also means when I run a shell command or compile etc. from Emacs with a buffer open on a Linux file path, the command runs on the guest, which is perfect.
And yes, it works with Windows: https://virtio-fs.gitlab.io/howto-windows.html
The main advantage of this approach is that your code resides on a native file system. This allows you to take advantage of this performance gain where it matters most, which for me is during runtime (I develop web applications and IO performance is of paramount importance). I care much, much less about performance on the host side.
I'm not backing up entire VMs through ZFS snapshots nor VM snapshots, and don't need a ZFS-enables backup target either. I'm merely backing up the relevant files (e.g. compose stack definitions, config files, data files) that are in the chroot, from the VM host, to GDrive, using popular tooling such as duplicacy.
Yes, virtio-fs is an option but I should have added that one of my goals was to not lock myself into a platform… I want to have the possibility of running my VMs on Windows as well and Samba seems like the easiest choice there. Just spin up a Samba share inside the VM, mount it on the Windows host and you’re good to go. No need to install any additional dependencies on the host.
I’ve tried the Windows NFS driver for Vagrant a while back and my conclusion is that it’s best not to force a square peg into a round hole… Not to mention that these custom solutions can experience software rot throughout the years.
Many years ago, I experimented with all sorts of syncing mechanism from my windows desktop to a remote linux VM and the sync lag was just unbearable. I'm much happier when the files live on the development VM.
I kind of went overboard and tried to plan for a disaster recovery plan for those development VMs, even though there's not really a very strict SLA required. I can provision those VMs using ansible, but any work in progress which wasn't pushed to github is at risk. To address that, I've got hourly rsync jobs of my home directory replicating to a rsync.net account in Zurich (why Zurich ? gives me the lowest latency from the hetzner server in Germany). I then asked rsync.net to enable 48 hourly snapshots on that account (with additional weekly, monthly). It has become essentially impossible to accidentally lose data and I feel very confident working for days on a git feature branch just committing locally, or perhaps not even committing, given i've got those hourly snaphots. My bare metal server could vaporize and i'd be back up and running in hours. Even if the hetzner datacenter was hit by a nuclear weapon, I would instantiate a VM on any cloud provider, provision it with ansible and be back up and working within hours.
I've spent WAY too much time working on my little disaster recovery scenario but it's been insanely fun.
Syncthing is my key to keeping my ~/git tree in sync between the two environments. Just have to be careful not to try to flip back and forth quickly between local/remote as syncthing delays can be ~10s.
Why don't you use git directly for that?
I gave this setup a go with PhpStorm. Host-to-VM sync is pretty seamless but you always have to manually sync changes from the guest to the host, so there’s that. Additionally, syncing can be very slow on large directory trees. I often found myself staring at a loading animation waiting for the IDE to catch all the changes and sync them to the guest.
AFAIK it is not recommended to open a project located on a network drive in a Jetbrains IDE but my experience with it so far has been great. Latency is not a concern because the VM is local.
Most importantly, I wanted to avoid a syncing situation where you have the source and the destination with two possibly differing states. This difference implies the possibility of conflicts which complicates things even further.
Edit: there is also JetBrains Gateway as a solution now but I find it less of a seamless experience than vscode remote, e.g. it requires using the large r 4 CPU / 32GB VM option via GitHub Codespaces and it doesn't sync plugins.
If it's a database or otherwise expensive to recreate file (like a certificate that's subject to renewal rate limits) keep it somewhere you can get to it if the VM needs recreated. Typically this is a mounted volume on a SAN in larger self-hosted environments.
Desktop eDLP agents run on the hypervisor and MITM the connection between the endpoint and the destination (and/or monitor the filesystem). But if you don't have the agent on your VM...all the agent sees is an encrypted session passed through the shared adapter.
sshfs is surprisingly good if you get the parameters set just right.
Edit: the point here is about shifting the area of compromise. I would very much prefer to compromise on IO performance when writing/reading files with my IDE rather than sacrifice the web server performance on the guest.