I work at a company that provides backups for both Linux and windows. The entire concept was around block level backups. You could just open up the block device and copy the data directly but it would quickly become out of sync by the time you finished copying it. We did not want to require LVM to be able to utilize snapshots to solve the sync problem. On top of that we had a strong requirement of being able to delete data from the backup.
This resulted in me learning how to build a kernel module and then slowly over about 6 months creating a kernel driver that allowed us to take point in time snapshots of any mounted block device with any fs sitting on top.
Other requirements also dictated that we keep track of every block that changed on the target block device after the initial backup (or full block scan, after reboot).
I wish I could release the source but my employers would not like that :( So at least for me learning how to write kernel modules and digging in to some of the lower stuff has keep me gainfully employed over the years. It is still in use on about 250k to 300k servers today (it fluctuates).
The hardest part was not writing the module, but getting others interested in it enough so I don't have to be the sole maintainer. I like working on all parts of the product and don't want to just be the "kernel guy".
One other time I wrote a very poorly done network block device driver in about 8 hours. You can find it here https://github.com/mbrumlow/nbs -- Note I am not proud of this code, it was something I did really quick, wanted it on hand to show to a perspective employer -- I did not get the job, I am also fairly sure they did not even look at the driver, so I don't think the crappy code there affected me.
EDIT: Thanks for both replies! :)
So, there was a thing that added this feature to the Windows kernel in the product to make this work. Aside from the Linux stuff, which was totally separate. But if you don't need the writing capability, Shadow Copies are good enough, sure..
(Source: I used to work with the guy who made the above post)
Windows Volume Shadow Copy has the advantage of being integrated with he FS a bit closer. So in Windows VSS can avoid some overhead by skipping the 'copy' part and just allocating a new block and updating the block list for the file.
For the Linux systems we had the requirement to work with all file systems (including FAT). So we could not simply modify the file system to do some fancy accounting when data in the snapshot was about to be nuked. So that resulted in me writing a module that sits between the FS and the real block driver. From there I can delay a write request long enough to ensure I can submit a read request for the same block (and wait for it to be fulfilled) before allowing the write to pass through.
> (Why couldn't you use the existing feature?) We did on Windows, VSS is used with a bit of fancy stuff added on top. For Linux there is no VSS equivalent (other than the one I wrote, and maybe something somebody working on a similar product may have written). And even if one did come about (or is and I am just not aware of) it for sure was not available when I started this project.
EDIT: ok so it looks like the management of the snapshot space is a bit different. still, you could probably have wrapped LVM management enough to make it palatable, in less than the time it took to write a custom module
Also at the time LVM snapshots were super slow. I don't have the numbers but even with the overhead my driver created I was able to have less impact on system performance.
I was able to do some fancy stuff to optimize some of the more popular file systems by making inline look-ups to the allocation map (bitmap on ext3). This allowed me to not COW blocks that were not allocated before the snapshot. This was a huge saving because most of the time on ext3 your writes will be to newly allocated blocks.
Wrapping LVM would probably not work, and still require a custom module to do, the user space tools don't do much. LVM really is a block management system that needs manage the entire block device, so existing file systems not sitting on top of LVM would get nuked if you attempted to let LVM start managing those blocks, and you still had the issue that reads and writes were coming in on a different block device. Asking people to change mount points was not a option. There were also some other requirements like block change tracking that LVM does not have the concept of doing. This is for incremental snapshots. Without this sort of tracking you will either have to checksum every block after every snapshot if you wish to to only copy the changes. This module also was responsible of reporting back to a user space daemon that keep a map of what blocks changed. So when backup time arrived we could use this list (and a few other list) to create a master list of blocks that we need to send back. This significantly cuts down on incremental backup time. Some companies call this "deduplication" but I feel that is disingenuous -- to me deduplication is on the storage side and would span across all backups.
So yes, requiring a module is much easier than telling a customer they can't trial, or use this product until they took their production system off line and reformatted it with LVM. Many people hated LVM at the time, it was considered slow and caused performance problems, this was like 8 years ago... LVM has vastly changed and does not have these type of complaints any more. But I can tell you people still are going to scream bloody murder if we told them they had to redo their production images and redeploy a fleet of 200+ servers just to switch to LVM so they could get a decent backup solution.
Also shout out to aseipp! Miss working with you. Have yet to find a bug in the code you wrote :p
Other interesting stories would include:
- Instrumenting certain filesystem operations (all modules share the same memory space; it's possible for one module to take a sneak peak at another's internal structures). This was back before dtrace & friends were a (useful) thing.
- Real-time processing of instrumentation samples. Doing it in the kernel allowed us to avoid costly back-and-forth between user and kernel memory -- but we only did relatively simple processing, such as scaling, channel muxing/demuxing and the like. If you find yourself thinking about doing kernelside programming because carrying samples to and from userspace is too expensive, you should probably review your design first.
Actually, that's not true at all! Just look at Mesa3D and any number of GL proprietary blobs from graphics vendors. Userspace drivers are actually quite common, they just aren't what folks think of immediately when they think about 'drivers'.
This whole communication process had to happen within a small time frame, something like 50 ms. A kernel module handled the communication, decoded the packets, and automatically relayed peripheral messages to other peripherals. The kernel module made it easier to achieve stable and real-time operation within the time constraint.
When writing the same function in userspace, there was absolutely no guarantee that messages would be sent or received in time.
I was working on an embedded system and needed to have fast I2C access (mostly just wanted very small latency bc it's an instrument). I2C from Linux userspace (using ioctls) adds a lot of overhead. I started looking into kernel modules but after a day of research I found out that you can access hardware registers from userspace using mmap("/dev/mem") which is even faster than a kernel module.
Edit:typos
but that would only work if some kernel module or the kernel itself hasn't mapped that IO space, right?
Two lessons learned:
1. dealing with tangible things for once is incredibly satisfying.
2. next time, download the fscking docs instead of relying on an intermittent 3G connection.
It'D be great if My small module Could be published As is, but I'd need to strip all addresses/filenames, and also add more proper locking (but yes, I want to publish it anyway).
I was in Australia for that work too...
The one that I'm most proud of delegates decisions on whether executables can be run to .. userspace. Which is simultaneously evil and brilliant.
i.e. User "foo" tries to run "/tmp/exploit" and the kernel executes "/sbin/can-exec foo /tmp/exploit". If the return code of that is zero then the execution is permitted otherwise it is denied.
This gives you ample opportunity to log all executables, perform SHA1-hash checks of contents, or deny executables to staff-members outside business hours. There's a lot of scope for site-specific things, creativity is the limit!
How did the user space redirection handle potentially high spikes of execve's?
And btw, thanks for all the years of debian-administration.org.
Even though it has some annoying gotchas (such as the fact ARM cores can sleep/frequency scale on demand with no forewarning, meaning cycles aren't always the most precise units of measurement), and is very simple -- this thing ended up being mildly popular. Even though I wrote it years ago, someone from CERN recently emailed me to say they happily used it for work, and someone from Samsung ported it to ARMv8 for me...
(I should dust off my boards one of these days and clean it up again, perhaps! People still email me about it.)
In case that doesn't make you want to run away screaming, I put a post in the "Who is hiring?" article:
i wrote a kernel module that allowed us to keep logs in RAM even across a soft reboot (i.e., if the device crashes or does a software update). it basically reserved a chunk of 1M at the top of RAM (using the same physical address each time). there was a checksum system that allowed it to tell during boot whether the data is still valid or if the bits had decayed.
i've also done a couple i2c drivers for temperature sensors or LED controllers. kernel work is fun.