How do I delete bytes from the beginning of a file? (2010)
devblogs.microsoft.com
devblogs.microsoft.com
I hovered over the link, and there it was.
Raymond Chen is a machine, the T-1000 of Windows development. I've never developed Windows applications using C++ and I still read his blog articles.
br { display: none; }
fixed it for me...But I wouldn't call vim line oriented. Unless you also want to call Notepad line oriented, since you can use Shift + End to do operations on an entire line :D
Sometimes I wonder if keeping the blog up and the permalinks alive is part of the contract Raymond has with Microsoft.
Of course, for most filesystems (and especially ones with sparse-file support, and especially ones with copy-on-write support), such an on-disk data structure is exactly what's already underlying the byte-stream abstraction. So this would just be a passthrough to allow people to directly manipulate that data structure. (In the process probably breaking certain preconditions the filesystem relies on, though, so it would need to track these "low-level extent vectors" as a separate filesystem object type. Programs would still be able to use regular file-abstraction syscalls on them, though.)
Interestingly, despite filesystems themselves not exposing the lower-level extent-vector abstractions, in some other systems that have file-stream-like abstractions, you can operate on "files" this way.
Postgres's BLOBs, for example, are seekable byte-streams, which also (at least theoretically) allow you to insert into the middle of them. (I say theoretically because it's not an implemented API, but it's not exactly hidden from the user, either. BLOBs just get broken out into records in a table representing their extents; you can rewrite the keys of said records in that table to do whatever low-level operations you like.)
Or, for another example, S3 and its competitors let you arbitrarily compose objects as “components” of other virtual objects, which then read back as the concatenation of the objects they contain. (Sadly you don’t get any other vector-manipulation ops than this, but if you keep references to the leaf extents around, that’s often enough to rebuild your extent-vectors any time they change.)
Also, of course, back on the real filesystem, you can just manage your "vector of parts of a file" as multiple files in a directory, and then abstract over that directory that by using a FUSE server which exposes a view where the files are one contiguous file.
Whereas, with these vector-of-buffers objects, either each extent would need to reserve two bytes at the beginning to represent the virtual size of that extent, or the vector would be a vector of disk slices at byte granularity, rather than a vector of plain disk offsets at sector granularity.
I can see the arguments for both.
The former is extremely efficient, and would probably be the go-to for databases that want to be performant without needing to dedicate a raw block device to them.
The latter approach, though, you can implement entirely "on top of" the filesystem (in the same way that the ext3 journal, or APFS containers, are implemented "on top of" their filesystems.) You could have a regular file representing the storage reservation (in full extents) for the vector-object; and then the vector object could be a tree (or something else efficient to rewrite) of byte-granular slices of offsets into that backing storage-file, with logic to ensure that such a reserved slice never straddles two extents, and a cache in the entry for a pointer to the actual extent. These would be perfect for, say, generating archive files.
In fact, since these are such different use-cases, there's no reason to not just give users both.
It'd be similar to—but faster than—what you'd get if you could both talk to, and past, the flash controller for a regular SSD, allowing you to both see an LBA block device, and see the physical extents + allocation map managed by the flash controller underlying the block device. (But you'd need to be able to write to the allocation map yourself, without making the flash controller fall over, which, well, good luck.)
Like you say, this is already how filesystems are implemented with inodes and superblocks. But users don't care - all they need is a stream of bytes.
If you really needed something like this, you could implement it yourself "inside" one preallocated file.
The simplest obvious use-case for pure vector objects: storing Unix text-based file formats as vectors of line-length extents, rather than a single byte-stream. wc(1), head(1), tail(1), etc. would now be O(log n) operations. Any parsing operation that can proceed separately for each line (like the context-free phases of syntax highlighting... or like grep) could be parallelized. The .mbox file format would be relevant again. This would essentially do for Unix what record-orientation does for z/OS—while keeping things "Unix-y."
Another powerful use-case: concatenation. The direct ability to manipulate extents means the ability to compose sequences of extents together into new files, without having to copy the underlying extents. So you can download some large document in parts (either because it's explicitly chunked, ala partial RARs; or because you used range requests to retrieve the file one chunk at a time, perhaps in parallel, as in HTTP or BitTorrent) and then just concatenate the chunks back together in O(log n) time. No need to pre-allocate a file, target your chunk-writes, and keep an external index of which chunk-jobs have succeeded. Just keep a downloads directory and concatenate when you're done.
It's funny; modern filesystems have the theoretical capacity for many operations like O(log n) concatenation, but they only bother to expose one or two (hard-link, and sometimes copy-on-write clone/snapshot). Filesystem authors are seemingly unimaginative and conservative where exposing new APIs is concerned; I wouldn't expect them to ever directly expose the ability to O(log n) concatenate files. But if they exposed a vector-object, it would give you these features for free!
But maybe that doesn't seem useful enough. Let's add one more feature: give the ability to store a "table of contents" with this vector-object, i.e. a labelled interval tree where the extents in the vector are the leaves.
Such a filesystem object would entirely obviate specialized tooling for container/record file formats; rather than mkvtoolnix(1), or avro-tools, or LevelDB, you'd just need a program that took the raw byte stream of one of these container files as input, and spat out a labelled-interval-tree-extent-vector file-system-object representation. You could cd into it, put things in, take things out... it'd seem a lot like a directory. But it'd still fundamentally be a file, in the sense that it'd still fundamentally be a bijection with a seekable byte stream. (Unlike actual directories, where dirents have no standardized collation order; where there's no obvious semantic answer for what part of the "content of a dirent", if any, goes into a treewise cryptographic checksum; etc.)
Certainly, though, if you had this type of object in your filesystem, you could implement your filesystem's directories in terms of it.
Whereas, what I'm talking about, is more like what you get by having a sorted-set data structure in Redis, and putting arbitrary strings in it. Or storing data in LMDB with BASIC-like numeric labels as keys. Rather than a tabular representation, your data store is a heap, with a tabular index set aside to point into it.
(Of course, you can emulate arbitrary-length records over fixed-length records, as Postgres does with its TOAST tables, or—in the small—as UTF-8 does with runes. But in the record-orientation implementation, that’s something that would be exposed to the developer, requiring them to solve the problem themselves or through a library; rather than the file system just showing you arbitrary-length records from the start. And, if you can’t rely on other users doing that work, you can’t really use it as an interchange format.)
'Introduction to the New Mainframe: z/OS Basics' http://www.redbooks.ibm.com/abstracts/sg246366.html?Open
This, and the 'ABCs of Systems Programming' series were pretty good I thought.
I always thought that was intentional to maximize the consulting opportunities.
That sounds similar to mainframe stuff like VSAM, where files are a stream of records instead of a stream of bytes.
This approach is already implemented in ZboxFS (https://github.com/zboxfs/zbox) to do content-based deduplication.
> A file should be seen as a low-level data structure with a limited set of operations.
Why should? For many people, a file is a basic way to permanently store data. Changing the data at any point of the file should be easy.
If you often want to strip a logfile from old data, a solution might be to create new logfiles at a regular intervals and delete old logfiles. Or maybe even better, use a database to store the log messages. Gives you query functionality and allows you to perform more advanced operations, as removing certain types of log messages.
Seems it's limited.
https://gist.github.com/minaguib/1cbe29922b06d50755a2f580b8c...
In general, if you intend to create structures that allow one to reduce file size, reclaim the data, defragment - that sort of thing, a rewrite/coalesce happens to be the viable solution. In general, if allowed to increase, fragmentation will worsen with time.
Edit: then delete the last ten bytes.
There is no function provided for shifting the byte offset of data in a file.
You can write an entirely new file, of course, but the question has an implicit goal of doing this efficiently.
dd skip=10 iflag=skip_bytes if=original.file of=first10ktrimmed.file
despite being old, and the amount of insanely horrible things you can do when you mess up the syntax or mistype something using it, dd is fucking amazing.
to be fair, its not directly editing the file, but making a copy with the first 10k skipped.
you could have a jump/pointer that points to the shortened first block of the file and then at the end of that, a pointer/jump to pick up where the rest of the file continues, but you'd wind up with misalignment and fs fragmentation.
What if you need to chop the first 10 bytes from a file that is more than half your disk?
The question is to how to trim the file head _in place_.
...so missing the point of the question entirely.
I can't recall any times where this would have been useful to me or think of a situation where it'd be a deal breaker. I could contrive a bunch of other operations that are inefficient on file systems like this (remove every other byte from the file, why not?), but doesn't mean that literally any idea we can think up should be supported by the file system apis.
From Unix and Windows we are used that we can insert and delete the end of a file and that we can read, write and seek anywhere in the file, but a filesystem could potentially do much more: Inserting/deleting at the beginning or inserting/deleting in the middle.
Deleting at the beginning of a file would be handy for queues or logs and inserting and deleting in the middle is a common operation for text editors.
There are even more questions you could ask:
- Why not have tabular files that allow indexed access?
- Do we need hierarchical directories or do we need directories at all?
- What about adding tags to files?
- Why having a filesystem at all?
> And I suspect the fact that you can simulate the operation yourself is a major reason why the feature doesn’t exist: Time and effort is better-spent adding features that applications couldn’t simulate on their own.
The initial value is 0, if you wanted to support truncating at the head, you just add to this offset the amount truncated.
In a sector-based filesystem, whenever this offset goes far enough into the file that crosses sector boundaries, you reclaim those sectors as free space. When it lands within a sector somewhere, that sector is pinned and you waste some space.
The problem is more that the userspace APIs don't expose well-supported mechanisms for doing this. Implementing it at the filesystem level is trivial.
What if you have an append-only system which is making remote backups. One service is writing the activity log append only, a second service is reading that, check-pointing, and then truncating the head of the file when the check-point has been committed. No need to do the tricky file-swap trickery.
Also known as a named pipe.
See ( https://gist.github.com/minaguib/1cbe29922b06d50755a2f580b8c... ) for some test notes I took a couple of years ago.
He makes that point at the end.