There are some really interesting changes that can potentially happen when the memristor type of memory becomes available, possibly with its own problems too, but with the huge benefit of moving less data.
There are some really interesting changes that can potentially happen when the memristor type of memory becomes available, possibly with its own problems too, but with the huge benefit of moving less data.
You can do this in C++ right now: mmap a file to a memory region, and just create structs in that region.
This has two problems:
- normally you want to save at controlled points in time, otherwise you have to worry about recovering from states where some function updated part of the data and then crashed (phone ran out of power etc)
- just writing a bunch of internal data structures into a file used to be moderately popular and has great performance, but it's a major headache when you ever update them. You end up implementing a versioning scheme and importers for migrating old files to your new application version. At that point a file format that is designed for data exchange is less headache.
In general mmap already offers you a way to treat your disk like memory with decent performance (thanks to caching), and the number of good use cases turned out to be somewhat limited. I doubt just making that faster with new technology will change much.
So the whole structure in its native form rather than the contents of the json file, for instance a graph would reside in memory and could be operated on directly.
Hence the 'serialized' in the part that you quoted. Once you serialize it the whole thing becomes hamburger and needs to be parsed again before you can operate on it.
This is a more common technique than most people suspect. It's taught in most operating system courses [1]. It's the basis for how SSTables (the primary read-only file format at Google, and the basis for BigTable/LevelDB) work, as well as for indexing shards. It was how the original version of MS Word's .doc files worked, and was also why it was so difficult to write a .doc file parser until Microsoft switched to a versioned serialized file format sometime in the 90s. I think it's how Postgres pages work (the DB allocates a disk page at a time and then overlays a C structure on top of it to structure the bytes), but I'm not familiar enough with that codebase to know for sure. It's how zero-copy serialization formats like Cap'n Proto & FlatBuffers work, except they've been specifically engineered to handle the backwards-compatibility aspects transparently.
It has all the problems that wongarsu mentions, but also huge advantages in speed and simplicity: you basically let the compiler and the OS do all the work and frequently don't need to touch disk blocks at all.
You could do it with relative offsets if you wanted, that's getting pretty close to pre-heating a cache from a snapshot. That way the file would not contain pointers but you could still traverse it relatively quickly by adding the offsets to the base address of the whole file.
Base+offset segmentation does have an overhead since you'd need some extra CPU registers for that if I understand it correctly.
History turned otherwise, unfortunately...
As noted in the article, HP says that memristor memory may be commercially available by 2018, so I guess we'll see then....