Leaky Abstractions
textslashplain.com
textslashplain.com
I think it's so interesting to see how different OS' have approached these long lived operations over the years. Arguably, *nix OS' have been able to move past these problems since it appears that package management is much more of a first party concern. "Oh, you want to upgrade this utility? Well, I rely on these other things so you should update them too."
I'm a webdev so a lot of this is foreign to me w/ how it actually all functions but it's cool to read about these problems and speculate how/why they came to be and why they may only exist on certain platforms.
Are people widely choosing computer science over software engineering and then being surprised they aren't studying software engineering or is it rare for institutions to offer software engineering as an option?
https://www.newcastle.edu.au/degrees/bachelor-of-software-en...
https://www.newcastle.edu.au/degrees/bachelor-of-computer-sc...
Looks like there's even a "computer systems engineering" as well.
https://www.newcastle.edu.au/degrees/bachelor-of-computer-sy...
Some Unis also offer different tracks for Masters degrees - MA (Master of Arts) in Software Engineering, or MSc (Master of Sciences) in Computer Science. Those differences are a bit more meaningful. The MA will be softer, with less research focus, and is (ostensibly) explicitly designed as preparation for industry.
Students can pick their final subject; there is no difference in subjects between the two programmes, though we obviously expect some difference in result. After all, you should show some proficiency in both masters if you're doing the 2yr programme.
I'm currently supervising a student (SE) on file recovery, to whom I'll definitely forward the link. It's right up his project's alley.
Side remark: had he been enrolled in the 2yr programme, the subject would have been the same. I'd just expect him to extend his current work with a research question towards information science.
It was interesting seeing what the EE's and ME's were learning about, standardized practices, vs what CS was taught, the "figure out a system" method.
They are punished in a way. It hurts a school to get a bad reputation among local companies for poorly-prepared grads. R1 research schools, particularly if they are public, may feel it less. But even then, the ones I've dealt with have always chased all the fads in trying to improve learning and metrics. A tenured professor who's just there for the research may choose to ignore it all of course.
I was CS, but had a lot of SE friends, and we all agree that the best SE course was the one that was mandatory for all computing majors. It covered version control and branching, testing, design process, abstraction design, etc. That class was probably one of the most daily applicable things I learned, I took it in my third semester. If my university offered more courses of that caliber, that didn't rehash and force memorization of buzzwords, then that program would've been incredible.
In that context I wonder if you could cut from rewritable CD-RWs as well back in the day (can't remember) - that seems like another abstraction that's similarly slow in reality.
AFAIR when you "copied" into CD-RW the files would show up semi-transparent (pending) and you'd have to click a button to process with burning. Probably same for cutting I guess.
I think you were supposed to write a pointer at the original place, so the driver would know to look again, but outside in. Video CD players often didn't support that pointer.
But actually, Windows was quite a latecomer on that feature. It's just that its UX was horrible so people would never know what option they choose, on other OSes (and other Windows software) people had to actually decide to use it.
https://en.wikipedia.org/wiki/Live_File_System
which is apparently the same thing as UDF.
It'd be interesting to now what lead to the architectural decision of wear to cut that abstraction. It feels like a Microsoft team balkanization issue (the explorer team and POs being different from the NT kernel ones, I assume?), but it's near impossible to know without inside knowledge.
I just want to create, read, update and delete some resources :)
"Why can I see the files in explorer and not this program I run to use those files?" isn't an answer you have to burn support time on because it's not ultimately implementing an FS layer in special GUI components.
It's not that I don't like their FS abstraction (there's a lot of abstractions I don't like; as an engineer that's not an interesting topic for discussion but instead basically expected). It's that they broke their own abstraction a decade later by implementing the multiplextion in a layer totally contrary to what the user expects. Like if FUSE was implemented entirely in Gnome components and anything that used open(2) broke as not seeing the veneer files but you could see them in every system application.
Oh, you mean exactly like Gnome GVfs?
Of course, as you point out the right abstraction in many cases is an abstract filesystem, and for whatever reason pluggable filesystems aren't really a Windows Thing (though I think they may have been added recently for features like OneDrive and lazy git checkouts). The existence of shell namespace extensions probably stopped that valuable feature from getting put in...
E.g. at one point (long time ago), MySQL's C client library would call read() for 4 bytes to read a length indicator, and then read exactly the number of bytes indicated by the protocol, which led to a lot of time spent context-switching vs. reading into a larger user-space buffer and copying out of that.
On Linux simple, plain "strace" remains one of my favourite first steps to find performance problems in part because these things are near immediately apparent if you strace a process because the really pathological cases often ends up dominating the output so much that you'll spot them right away.
Another "favourite" issue that shows up often when you use strace like this is e.g. excessive include paths - running strace on MRI Ruby with rubygems enables and lots of gems pulled in is a good way of seeing that in action - like this zip problem it's an example that seems totally reasonable when the number of gems is small, and that first becomes apparent when you test with lots of gems, and look at what's actually happening under the hood.
A similar example of testing with too small datasets and/or on a large enough machine to not spot what becomes immediately obvious if you trace execution was how a type 1 font reference library made available by Adobe (no idea if they wrote it or if it was from another source originally) which would exhibit another typical pathological behaviour and call malloc() hundreds of times on loading a font to allocate 4 bytes at the time instead of allocating larger buffers. (strace won't catch malloc(), but ltrace does)) To avoid messing too much with the code, we replaced the calls to malloc() with calls to a simple arena allocator and load speed shot through the roof and memory usage dropped massively.
Too few people trace execution of their code. I know that because if most developers did, issues like the above would get caught much sooner, and would be rarer in published code.
When you add e.g. Python's standard library as a ZIP to the Python search path, the thing will open/read/read/read/close that file approximately a gazillion times on startup. That's OK on Linux/Unix, where that is fairly cheap. Guess which OS doesn't like that pattern at all?
https://www.johndcook.com/blog/2009/04/06/numbers-are-a-leak...
and yet, when people start adding unicode text under this representation, and the program fails to work properly. Or when counting characters, and assume the glyph count would match the char array length.
or any number of other string issues.
>And you can’t drive as fast when it’s raining, even though your car has windshield wipers and headlights and a roof and a heater, all of which protect you from caring about the fact that it’s raining (they abstract away the weather), but lo, you have to worry about hydroplaning (or aquaplaning in England) and sometimes the rain is so strong you can’t see very far ahead so you go slower in the rain, because the weather can never be completely abstracted away, because of the law of leaky abstractions.
> leaky abstractions > > hydroplaning >
hydroplaning requires:
- water (“leaky”)
- sliding (“...traction”)
- lack of _brakes_ (“abs...”)
Well, and speed too.
Right... that's why they're called "leaky abstractions". Because when you don't pay attention to the implementation details and their constraints, things break.
If it weren't leaky, you'd call it a "specification" instead of an "abstraction".
Iterating works. SQL works. Windshield wipers work. The iteration abstraction isn't leaking - the hardware and memory architecture isn't part of the abstraction. The SQL abstraction isn't leaking - the performance isn't part of the abstraction. The windshield wipers abstraction isn't leaking - the wipers were designed for rain, not a hurricane. Abstractions ignore details - that's the point. Reality doesn't cease to exist.
If your car isn't fast enough for the race, the car, steering wheel, accelerator pedal are not a leaky abstraction. You just don't have a fast enough car! If you have to worry about stuff like shifting gears, then an automatic transmission isn't the abstraction for you. The abstraction isn't leaking, you just chose the wrong one.
Abstractions are great.
I've never had to worry about CPU instructions or assembly code in the code I've written - the programming language abstraction is perfect as far as I'm concerned. Somebody handles it. I've never had to worry about the electricity powering my machines in the Cloud etc where I deploy. Perfect abstraction. Somebody handles it. My web browser chugs through hundreds of megabytes of JavaScript daily. Perfect abstraction. I've never needed to concern myself with the implementation of say V8. I've had my share of SQL performance investigations, but most of the SQL I've written didn't need any performance tweaking (there's a different abstraction for performance that works great - indexes).
Leaking is a matter of perspective.
Likewise, in an enterprise setting a (Windows) user shouldn't have to worry when saving a file to X:\ what kind of drive/storage X is. And 99.9% of the time, that abstraction works. But then someone tries to save a 500GB video and suddenly it does matter what kind of drive X is, where it is, how it's connected, the filesystem it's using etc.
I think the issue in both cases is that the distinction is fairly obvious if you're in the know but not necessarily for anyone else (until they try it).
I also have a gut feeling that doing similar operations to a connected android phone (e.g., moving photos from your phone to your PC over USB) is also slow, probably for similar reasons.
MTP is terrible[1].
[1]: https://en.wikipedia.org/wiki/Media_Transfer_Protocol#Perfor...
It's sad that with cloud being the solution for everything those days, this will probably never be improved within next decade.
I've had multiple experiences of issues on the device side causing the entire explorer.exe process to crash!
Explorers handling of the MTP protocol is not resilient to badly behaved devices, would not at all be surprised if there are security implications where a badly behaved MTP device can get an RCE in explorer.
These determinations can be made with few tests each. Start with some reasonable amount of data, keep doubling it until the test fails. Continue binary-searching between last success and last failure.
The results may come out surprising. Performance does not scale linearly - just because your program takes 1 second for 1 unit of data, and 2 seconds for 2 units of data, doesn't mean it'll take 20 seconds for 20 units of data. It may as well take an hour. Picking a threshold far above what's acceptable, and continuously monitoring how much load is required to reach it, will quickly identify when something is working much slower than it should.
As to slow operations... that is more likely because of synchronous implementation.
The popular, naive implementation is, as above, to repeat same simple operation over and over again: read from source, write to destination, read from source, write to destination.
A better implementation (what I would do) would be to pipeline operations. Basic pipeline would have three components, each streaming data to and/or from buffer
1: Read from source to pipeline buffer
2: Read from pipeline buffer to write to destination, write information about files to delete to another pipeline buffer
3: Read files to delete from pipeline buffer and execute deletions.
Using *nix shell you could do something like that in single line
1: tar -c files to output
2: pipe output to tar -xv, write files to destination producing list of written files, pipe written files to output
3: read piped list of written files and remove them from input dir
Now, this is not perfect because we are wasting performance on creating tar file when we immediately discard it, but you get the picture.
Litte bit of a rant, but I see this SO MUCH in database layers in applications. Implement an operation for one row, slap a for loop around it, it works kinda quickly on the small test data set... and then prod has an intimate conversation with a brick wall. It has been 0 days at work since that happened.
> As to slow operations... that is more likely because of synchronous implementation.
Interestingly, I think the answer is a solid maybe and depends on the storage and how you issue your i-o operations. A flash storage will increase performance if you increase parallel operations, up to a point. However - and this code apparently was written 20 years ago - on spinning drives, parallel io-operations slow you down if the OS does not merge those operations. So it's entirely not obvious.
There are obviously improvements you could do. For some filesystem you can just remove entire folders rather than remove the files individually just to remove parent folder.
Reminds me of the kind of patterns functional programming languages introduce, where you process data by describing operations on individual items and assembling them into "a stream". I'm always wary of those - without a good implementation and some heavy magic at the language level, they tend to become the kind of context-switching performance disaster you describe.
I am personally of the opinion that, to be a good developer, you have to have mental model of what happens beneath. If you are programming in a high level language it is easy to try forget about the fact that your program runs on real hardware.
I know, because I work mostly on Java projects and trying to talk to Java developers about real hardware is useless.
I have an interview question where I ask "what prevents one process from dereferencing a pointer written out by another process on the same machine" and I get all sorts of funny answers and only 5-10% candidates even have beginning of understanding what is going on. Most don't know what virtual memory is or are surprised that two processes can resolve different values under same pointer.
Yikes. How many of them have a degree in computer science?
What happens is they know what virtual memory is but can't connect the concepts. It is knowledge without understanding.
But that was incredibly hard, and error prone (and comes with tonnes of limitations). The fact that today, it's very easily possible to write working programs without knowing any of the underlying details, is a marvel.
If i were to hire a truck driver, i wouldn't expect to have to ask him to understand how the truck itself worked (e.g., fuel injection when he presses the accelerator). He only needs to know how to operate the truck from the interface (steering wheel!). Why isn't this the same for a java programmer?
No, it wasn't. If it was incredible anything it was tedious. But nobody really expected anything super complex from you. Just look at the kinds of programs that were produced in 80s or 90s.
It would be incredibly hard today. Machines got very complex, operating systems got very complex.
But it wasn't so complex in 90s. I think I stopped writing assembly when I started using WINAPI because that was the point when assembly stopped being practical.
Sometimes I wonder how I would fare if I was 18yo again today and had to start in development and learn everything from scratch. I feel being able to learn everything as the technologies were evolving is a huge advantage I enjoy.
I pick Rust every time.
I have this concept of easy problems and hard problems. Every decision to choose technology is a compromise and comes with its own problems. It is your job to know whether these are easy or hard problems.
Compilation time is easy problem. Just put more hardware to it or modularize your application or schedule your coffee breaks correctly.
Building reliable abstractions to prevent hard to debug problems is hard problem. Building a large, complex, reliable application in ANSI C is hard problem.
I try imagine, what would you rather spend your time on: coffee breaks or debugging complex bugs?
This is not that easy as it seems.
Rust compilation is not embarrassingly parallel as a lot of what's going on seems to be deferred at the link stage.
Modules inside a crate can't be compiled in parallel because of some reason I don't truly understand, please educate me.
Crates can be truly compiled in parallel, crating projects with hundreds of crates is quite a pain for other reasons.
The result is that a change single line of code in the project I'm working in can take more than 2 minutes on my last gen MacBook pro.
I have a M1 Mac mini where the compilation is twice as fast but I don't have enough RAM there and I'll have a hard time to convince my manager to make an exemption to my companies laptop refresher policies "because of rust".
If I have to take a coffee break every time I need to wait for an incremental compilation my heart would explode.
My current solution is:
1. Multitask. Do somethings else while I'm waiting.
2. Split the code that I'm iterating on in a new project. With minimal dependencies and copy the code back in the main project when I'm done (or just import it as a crate if it makes sense)
Both options suck and I wish they weren't necessary. Please let me know what else could I do. What hardware should I buy etc
(I already use zld on Mac and use minimal debug symbols, iirc debug=1)
In practice we only batch when it starts hurting. It doesn't hurt to delete files one by one on a normal file system. It's made for that. So the API wasn't "many by default" and that's how it works for zip files as well.
The same could be true here where you're moving from a zip file to probably the same filesystem the zip files is in; if removing a file from the zip file is actually an in-place move data then truncate. The problem, of course, is that removing a file from the zip file is tremendously expensive. Reading the file with one syscall per byte doesn't help (especially post-Spectre workarounds that make syscalls more expensive).
But ideally I’d want the system to delete at the end if possible, and to otherwise delete as needed, instead of either doing only all at end or only after every single file.
Wait, does that mean that a file which is larger than half of the partition in fat32 cannot be moved? Or not even be renamed?
Dave's Garage - Secret History of Windows ZIPFolders
What I can’t remember is what they stood for - ‘w’ was a word, and I think ‘h’ a half-word. But was a word 16 bits when I was writing that code?
Thanks for the PTSD.
That fact that the issue can be easily resolved without changing the abstraction (the plugin system) is proof that the abstraction is not leaky.
A leaky abstraction would leak implementation details from the plugin to the caller and thus make it impossible to replace or modify the plugin without modifying the caller's logic. It doesn't appear to be the case here. It seems like the issue can be resolved just by changing the plugin's logic; the caller's logic would stay the same. This sounds like an excellent abstraction and could be the reason why the core logic hasn't had to be changed in 1998... It should be hailed as an achievement, not shunned as 'leaky'.
There are other places where the abstraction leaks as well (try to add a file with a Unicode name to a Compressed Folder).
It's true that you could add more and more code to try to make the abstraction leak less.
Is it the abstraction which is forcing the files to be copied 1 byte at a time? It doesn't seem like it. It's an implementation issue. The interface allowed the plugin to have been implemented in many other ways, the implementer of the plugin just happened to go about it the wrong way.
For example, the zip folder could realise that deletion is an expensive task, and defer the job till later. It would tell the kernel that the raw bytes of the zip file are only available to other applications after this operation is complete (so that emailing someone the zip file can't represent the pre-deferred-operations state).
The tree of deferred operations can then be optimized - for example, deleting multiple files in a zip file could be combined into one. Deleting stuff in a zip file, and then deleting the whole file can likewise be combined.
Kinda like peephole optimisations for stuff your OS does.
System services is a virus.
When I saw this timestamp sometime ago on my PC I thought it was a joke?! C'mon 1998 like WTF!