OpenBSD's safe, new file(1) implementation
marc.info
marc.info
You can also find a git mirror of the referenced commit at: https://bitbucket.org/braindamaged/openbsd-src/commits/c421d...
I'm not so familiar with FreeBSD coding standards and design practices but it looks like that could still be refactored even more by getting rid of all the extra memory allocations (and deallocations), making future changes less likely to introduce bugs. For example it's allocating one structure for each file given as input, and making copies of all the paths... which I'm pretty sure isn't necessary as file tests each of its input files independently; the code misleadingly implies, however, that they are somehow dependent. The old version was better in this respect.
That said, compared to the much more complex old version, the rest of it is an improvement.
This[0] is a good place to start. Other than than, just read the code.
[0] http://www.openbsd.org/cgi-bin/man.cgi/OpenBSD-current/man9/...
That's a syntactic style guide; I was referring to higher-level issues at the design level, like "avoid dynamic allocation unless absolutely necessary", which are between syntax and the general very-high-level "follow the UNIX philosophy" type of guidelines.
To answer the sibling comments about preparing a patch and sending it in: this was just an observation I had from a brief glance at the code; I don't have the resources to delve into the research and find out if this is actually something they would want to change, if there is a reason behind doing it this way, or only the norm for the project (e.g. the Java community largely has a very different notion of what it means to "simplify" code) - so if someone wants to take these ideas and follow through, they are more than welcome to.
https://github.com/OpenGrok/OpenGrok/wiki/OpenGrok-installat...
EDIT: Not all of them are working, though.
This new implementation was created from scratch, carefully, with modern coding practices by a very proficient programmer.
Cool. What happens 10 years from now, when we've wanted all the missing features back, and when writing C code (let alone C code without tests) to run on untrusted input is no longer "modern coding practices" but firmly legacy?
This seems like a great way to have a version of file(1) that works great in 2015, on the subset of filetypes that are in-scope for the initial version, on this version of OpenBSD, but I'd like to know how it becomes sustainable.
But it's sort of like living in a world where modern medicine has been discovered, and celebrating that one city has started moving away from leeching to treat fevers, while every other city is still practicing it, and that one city still uses leeching to treat many other diseases. Strictly speaking, it is an improvement, but....
I was going for, let's quit bloodletting, but let's also not celebrate how we're occasionally no longer bloodletting, because it's still embarrassing how far behind modern medicine we are.
Could you please elaborate on what "bloodletting" and "modern medicine" represent here?
We have languages like Rust and we have proof tools like Coq which can be leveraged to become more memory safe or more provably correct... using those tools would be like using modern medicine.
I don't think the analogy is really that great because those tools are still incredibly immature in terms of actual usability.
If you absolutely need speak in parables, do it with something that represents an earlier common practice, for example with producing houses with hammers and nails instead of modern pre-fab concrete blocks. But even that is stretching it a bit too far, the productivity difference simply isn't comparable.
Of course, some people did bloodlet for reasons that didn't help, but without knowledge of bacteria, many were just trying their best with what limited knowledge they did have.
New implementation of the file(1) utility.
This is a simplified, modernised version
with a nearly complete magic(5) parser
but omits some of the complex builtin
tests (notably ELF) and has a reduced set
of options.
ok deraadt
What is significance beyond the obvious?What options are now missing?
What are the tradeoffs? Beyond a simplified API, are there new capabilities, functionality/use-cases, and/or improved performance?
http://brynet.biz.tm/pub/file.1.pdf
It seems to have all the options I'd ever use. In fact I've never used any options to file (though I do use it to quickly identify a handful of files, often enough).
The only reason OpenBSD does this kind of thing is because they got annoyed with the vulnerabilities in the classic file/libmagic source which everyone seems to use. They probably disliked the code style / organization of the source enough that they didn't want to patch it up (which they also do fairly often for various ports). That said I've never personally looked at the source.
I'd say very few people do, which is kind of scary for a utility like file(1), which people feed any random file without giving second thought.
My point is, the degree to which you can expect "safety" from a program (whatever that word means to you) is predicated on the degree to which you can trust the competence and practices of the developers, and the degree to which your definition of "safety" applies to the program's specifications. Remember, all programming languages are Turing-complete, which means you can blow your foot off with any of them if you're not careful.
This is a simple implementation that includes a safer new magic(5) parser and excludes many of the "built-in" parsers that the contemporary file had, i.e: for executable formats, which are probably better examined with specially tools like objdump(1).
The original author shared this very fitting sentiment on the news the new file(1)...
http://marc.info/?l=openbsd-cvs&m=142989483913635&w=2
"The Albatross fell off, and sank Like lead into the sea." — Ian Darwin
The Rime of the Ancient Mariner- using language with more safety guarantees than C
- including seccomp (either bfp, or even just mode 1) by default
- in the announcement saying: tested with XXX, YYY, ZZZ; fuzzed with ABC, DEF
It would be a great opportunity to create a poster-child people could point to and say - this is the way we should be (re)writing secure software, these were the awesome tools used to help it. For OpenBSD the first two couldn't happen (part of base system so C; there's no seccomp support), the third could but didn't - at least not publicly (valgrind, clang-analyser, [am]san, afl, etc. could all be listed)
Maybe the next time something important is broken.
Basically I'm saying that we can try very very hard to write secure code in languages which invite issues, or... try to eliminate whole classes or issues at a time. Is it such a terrible idea?
You mean decades-old languages like Perl and Python for example? Technologies like the ones we can finally now afford to implement - actual syscall filtering and selective capabilities dropping? Supporting utilities which only matter at development time? How do those go away?
And I wrote in the first message - OpenBSD integrates `file`, has to use C and doesn't implement seccomp. They couldn't become the poster child. But the next `file` reimplementation probably won't hit the news anymore.
People can improve the way we write software right now. But choose a project at random and they don't care or don't use what's available for free. Some promotion would be great when everyone's looking.
Seccomp is almost unusable for most purposes with anything resembling tight filters as you do not really know what syscalls glibc especially will add to your program. Capsicum in FreeBSD is much more usable, and they are starting to priv-sep programs with it, they are good examples, I would look at them. Allegedly Linux will get Capsicum one day. OpenBSD has many examples of priv-sep such as openssh, just using seperate processes.
Linux is disadvantaged by not having a base system in teh same way - you can make a new "file" but how to know if anyone will use it. You can try to rewrite the Gnu tools better, but many are a horrible mess. You probably have to start a new Linux distro...
OpenBSD's new, untested file(1) implementation.
The only thing "safe" here is that the author was completely unharmed by the unit tests that he failed to write. Where I work 2500 lines of new C code without a single line of tests would be laughed off the code review system and then probably revisited around performance review time.
Heroic programming is how we got to where we are today, running the world on gigantic piles of untested gotos and question mallocs and frees. There are no real heroes in programming, only people who haven't yet figured out how to write tests.
And the OpenBSD people are very aware of fuzzing, this new implementation of file is a direct reaction to Michal Zalewski findings...
I sincerely hope the lack of unit tests is because it's covered by existing tests.
http://lcamtuf.blogspot.no/2014/11/pulling-jpegs-out-of-thin...
Apparently, Ian no longer considers himself a maintainer, devoting such title to Christos Zoulas of NetBSD and Two Sigma.
1) Identify a harmful input through fuzzing 2) Reduce the input to minimal testcase 3) Contribute testcase and fix to existing code
Note the absence of "rewrite the whole thing without tests" in this process. From-scratch rewrites that may or may not have fewer bugs than their predecessors are known as CADT.
http://www.jwz.org/doc/cadt.html
"This is, I think, the most common way for my bug reports to open source software projects to ever become closed. I report bugs; they go unread for a year, sometimes two; and then (surprise!) that module is rewritten from scratch -- and the new maintainer can't be bothered to check whether his new version has actually solved any of the known problems that existed in the previous version."
The OP rightly points that there are no known tests that demonstrate the known problems in the older implementations and also no proofs that the new implementation passes them.
They are very different. Especially one being a core utility another a Desktop Environment.
I agree with that critique of the Gnome development process, but the OpenBSD development process prides itself on security and that's a first commit, new changes are being added (as it can be seen by the log of that file - http://cvsweb.openbsd.org/cgi-bin/cvsweb/src/usr.bin/file/fi... )
I'm sure I'll manage to write a secure file(1) implementation, if it's not required to produce correct output.
Here is the relevant directory for file: http://cvsweb.openbsd.org/cgi-bin/cvsweb/src/regress/usr.bin...
The files in that directory sure look like tests to me. Am I missing something obvious?
It does sound CADT-y.
http://www.jwz.org/doc/cadt.html
"It hardly seems worth even having a bug system if the frequency of from-scratch rewrites always outstrips the pace of bug fixing. Why not be honest and resign yourself to the fact that version 0.8 is followed by version 0.8, which is then followed by version 0.8?
But that's what happens when there is no incentive for people to do the parts of programming that aren't fun. Fixing bugs isn't fun; going through the bug list isn't fun; but rewriting everything from scratch is fun (because "this time it will be done right", ha ha) and so that's what happens, over and over again."
Specifically, why didn't somebody just remove the ELF part from the existing code, if the goal was "reducing the attack surface by removing the features?" (which can be a reasonable strategy for increasing the security).
Instead there most recent test is 6 years old
I don't feel that this detail is relevant. Because file is a basic utility, and its function is not supposed to change much over time, a test written 6 years ago for file should still be relevant to file today.
Specifically, why didn't somebody just remove the ELF part from the existing code
I think this was not done because much of the file code followed outdated coding conventions, and would have to be replaced anyways. Take a look at this diff comparison of file.h: http://cvsweb.openbsd.org/cgi-bin/cvsweb/src/usr.bin/file/fi...
So no, the 6-year tests weren't sufficient. And yes, the updated tests and fixed old implementation is more important step than ""this time it will be done right", ha ha"
Additionally, replacing the "defined" constants from the file.h with the enum in another .h is hardly a reason for the rewrite.
However, I don't understand why you suggest that fixing the old implementation to adhere to new programming standards would be better than re-writing the program from scratch, because most of the program would have to be re-written anyways, and so the developer would barely save any time and effort by fixing rather than re-writing. Also, if more thorough tests were to be included, then either fixing the program or re-writing it would have the same impact security-wise, as the improved tests would account for any discrepancy in either route.
IMO, in this case, whether to fix the old software or to re-write a new version were both acceptable options.
I suggest extending the tests, fixing the problems discovered in the existing implementation and then doing something completely new, somewhere else.
Completely unfounded assumption.
> the author was completely unharmed by the unit tests that he failed to write
That's always valid, having written unit tests or not, and a lot of unit test lovers write thousands of irrelevant unit tests that let bugs slide
> Where I work 2500 lines of new C code without a single line of tests would be laughed off the code review system
A lot of software has been written without unit tests which doesn't mean they don't work (as opposed to what detractors might say). In fact there are a lot of them you're using right now.
> There are no real heroes in programming, only people who haven't yet figured out how to write tests.
Yes, sure, "Unit tests are the only true way", not to mention unit tests in C are much LESS frequently possible, especially in low level stuff.
I just find risible how people accepted this unit test fetichism without questioning. Uncle Bob fans always makes me laugh
The work he has been doing with AFL and googles big fuzz farm, focusing on utilities that are used daily without thought is insanely important, imho
Anyone knows how much impact this change can bring? How much systems rely on file?
Personally, every time I use it was as one-time script.
Despite using old tools, the OpenBSD guys release their updates every 6 months like clockwork, and have for the past two decades. I know of no other FOSS project with that level of project management.
CVS is frankly hideous. Saying "git is way too complicated" is another way of saying "CVS is way underpowered". However, a good reason to cling to this relic of a bygone age is that, as far as I know, they have everything in CVS, and it must be a lot more convenient to be able to checkout an arbitrary part of their dev tree than messing around with git submodules.
That said, if their attitude is really "if you complain about CVS, you are not worthy", that is bound to turn off many people, and for good reason.
Perhaps their goal isn't reaching those people.
Side note: I'd be really interested to see an OpenBSD designed VCS. I'd be very curious what the design would look like.
That's what I am saying. However, there is a big difference between "we use CVS for good reasons" and "we are not aware of the limitations of CVS compared to modern DVCS". That's the difference between "sane" engineering conservatism and a "get off my lawn" attitude.
I don't know. What's wrong with it? I've used it for decades now and it's never really been a problem. Sure, I like git, but the idea that CVS is some kind of show-stopper just never resonated with me.