It is good to know that your base64 encoding function is tested for all corner cases, but integration and behaviour tests for the external interface/API are more important than exercising an internal implementation detail.
What is the external interface of the kernel? What is its surface area? A kernel is so central and massive the only way to test its complete contract with the user (space) is... just to run stuff and see if it breaks.
TDD has some good ideas but for a while it had turned in a religion. While tests are great to have, a good and underrated integration testing system is just for someone to run your software. If no one complains, either no one is using it, or the software is doing its work. Do you really need tests for the read(2) syscall when Linux is running on a billion devices, and that syscall is called some 10^12 times per second globally?
Or it crashed and they couldn't be bothered to report it. Looks the same as no one used it but has a different cause.
You're underrating the fact that Linux's massive and diverse user base is a great filter for bugs. Only the weirdest heisenbugs can survive this filter, and because of their nature, they wouldn't be tested anyway.
I always assumed the fact I can reliably break Linux on Laptop A with Docking Station B by doing Action C was pointless to report, because how is anyone supposed to reproduce it if they don't own Laptop A and Docking Station B?
There's a good chance it's a firmware or hardware bug Linux cannot fix, but it's worth a shot.
But to answer your question, Laptop A and Docking Station B probably interface with each other with the same two chips/subsystems present in other devices. And if they work with Linux, the maintainers of these drivers, often the manufacturers, are publicly listed and are the ones that will try and troubleshoot it with you.
I'd say that's very common and not exclusive to some specific project.
I made a change of about 30 lines of code recently, and it had 140 lines of tests (and I, by far, do not cover every situation).
I think I've rarely encountered well tested pieces of code that were not way smaller than the tests that came with them.
And unit testing is relatively useless, what would be useful are end to end tests that start with userspace API but that's much harder task, altho hopefully those tests wouldn't need to be changed that often.
i.e. does not depend on external timing or fickle hardware.
Very, very difficult to unit test someting that runs using hardware interupts.
There's also the Linux testing project, which is technically third party. It's not clear to me how extensive it is but for a project as important as Linux I think it has to be graded as "needs improvement."
But testing costs a lot more when there is hardware present. It's a whole new dimension in your testing matrix. You can't run on cloud provider CI. Hardware has failures, someone needs to be there to reboot the systems and swap out broken ones (and thing break when they run hot 24/7). And you need some way to test whether the hardware did what it is supposed to do.
Although some kernel driver developers have testing clusters with real hardware (like GPU drivers devs), there is no effin way Linux could be effectively tested without someone paying a lot of money to set up real hardware and people to keep it running.
Of course the hardware can be simulated but simulation and emulation are slow (and potentially expensive). At work we run tests on simulation for next gen hardware which doesn't exist yet. It is about 100,000x slower than the real deal, so it's obviously not a solution that scales up.
Testing purely software products without a hardware dimension is so much simpler and cheaper.
Just crowdsource this testing: I am sure that there exist some people who own the piece of hardware and are willing to run a test script, say, every
* night
* week
* new testing version of the kernel
(depending on your level of passion for the hardware and/or Linux). I do believe that there do exist a lot of people who would join such a crowdsourcing effort if the necessary infrastructure existed (i.e. it is very easy to run and submit the results).
And having any kind of manual intervention required will almost certainly reduce the reliability of the testing.
This is further complicated by the need to reboot with a different kernel image. Qemu and virtual machines can't do all kinds of hw testing needed.
And in fact, the kernel is already tested like this. Just very irregularly and sporadically. The end users will do the field testing and it is surprisingly effective in finding bugs.
> And having any kind of manual intervention required will almost certainly reduce the reliability of the testing.
Perhaps I am underestimating the necessary effort, but the willingness and capability problem can in my opinion be solved by sufficiently streamlining and documenting the processes of running the test procedure.
If the testing procedure cannot be successfully run by a "somewhat experienced Linux nerd", this should be considered a usability bug of the testing procedure (and thus be fixed).
Dogfooding of builds is also testing, just a different kind.
Hardware testing takes up a pretty sizable portion of the development cycle. Why do we think we are special?
SQLite Test Harness #3 (hereafter "TH3") is one of three test harnesses used for testing SQLite. TH3 meets the following objectives:
- TH3 is able to run on embedded platforms that lack the support infrastructure of workstations.
- TH3 tests SQLite in an as-deployed configuration using only published and documented interfaces. In other words, TH3 tests the compiled object code, not the source code, thus verifying that no problems were introduced by compiler bugs. "Test what you fly and fly what you test."
- TH3 checks SQLite's response to out-of-memory errors, disk I/O errors, and power loss during transaction commit.
- TH3 exercises SQLite in a variety of run-time configurations (UTF8 vs UTF16, different pages sizes, varying journal modes, etc.)
- TH3 achieves 100% branch test coverage (and 100% MC/DC) over the SQLite core. (Test coverage of extensions such as FTS and RTREE is less than 100%).
That's not something you see called out very often. Correct code fed to a broken compiler can definitely give you a broken binary. Likewise a correct binary will pass on a correct simulator and may fail on broken hardware.
I take that as an argument not for or against tests but against making it easy for management to turn your project into a feature factory.
That weirdness is precisely the sort of thing you want automated tests for.
It's more useful to categorize tests by how hard to setup an environment. In reality cost/usefulness line lies on this boundary.
* in single process * multiple process using IPC * multiple processes using network * tests validate functions calling foreign services
I see many devs (including myself) call "integration" tests as "unit" tests because in your particular app/system they are easy to spawn locally even without any container.
You are free to write tests that compose packages, and I’ll do it for my own sanity and confidence. But that’s in addition to, not instead of, the minimum “not horrifically irresponsible engineering” standard of universal, exhaustive, very detail oriented mock driven unit tests within all packages.
It breaks my heart when I see a manager with some engineering background, whose idea of better is to make the code coverage number higher.
People here have great arguments against tests. I agree that they’re not always useful. Aiming for 100% code coverage is a waste of time.
But if a manager wants easily measurable metrics to hit and they have the budget and bandwidth for it, why should I complain?
Writing tests is easy, predictable work. My rate is my rate and if this is how management dictates I spend my time, I’ll take an easy couple weeks and paychecks any day.
But is it useful work? This is the problem: engineers have at some point decided that keeping busy writing tests is better than finding other ways to write correct code.
Now that TDD has died down as a religion the current fad became strong typing. I will repeat it again and again, I have yet to see an error caused by me passing an int where an array-type was expected in Python or JavaScript. I have been doing this for a long time. This too shall pass and we will see posts like “Show HN: new lightweight dialect of XYZ without the burden of types”.
You will discover that you reversed the order for parameters to a function by running your code with much less effort than annotating all your code. This type annotation is entirely useless:
def get_user(email: str, create_if_not_exists: bool)If you don't have tests, all that means is you or somebody else is testing it manually.
Ah, the Sufficiently Smart IDE, companion to the Sufficiently Smart Compiler. Sure, go ahead and wait for the magic autocomplete algorithm that can deduce precise function signatures in a large Python or JavaScript code base.
Meanwhile, some of us have work to do, and we'll use the best tools at our disposal for managing non-trivial code written by many people across teams. And those tools include modern static type systems.
Python was my first programming language(cliche, I know) and I didn't understand the whole "static typing is good" thing until I learnt languages like C, C++, C#, and rust.
Python also introduced type hints and made them a valid part of syntax for good reason. You'd typically end up writing a docstring with the return type of a function and the types of the parameters anyway, which is what the person replying to you was presumably referring to.
It's not about passing in an int instead of an array, but rather but being able to figure out what a function wants without needing a stackoverflow thread or having to search through tons of documentation for something that would be otherwise trivial in statically typed languages where the function declaration tells you quite a bit about a function.
Hence it's safe to say that types documented in code in some way are very useful for humans, and if you have issues with statically typed programming languages, there's a chance you were not documenting your code enough already.
We feel it "funny" because that culture is perpendicular to the Startup culture that abounds in this forum. But number-wise, it's a minority.
The great majority of people are not wholly in love with their craft. They just want to do their job to get paid. They dont read about Rust on weekends, and dont program yet-another-X in their spare time.
And that Ok. Thats what the majority of the world does.
I would rather my employees (and contractors!) work on what's actually valuable, and not just take blind instruction. The people closest to the code should be most empowered to improve it. Is that not a common management view?
"write as many tests as possible, 100% code coverage, in fact 100% branching code coverage. Use an absurd amount of mocks to achieve this"
to
"write tests, not too many. mostly integration". (Guillermo Rauch tweet)
And thank god because 100% code coverage always felt like an exercise in obedience to dumb process over good judgement.
Because otherwise, the more tests you have against implementation details, the more times your tests will break just because you've moved some code around. Tests need to go red when there is a bug, not because you've renamed a couple of private methods and now the mocks are broken.
Code coverage? Use it to determine what is dead code you can remove: if your tests don't go through some code, that code is useless.
How do unit tests stay green when you rewrite in another language? Are your tests an external API you call?
On the other hand, if you do that with application code, you're essentially submerging your code in a tar pit. Any change in behavior will require changing dozens of tests. Fixing broken tests will become so routine that they stop telling you anything about your code. Development will become "I changed from foo to bar, now let's update the 73 assertions that failed to whatever value the test runner says they have now".
I think I can agree with the sentiment that mock heavy tests should be short lived, and why you don’t want to keep those around. Not sure I can agree with the rest.
I aim for max of two stubs per test, and many of those end up with a TODO. Pure functions need no stubs, and most well factored code should only need at most one. One stub one test is pretty sustainable. It’s easy to rewrite such a test if the requirements change. Unit tests should absolutely be disposable, but you don’t dispose of them all at the same time. Just the ones that don’t fit the new rules.
I do run into a steady stream of people who can’t seem to understand that the tests should affect the structure of your code. “And then a miracle happens” is what you have there - long intervals where the system produces no verifiable state or output is bad. That’s not an architecture. It’s lack of it. A functional programming style makes this easier to avoid, but it’s not a cure, because the disease is in their heads, the code is the symptom.
There’s a substantial overlap between people with untestable code and people with undocumentable code. They can’t explain the code to the test framework any better than they can explain it to each other.
All that said, testing is hard. It shouldn’t be this hard and we need to keep looking for ways to improve that situation. But even here we have people who reach for the least expressive solution quite frequently, such as assert over more reflective matchers, which make for much more useful red tests. At least BDD style seems to be winning out.
People with a lot of clout have absorbed the virtues of automated testing in general and applied it to unit testing in particular. It’s hard to swim upstream on that one.
Just want to point out that advice is inspired by Michael Pollan's diet advice: "Eat food. Not too much. Mostly plants."
In my eyes this still hold true. Software that has expected behavior should be tested to make sure that the behavior isn't broken due to changes to any of the involved hot code paths or data formats.
I once wrote a code library for some boring business system that handled integration between the system and a JWT library, which would make sure that certain requests should be serviced. Using a library without tests wouldn't be acceptable (security related/adjacent code is perhaps the best example of such circumstances). Neither would my code not having tests, either, given that this library would be used across multiple services within the system.
Thus, I wrote tests until I got pretty close to 100% coverage and doing that actually helped me discover a few bugs while I was actually writing the tests! Not only that, but once the need to refactor something arose due to changing requirements, the tests breaking told me exactly what I had overlooked while doing those changes. Not only that, but if I'm long gone and someone comes to make changes to the library, the CI will tell them about the things they might overlook themselves, aside from any boring Wiki that they wouldn't read or other docs. The tests also demonstrate all of the ways how the code can actually be used, so aside from the occasional code comments, they also serve as living documentation.
There absolutely are cases where testing something won't be viable (e.g. different file systems the code implementations for which depend on the runtime that's installed on the system, whereas all you get is a leaky abstraction in front of these and your test setup doesn't contain every covered platform, for example, checking which file paths are parsed as valid and which aren't across different file systems on different platforms), but in most systems they're not the majority.
> Use an absurd amount of mocks to achieve this
You also hit the nail on the head here - this is a problem and a symptom of us perhaps developing all of our systems wrong. The main reason for not writing tests (one that I can understand) is the fact that it's not easy to do so. You end up with various mocking frameworks and libraries that try to take away some of the pain caused by the fact that your entire system is not testable, but end up with more complexity to dance around in the end.
I think the only way around this is to do data driven design that's coupled with functional programming in ample amounts, with as many pure functions as you can get. This would be completely un-idiomatic for many of the languages out there (e.g. those that rely on injecting services/repositories/whatever in fields, instead of passing everything a function needs in the parameters), but is also the only way how you could make testing easier. Maybe passing interfaces to "services" (many seem to use the service) pattern would be wrong and instead you'd need to pass in separate methods that your code will use. So instead of passing in UserService you'd pass in UserService::getUserById.
So in a sense, it's a struggle to find the balance between code that is absolutely untestable without being in mock hell, to the point where you test mocks and your tests are useless and ending up having to write code that goes fully against how things are done in any given language and the frameworks you'll use within it, probably ending up with more code meant for decoupling those parts than you have the time to maintain.
> write tests, not too many. mostly integration
In an imperfect world, I guess we can just pretend that this is okay, because it will give you the most results, compared to the amount of work you need to put in. At the end of the day, nobody wants to pay 10x more for systems that are nearly perfectly tested, they just want something that is vaguely decent and will accept hand-wavy apologies for everything constantly breaking, as long as the breakages are small and non-critical enough. Devs also don't seem to typically enjoy writing tests, in part due to some systems not being testable easily, but also because of many tools, in particular mock frameworks and even integration testing tools (like Selenium, which thankfully has more and more alternatives), just being unpleasant to use.
wow. that blows my mind. is that the most frequently invoked piece of code in the world?
It would be cool to know which is the most frequently called piece of code in the world. Maybe something that hasn't been changed in decades.
Yes, you do, because otherwise you will find that some change or new hardware will break it.
If a change breaks it, some one will probably have called read(2) a few times before that kernel is shipped to you. A change is filtered through multiple developer and the kernel has some kind of smoke and consistency tests it is subject to. Also there are people and companies that compile their kernel from the master branch and run software on it.
No one's gonna ship Linux with a broken read syscall. And if there were tests, one could still ship a broken read syscall anyway.
Tests do not guarantee the absence of bugs.
Genuinely curious: why would a company ever do this?
Distro kernels are general purpose and include a lot of stuff you do not need.
A custom-configured kernel is much smaller and (consequently, theoretically) more secure.
A commenter below gave a good example of a company that would need to run a bleeding edge kernel (Red Hat). The vast majority of companies would never need to do that though.
This may work for most people, but not all. It's easier to get things fixed while developers still work on something, than weeks/months later. Also if you have not most-popular hardware/software configuration, testing done by others isn't necessarily sufficient for you.
And it sometimes is.
> If a change breaks it, some one will probably have called read(2) a few times before that kernel is shipped to you.
I've seen to many kernel bugs on new hardware to believe this is enough. Maybe not in read(), but still in something that should work. Throwing testing on users may have been acceptable in the past, but not anymore.
I think that this heavily depends upon the part that is tested or not tested. For certain parts in the kernel, such as the firewall, I truly believe, that such tests cases (including corner ones) shall be present.
Maybe not, but isn't it true that any new code that is being added to the kernel has been run on exactly 0 devices? And new code is being added all the time.
Maybe that's why the most recent commit as of right now is loaded with fixes of broken things: https://github.com/torvalds/linux/commit/08ad43d554bacb9769c...
It's one thing to say that it's impossible to test effectively locally before release (which I'm not sure if that's true or not). But you're saying it's just not worth testing because it'll break in real life and that's even better, which I"m not sure I can agree with.
Are you saying that if Linux had tests, there would be no bugs?
https://github.com/torvalds/linux/commit/c60c152230828825c06...
The fix is to change && to ||
This seems like the exact type of bug that a unit test could prevent.This attitude blows my mind. The point of tests is to find bugs before you ship them to the user!
Writing tests after you've already verified that some piece of code is working correctly (e.g., by executing it 10^15 or whatever times on different systems with different inputs) is useless, I give you that. But what if you want to make a non-trivial change to the code? How are you going to make sure you didn't introduce any bugs without tests?
Automated tests are not the same as TDD.
The fact that systems level programmers, generally speaking, have a disregard for testing automation is apparent. Over the years if somebody had the "pleasure" of running any type of linux distro, how many times would things break randomly for reasons that the end user could only imagine? Many of those issues probably could have been surfaced and fixed before the code shipped had there been any testing automation in place.
But let's forget linux for a minute. Look at Windows. For the longest time the quality of windows was the biggest joke in software. And those engineers working on it, sorry to say, were cut from the same cloth as the Linux kernel folks and all other systems programmers. We don't need no stinking tests. Well eventually somebody that cared about the reputation of the company forced testing upon these teams and lo and behold things have gotten way better. Apple is the same way.
To me it is honestly astounding that there is still such a strong anti-testing sentiment in the industry. When you have more than a handful of people working on a complex system I greatly prefer to know that there is a robust test suite looking at important functionality, seeing if performance degrades, checking for static analysis errors, etc. And when something does break, since I already have a testing solution in place, it is generally easy to add a test which covers the regression and ensures that it never comes back.
Having never interacted with the kernel, I have to assume it isn't just one massive file right? It's broken into separate files and components? And if you ever want to modify or refactor one of those components its nice to have the confidence that your code changes are safe without having to rebuild the whole thing and run your integration/behavior/e2e tests. Especially if you aren't the one who originally wrote the component you're modifying and don't know what the intended behavior for edge cases was.
Obviously if you're mocking everything then the tests might be pointless(there can still value in preserving in the test what you assumed the behavior was), but I feel like in some ways people have over-corrected on TDD and claiming unit tests are pointless.
Only time I've written 100% branching coverage unit tests was when I needed to implement a spec, with many alternate representation formats, and a plethora of test cases I could run through as a litmus test of conformity.
1. Are there pieces that can be tested independently where verification is valuable?
2. What is the impact of a bug or regression?
3. What type of testing is most likely to uncover issues in my product given its architecture and domain?
4. Am I confident to refactor the code base without creating new bugs?
etc. etc. Iow, testing isn't magic, it should be driven by business goals. For example, I use Rust which due its great type system makes some types of code "just work", however, Rust cannot prevent logic errors. This means when I'm writing 'algorithmic code' I tend to write heavy unit tests. When I write API driven code, I tend to use more integration/end-to-end style testing. Do what makes sense for your code and goals. Tests take time and need to be refactored, so they aren't free, but can be valuable.
The hard part in adding unittests is deciding what a unit is, and when a unit is important enough that it should have its own battery of tests. Choosing the wrong boundaries means a lot of wasted time and effort testing things that likely won't break or change so fast that put a drag on refactoring.
I disagree that a kernel can't or shouldn't be unit tested. At the very least, it has a strong interface in the userspace system calls. Of course you should unit test the system calls. Especially because Linux's motto is not to break userspace.
The other benefit from unit testing that is overlooked is that it accelerates optimization. If you have a good set of unit tests that test all observable behaviors of a system means you can optimize the logic inside and constantly rerun those tests to make sure it's not broken. This speeds up the experimentation process and clearly delineates the interface so that you can see when and how you can "cheat" on things that aren't observable.
So, hard disagree. Test the bejesus out of the kernel. It will harden it and open up the ability hyperoptimize the units inside.
edit to add: well-written (i.e. terse and readable) tests are excellent documentation on how a unit should behave.
However, I think there are parts that probably should NOT be unit tested, because unit testing isn't free and slows down change by getting very close to internals. As an example, device drivers that are tied tightly to hardware. While it would sound nice to have an nvidia GPU fake/mock, in reality it is probably not possible to create one, keep it up to date, etc. The complexity would be enormous and might not show anything as you wouldn't know if it was a bug in the fake or driver.
As such, I retain my answer of: it depends what you are trying to accomplish. Different testing strategies for different types of code and domains.
Ok, but then again: How would you make sure your Nvidia driver is working correctly?
"Deciding what a unit is" isn't the hard part. The hard part is finding the units that benefit from unit tests (and convincing other religious people about this, as I am trying to now, which in this case - not many).
The Linux system calls are "deceptively simple". Simple to test, right? How complicated can say `write(2)` be? But if you actually tried doing it, I'd be surprised if you can write reliable "unit" tests beyond writing to /dev/null.
The practical way to test system calls are to test it with real world usage. Unit tests here might catch some problems, but the vast majority of kernel bugs aren't those that can be caught with unit tests. (If they're bad enough, the bug will cause the OS to crash before it completes booting, no unit tests needed.) The more subtle bugs are often hard to reproduce, only happens in certain loads and hardware.
As to your final edit-to-add: you're joking right? Like Linux syscalls need more documentation. POSIX (I'm aware this is not exactly Linux documentation, but still) was a spec literally before Linux was started.
edit PS: https://tenor.com/view/unittest-unit-test-gif-10813141
GP is using the original definition of unit test, GGP is using the religious implementation-level definition TDD advocates have come up with.
From GGP:
> but integration and behaviour tests for the external interface/API are more important than exercising an internal implementation detail.
They call this an integration test, but by the original definition, testing the interface is a unit test. An integration test is how multiple units interact with each other. Testing the implementation details is what TDD advocates have turned unit testing into, by changing the definition of "unit" from semantics (a self-contained module that does one thing from the business perspective) to syntax (a function).
> edit PS: https://tenor.com/view/unittest-unit-test-gif-10813141
Seems like you agree with the original definition. The TDD version of this GIF would be a test for each part that makes up the handle (the implementation details instead of the interface).
Using that definition, that's just basically saying "the kernel can benefit from ... tests". Of course the Linux kernel doesn't have "TDD-religion unit tests", but there's already more than enough "tests" for the Linux kernel, and there's more than enough benchmark tests for the Linux kernel (for the optimization argument). They're just not "unit" tests in any meaningful way.
From the pedantic side of things, the original comment said "(TDD-religion) unit tests are overrated", then the reply was either "no they're not", or "no, unit tests (as originally defined) are not overrated". The latter is somewhat inconsistent with the "hard disagree", so either they were confused as to which definition they wanted, or cherry picked properties from both unit testing regimes.
This is a naïve utopian world view. The word "good" is doing all the heavy lifting for you. In the last 15 years I haven't seen a single company that had "good" unit tests. I'm inclined to believe they don't exist.
Some of the tests I saw turned out to be good. Most of them weren't.
The reality is that if you're committing to unit tests, a big chunk of your tests will be shit. And when that's the case, accelerated optimization is far from guaranteed.
I’m not sure why you’re expecting the answer “no” here. The biggest value of automated tests is non-regression. So with billions of users, yes, you absolutely need reliable and reproducible tests. Maybe not unit tests (I agree they’re overrated, you want to test contracts and behaviors, not implementation details) but a strong integration test suite is a must have. To ensure you don’t break anything your billion users might rely on.
Few unit tests in the kernel would be able to compete with so many chaos monkeys (a reference to Netflix's https://github.com/Netflix/chaosmonkey).
Okay, but Linux has internal implementations of base64, assorted hashing algorithms, some compression, probably other algorithmic tidbits; wouldn't those be easy to test?
Edit: And downthread people are saying that Linux does do that:)
Additionally the in-tree Realtek R8169 driver fails to redetect Ethernet being plugged in after I unplug it for 5-10 seconds, and I had to resort to the out-of-tree R8168 driver as a workaround.
But really, most code does not need to be tested. The parts that are likely to break when others change them should be. Testing can mitigate risks for complex, esoteric, highly dependable, domain-specific, and technical code. In such cases, tests can save a lot of time that would be spent fixing defects later. In contrast, testing as a religion is entirely counter-productive. It's a time sink in an effort whose whole purpose is to save time. And what is worse, poorly written tests make the codebase less maintainable, more complex (to pass wrongly written tests), and degrade code quality (through a false sense of security).
Functional/integration testing can suffer from the same problems. There's nothing more annoying than fixing a defect in a system only to uncover a mess in functional tests that relied on the bug. Of course, some of this is unavoidable. However, cargo cult-style testing is avoidable and entirely self-inflicted harm in many teams.
In short: testing is a tool to save time and increase code quality; such a tool should be used where it saves time and improves code quality, and it should not be used where it just results in the opposite.
So if tests don't cut it, what does the Kernel community then do instead of testing to verify correctness?
> Do you really need tests for the read(2) syscall when Linux is running on a billion devices, and that syscall is called some 10^12 times per second globally?
That's exactly the kind of code that needs tests.
What happens if someone changes some code that affects the read call and expects it to run on said billion of devices?
Another huge source of issues is due to hardware doing something unexpected/not-to-spec - another thing that unit tests would very poorly verify given that, any unit test will simply reproduce how the developer thinks some piece of hardware works, rather than what it does in real life
Production kernels are usually better served by long-running stress tests, that try to reproduce real-world use-cases but inject randomness into what they are doing. And indeed, both NT and Linux kernels are extensively tested in this fashion
It's actually causing an unexpected problem: It finds so many issues that kernel devs have troule keeping up fixing them.
Many many companies run tests against the various kernel trees/branches, and will flood you with emails if you break the build.
I can't remember who (I think IBM maybe?), but Greg KH said one company made a fantastic test suite out of nowhere. If you submit six patches, it will try each one, tell you what you broke in patch 3, and recommend a way you can fix it. And this just gets emailed to the person who submitted the patches, and the maintainers if they actually merged it.
He said it was insanely advanced, and was made without the knowledge of the kernel maintainers. A company just made it and started sending emails one day.
And that is what xfstets is used for.
And many other drivers are "some vendor made it, and they won't tell us how the chip actually work, just give us driver code".
Then as other people mentioned hardware isn't exactly 100% predictable, HW bugs happen, firmware upgrades change how stuff works etc. so your test suite might be entirely unable to catch that.
0. https://en.wikipedia.org/wiki/Mock_object#Use_in_test-driven...
In hardware, you often have rigid boundaries with well-defined behavior, and almost no state in each unit. A big part of hardware verification is just testing that you did the basics of the interface right and that you won't deadlock anything. The rest of it is testing the behavior of your nearly-stateless thing.
Software testing tends to involve exponentially larger state machines, and fuzzier abstraction boundaries. The basics are mostly done for you, and you tend to do complicated logic that holds and manipulates a lot of state.
"Fuzzing" in software is very similar to constrained random testing in the hardware world, but it still doesn't catch a lot of basic issues due to the size of the state machines involved.
The testing philosophy you need to apply is not the same at all.
If I'm writing a library which contains algorithms or something like that it's pretty easy, seems useful or necessary.
But a lot of the time I'm gluing things together through API calls and it just seems pointless ?
An algorithm for, say, quick sort, will “test positive” on the same unit test that uses a brute force approach. A performance test might pick some differences up, but at that point that’s not a unit test any more
Is this not due to the legacy of a language that allows race conditions? Take Loom [1] in Rust which is a tool to test your concurrency primitives, ensuring that deterministically all interleavings are tested and verified to be sound. Then, when you expose a safe interface, you know that the only potential issues you have are user logic errors or not covering the entire problem space with your tests.
Still, reducing the problem space to such problems instead of any time you reach for atomics or a mutex, I imagine, removes the majority of bugs, so lets not throw the baby out with the bathwater because it's not perfect.
This assertion calls for links to data that backs this up I think.
Published counterarguments seemed to be easy to find, eg this survey suggests that most bugs are reproducible ("bohrbugs"): https://xiaotingdu.github.io/publications/TR2018.pdf
(journal link: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=24)
Here is an article about the various ways Linux is tested: https://embeddedbits.org/how-is-the-linux-kernel-tested/ Unit tests are just a small part of the array of automated testing used.
It gets harder to test low-level driver code, but some subsystems have similar sets of tests, e.g. MTD has a set of tests for drivers & flash chips that they interact with. In 2016 I ported those tests to userspace as part of mtd-utils. The tests for the kernel crypto stuff I mentioned would also test hardware accelerators, if enabled.
Filesystems have the fstests project (formerly xfstests, as it was originally written when porting XFS to Linux), testing filesystem semantics and known bugs. There is something similar for the block layer (blktests). The downside being that those test suits take rather long to do a full run.
Those are just things that I can think of on top of my head, given the subsystems that I have interacted with in the past. Static analysis on the kernel code is also being done[1][2], there are also CI test farms[3][4], fuzzing farms[5], etc. As others have pointed out, there is a unit testing framework in the kernel[6], IIRC replacing an older framework and ad-hoc test code.
[1] https://www.kernel.org/doc/html/v4.15/dev-tools/coccinelle.h...
[2] https://scan.coverity.com/projects/linux
[4] https://bottest.wiki.kernel.org/
This is especially true when we're talking about reverse-engineered device communication protocols, where you don't have a clue how the device actually works internally, and thus lack the basic means of constructing a mock that does more than just implement the exact same assumptions you've already based your driver code on (resulting in always-green tests that never find actual bugs). But also in cases where you have a protocol spec, the vast majority of ugly bugs in drivers usually originate from differences between that spec and the behavior of the device in the real world.
That would create a good description of the hardware that could be used for emulators and to test whether actual hardware behaves the way it was specified to behave.
> the vast majority of ugly bugs in drivers usually originate from differences between that spec and the behavior of the device in the real world.
Having the model hardware encoded in software would make it possible to test its behavior against real hardware for comparison and allow making adjustments.
> Having the model hardware encoded in software would make it possible to test its behavior against real hardware for comparison and allow making adjustments.
So you waste weeks to set up a hardware test jig, trying to match it with software and next firmware update changes some stuff that makes model incorrect again.
I mean I would expect that if you're a company that wants to put a driver for your hardware into kernel, but many drivers are just RE stuff and setting up a test jig for hardware that allows for fully automated testing can be quite an effort
When did you acquire such taste for luxuries?
I’m convinced that fully 75% of engineering effort in Silicon Valley is exactly this kind of testing. At least it is in my company, and we get people from all the big names, and none of them think it’s weird. Sometimes I wonder if I’m insane.
Not saying this is not true but has Linus stated this anywhere?
Because once unit tests have been introduced, the question becomes: Who / what executes those unit tests? Who makes sure no one breaks them? You basically need CI pipelines and infrastructure and suddenly the tiny code change is not so tiny anymore.
Those already have unit tests:
https://github.com/torvalds/linux/blob/master/lib/list-test....
using an existing in-kernel test framework no less.
If you step up into the 'lib' directory (containing generic in-kernel utility code) you might notice that there are a whole bunch of C files already that have an "_test" suffix in their name (or "_kunit").
Fontenelle, golden tooth, etc..
Here's one prominent example: https://github.com/linux-test-project/ltp
A while, back on hacker news, there was an article about a company developing a database. Their entire testing methodology came down to producing a deterministic kernel and thread scheduler. This allowed them to simulate every possible permutation of their concurrent code in scientifically speaking reproducible manner.
Developing this test framework was actually the majority of what the company did. Testing the kernel would be a similar level of effort.
Did they test against weak memory models? Because weak memory models do not necessarily have any equivalent interleaving behavior (stores can appear in different orders on different cores).
Atmosphere is but air, and production is but testing: a (hopefully) very long test run, which will only end when the project will be dismissed.
- https://en.wikipedia.org/wiki/Linux_Test_Project
- https://github.com/linux-test-project/ltp
Others have already commented on the testing situation for the kernel in general (historically, the tests were all in separate projects outside the kernel, and more recently, the kernel itself has a testing framework), but for filesystems in particular, there's an external project called "xfstests". The name might imply they're only for xfs (and they were originally made to test xfs in particular), but they're used for all filesystems (and also for the common filesystem-related code in the kernel).
x86 in particular has quite a few, although they run from userspace and exercise stable ABIs, so one might argue that they’re really integration tests.
In software real troubles comes from integration, not much unitary stuff (unless you praise Monkey coding).
And for Linux, it is not just a plain logic application but a kernel that runs on different hardware and has huge base of users and applications. Testing the thing as whole is the only thing that matters. Releasing alpha and beta software is a much more sound, rational and efficient approach. IT industry is (or at least was) organized to test before production.
I am certainly not against unit testing. It remains a wise approach for piece of software that need to be glued in stone forever or have super bounded in/out outcome (eg: bank transfers). But it is nothing more than a tool that you might use or not based on situations. Certainly never a one fit all solution !
There are some things that are absolutely worth testing. For example my current project is a library that calculates payroll taxes for US employers. This stuff is self contained, stateless, idempotent, but complex. And yet I know what kinds of inputs are non-trivial and can calculate the outputs by hand, so I can test it.
On the other hand, testing that CRUD form that submits data to S3 as well as saving it in the local database is nearly useless. Run through the form when you update it manually and make sure it works. S3 semantics change rarely, your own database changes would necessitate you changing the corresponding form.
When you have a complex monolith, unit tests and static typing allow you to create features and refactor the code fearlessly ("Computer says no" is a good thing here).
For example, a decades old cad product with millions of lines and tens of developers doing parallel changes. The unit tests allow the continuous integration system to dish out isolated reports on failing tests before pullrequest is merged to main, instead of integration test just bombing at some random place.
Unit tests save time.
Is this an evolution of English of simply a sign the writer is not a native English speaker?
(I mean no judgement with this question)
So the only approach that has worked is solid code reviews. There was a very interesting email exchange between Linus and an engineer from Sun. Linus was rejecting his patch because there was a spelling mistake in a comment. That's how strict he is about code quality.
Linux definitely has tests.
And without tests, how do you know if you introduced a bug or not? Send it to QA and wait for a week?
Do you really think most code out there is fully end-to-end tested automatically?
Having integration tests are a good idea. Mandating TDD dogmatism and selling unit tests as the only type of test acceptable is what's wrong here
What was written first: A. Unit tests B. Binary code
(The answer might shock some of you…) -
Which one came first is irrelevant, the question is if they help us write better software.
For example, checking that your backoff works correctly in the case of dropped packets is really hard to do at a higher level, because that part of your API isn't exposed publicly but is still critical to the correct function in degraded environments.
The people who insist on unit tests for 1-line functions are crazy and dogmatic though, I agree. But if a large fraction of your codebase is 1-line functions I'd also argue your architecture is crap. A codebase where effective unit testing is useless or difficult is a distinct smell.
Is there a reason something like the scheduler, which should be mostly algorithmic, should lack unit tests? I can see an argument against some aspects of device driver testing, but the scheduler is in a different class of code, right?
It's like quarantine when you know a part of codebase is infected by "bugs"
For the most part we had Linus eyeballing every line of code before merging it. And if you wrote an extra if or did a boo boo using the wrong enum flag or you overran your buffer, you were flogged, berated and pitied in front of an international audience of engineers.
I don't know how things go now in the post-sensitivity world.
Linux Kernel coders are real programmers. Do you think Mel wrote unit tests?[1]
Real programmers are discrete mathematicians in their head. With pointers.
Their code isn't perfect.
By analogy, when Andrew Wiles published his some 109 page proof of Fermat's Last Theorem it had bugs. [2]
The mathematics community tested it and he eventually corrected it. The Linux Kernel is like that.
There are no unit tests for a^n + b^n = c^n because no integers above n=2 satisfy them.
You can't unit test your way to secure, correct code in the kernel either. Only a community of testers and verifiers can do that.
"given enough eyeballs, all bugs are shallow"[3]
[1] http://www.catb.org/jargon/html/story-of-mel.html
[2] https://en.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27s_...
You don't need testing for verified code interacting with verified code in an environment verified to run verified code, but testing takes half the time for "good enough" reliability so we don't spend the time.
The dollar figure of the cost of a bad proof, or of a kernel bug, will always depend on the specifics of the case.
Unit tests cannot show the presence of bugs that you are not teating for, but they can show the absence of bugs you've fixed and do not want to re-introduce / regressions.
Not having unit tests strikes me as either hubris or lack of resources.
Are you really going to have people run repeated tests even if some of the tests could be automated? If a test cannot be automated, then that's understandable, but why wouldn't you automate what you can with unit tests and integration tests?
I would further argue that even if unit tests, integration tests, and a community of testers and verifiers are working together there will still be bugs, some of which will be critical.
Why wouldn't you use every tool at your disposal in addition to community / people to find and fix the bugs?
The reason I see not to do it is it takes time away from writing correct code. I could spend half my time writing unit tests or I can spend twice the time writing the code in a verifiable language.
How are you verifying that interface A and interface B are correct?
You would either need to test manually, test automatically (i.e. test automation, unit tests, integration tests, fuzzing, and perhaps even mutation testing or design by contract), or maybe you are using formal methods (proving correctness with mathematical rigor). If you are not testing, how are you verifying that the code is correct?
"It works on my machine" or "my teammate peer reviewed it and did not see anything wrong" are pretty low standards and how a lot of bugs get shipped and discovered in production.
>> The reason I see not to do it is it takes time away from writing correct code.
Again, how do you know the code is correct? Spending more time writing code is pointless if the code has bugs and exploits.