And unit testing is relatively useless, what would be useful are end to end tests that start with userspace API but that's much harder task, altho hopefully those tests wouldn't need to be changed that often.
i.e. does not depend on external timing or fickle hardware.
Very, very difficult to unit test someting that runs using hardware interupts.
There's also the Linux testing project, which is technically third party. It's not clear to me how extensive it is but for a project as important as Linux I think it has to be graded as "needs improvement."
But testing costs a lot more when there is hardware present. It's a whole new dimension in your testing matrix. You can't run on cloud provider CI. Hardware has failures, someone needs to be there to reboot the systems and swap out broken ones (and thing break when they run hot 24/7). And you need some way to test whether the hardware did what it is supposed to do.
Although some kernel driver developers have testing clusters with real hardware (like GPU drivers devs), there is no effin way Linux could be effectively tested without someone paying a lot of money to set up real hardware and people to keep it running.
Of course the hardware can be simulated but simulation and emulation are slow (and potentially expensive). At work we run tests on simulation for next gen hardware which doesn't exist yet. It is about 100,000x slower than the real deal, so it's obviously not a solution that scales up.
Testing purely software products without a hardware dimension is so much simpler and cheaper.
Just crowdsource this testing: I am sure that there exist some people who own the piece of hardware and are willing to run a test script, say, every
* night
* week
* new testing version of the kernel
(depending on your level of passion for the hardware and/or Linux). I do believe that there do exist a lot of people who would join such a crowdsourcing effort if the necessary infrastructure existed (i.e. it is very easy to run and submit the results).
And having any kind of manual intervention required will almost certainly reduce the reliability of the testing.
This is further complicated by the need to reboot with a different kernel image. Qemu and virtual machines can't do all kinds of hw testing needed.
And in fact, the kernel is already tested like this. Just very irregularly and sporadically. The end users will do the field testing and it is surprisingly effective in finding bugs.
> And having any kind of manual intervention required will almost certainly reduce the reliability of the testing.
Perhaps I am underestimating the necessary effort, but the willingness and capability problem can in my opinion be solved by sufficiently streamlining and documenting the processes of running the test procedure.
If the testing procedure cannot be successfully run by a "somewhat experienced Linux nerd", this should be considered a usability bug of the testing procedure (and thus be fixed).
Dogfooding of builds is also testing, just a different kind.
Hardware testing takes up a pretty sizable portion of the development cycle. Why do we think we are special?
SQLite Test Harness #3 (hereafter "TH3") is one of three test harnesses used for testing SQLite. TH3 meets the following objectives:
- TH3 is able to run on embedded platforms that lack the support infrastructure of workstations.
- TH3 tests SQLite in an as-deployed configuration using only published and documented interfaces. In other words, TH3 tests the compiled object code, not the source code, thus verifying that no problems were introduced by compiler bugs. "Test what you fly and fly what you test."
- TH3 checks SQLite's response to out-of-memory errors, disk I/O errors, and power loss during transaction commit.
- TH3 exercises SQLite in a variety of run-time configurations (UTF8 vs UTF16, different pages sizes, varying journal modes, etc.)
- TH3 achieves 100% branch test coverage (and 100% MC/DC) over the SQLite core. (Test coverage of extensions such as FTS and RTREE is less than 100%).
That's not something you see called out very often. Correct code fed to a broken compiler can definitely give you a broken binary. Likewise a correct binary will pass on a correct simulator and may fail on broken hardware.
I take that as an argument not for or against tests but against making it easy for management to turn your project into a feature factory.
That weirdness is precisely the sort of thing you want automated tests for.
I'd say that's very common and not exclusive to some specific project.
I made a change of about 30 lines of code recently, and it had 140 lines of tests (and I, by far, do not cover every situation).
I think I've rarely encountered well tested pieces of code that were not way smaller than the tests that came with them.
It's more useful to categorize tests by how hard to setup an environment. In reality cost/usefulness line lies on this boundary.
* in single process * multiple process using IPC * multiple processes using network * tests validate functions calling foreign services
I see many devs (including myself) call "integration" tests as "unit" tests because in your particular app/system they are easy to spawn locally even without any container.
You are free to write tests that compose packages, and I’ll do it for my own sanity and confidence. But that’s in addition to, not instead of, the minimum “not horrifically irresponsible engineering” standard of universal, exhaustive, very detail oriented mock driven unit tests within all packages.