Nowadays with so many open source projects and so much distributed computer power, it is even easier.
Then i saw that they try building those projects against their library as part of their test pipeline. If you want to add your project to the list, you can submit a pull request. No wonder the list is so substantial - it gets you free testing! Highly ingenious!
So plenty of programs rely on UB even without realizing it.
Anything less is asking for trouble.
Also, every test that was ever written for a compiler is still out there, being run, somewhere.
And, for the record, the real world is pretty big.
Back when I worked on a commercial compiler, we had a smoke test which we ran before committing. It would take ~30 minutes to run. We had a longer, roughly eight hour, set of tests which we ran daily. And we had machines running tests 24/7, which managed to clock up roughly 20 million distinct test cases (combinations of code and options) over the course of a year.
Every single compiler bug we'd ever seen was present (in a reduced form) in the conformance tests we'd run before release. That took about a week.
Depends who you work for and how willing they are to let such justifications stand in a court of law.
I’d much rather have that than a compiler that evolves slower because its developers worry about existing non-conforming code.
I can always use the old version of the compiler if I don’t want it to break.
Same with a list sort that is stable but never guaranteed it would be, then turns unstable much later. I was bitten by this one in .NET as were many others, and as a response the developers added hacks into the code to simulate the old behavior it detected the code was built against the older lib. That didn’t make the problem easier to find...
Sometime practicality wins out and the change is reverted when it breaks everything. It’s nice to come up with purist views, and in this case being able to point to documentation and say that the entire world needs to change is the morally high stance to take, but it’s not always what happens. (Some may say that this is what needs to happen to undefined behavior, but that one’s still up in the air. As it so happens I am on the “read the standard” camp…)
There also seems to be this weird misapprehension (probably not shared by you) that this is the compiler being "out of to get you" for UB. That's not what happens... the compiler ASSUMES that no UB can happen and optimizes accordingly. This can certainly have surprising results, but doing anything else would leave a lot of performance on the floor. What has happened over the last decade is that compilers have become better and better at optimizing leading to more and more surprising results.
To anyone (again, not you) ITT complaining that their program 'used to work fine': It just happened to work, but it's been buggy from the get-go. If that's a problem for your ego, then you need to adjust your ego.
Of course nobody actually thinks that the compiler is "out to get you". But if you do any work in a security context, that is the attitude you have to take: "If the compiler was lawful evil and trying to put security vulnerabilities into my code, what excuse could it use from the spec to justify doing so?" That's the only way to make sure that you won't get bitten by "well-intentioned" optimizations.
> To anyone (again, not you) ITT complaining that their program 'used to work fine': It just happened to work, but it's been buggy from the get-go. If that's a problem for your ego, then you need to adjust your ego.
This would be valid if what people were complaining about was random cowboy behavior. What people are complaining about is the fact that it is nearly impossible to write code without UB. The fact that UB can result from "trivial things like x+1 where x is signed" is exactly the problem.
Not perfect, but should work for most sane people running CI before merging changes.
Or phrased differently, the test suite is (or should be) the machine-readable language specification.
If you look at Rust, you cannot commit to master:
- without the resulting master merge passing the whole test suite, which would probably take a day or more to run locally, and - without a review of someone that knows the part of the code you modified well, or somebody they delegate to, and often, of a group of multiple people that know that part of the code well. These people verify that the code is correct, commented, tested, appropriate, etc. Often, a design team in charge of that part of the code that makes sure that the change goes in a direction that aligns with the project strategy/goals (i.e. that this is a change we want to make).
Even the best of the best software developers:
- don't always test that their code compiles before committing: for Rust, they really can't to, they'd need to cross-compile the compiler dozens of times, run these compilers under qemu on different hardware, and verify that they work - too much work for any single human - don't always test that their code passes the test suite: the test suite is _huge_. Its impossible for anybody to test anything locally, takes too much time. - don't always benchmark the compiler after a change on reasonable applications, takes too much time. - etc. point being, you cannot trust the best of the best, much less anyone else. As a Rust compiler dev, the first person that you should not trust is _you_.
The whole point of the process is identifying all possible sources of errors, and making them impossible, or at worst, extremely unlikely to happen.
Most of these errors are only quite unlikely to happen, e.g., because all the automated tooling that exists to prevent them is also written by humans that make mistakes.
But every time a new source of error is identified, there is at least a discussion about how to make it impossible to happen in the future.
Having a test suite, having reviewers, are just two things you have to do to avoid a large set of errors (see a comment below about how each release of Rust is required to succesfully compile all existing public Rust source code and run all their tests succesfully).
They are not barely enough. And even if your process makes errors extremely unlikely, Rust and LLVM have hundreds of PRs on flight and hundreds of developers working on the source code regularly. Throw enough developers at a project, and unlikely events become a daily thing.
Tests are fantastic for defined behavior. but running actual real world gigantic battle tested programs is where I found bugs. It _seems_ like you can just define the behavior. In my experience, the interpreter would mostly do what I wanted, unless like 4 complicated factors were in play. Often, it's so confusing, the user of your language is just trying to make something work. They're under a deadline, they find a way, and they make it work. They're not morons, hell, you have to be pretty sophisticated to find that weird edge to exploit.
I was super new, but I thought of a lot of my coworkers as surgeons. They'd see the problem, and find the smallest possible change to fix a bug. They were very careful.
I dunno. it's super hard. I made some optimizations where I had a choice, I could break a working program (but relying on undefined behavior) in one way or another. Seemed like with those big projects, the optimization is totally worth it. But you have to pick. If you choose the semantics to mean 'a' then X breaks, if you choose 'b' then Y breaks. Our users had been around the block a few times. They were sympathetic, and we held a pretty firm line on undefined behavior.
I dunno. I think rust probably handles this the best, recompiling and retesting against every crate ever. But here's the thing, if you're relying on some weird edge case, the test probably is as well. So you don't actually detect the change in behavior.
I think it's a hard problem, and there's no silver bullet. Write tests. Have careful engineers. Run big programs. In spite of all that, you're going to have to go to your users hat in hand occasionally, because when you build big systems there are lots of subtle edge cases. Shit's hard man.
I’ll be a contrarian, and say the careful software engineers are the most important. They’ll stop feature work to recreate the test suite.
On the other hand, nothing can save a good test suite from being selectively disabled or regressed by sloppy engineers.
Similarly, having a process that dictates careful code reviews just slows everyone down. The careful engineers would have done the code reviews properly without the policy. The sloppy ones will spend more time to produce comparably bad (or even worse) code reviews.
That might have been K&R's original intention, but that went out the window pretty early on when optimizing compilers arrived on the scene. And eventually when the ANSI spec was written, it was explicitly written to allow all kinds of optimizations (the "as-if" rule).
The ironic thing about that is that when they wrote C89, they said they were trying to standardize existing practice. I don't think they had any idea that they were effectively declaring 90% of all programs buggy.
With hindsight, maybe K&R should have consulted with PL experts and researchers and designed a "proper" language with a better emphasis on soundness and safety. I'm not saying they could have invented something like Rust back in the 1970'ies, but even then the state of the art was far ahead of K&R C.
I agree. During the standardization process, there was a proposal (‘noalias’) that dmr described as “a license for the compiler to undertake aggressive optimizations that are completely legal by the committee's rules, but make hash of apparently safe programs”, and, that “[i]t negates every brave promise X3J11 ever made about codifying existing practices, preserving the existing body of code, and keeping (dare I say it?) ‘the spirit of C.’”¹
If he or anyone else at the time had realized that ‘undefined behaviour’ would turn out to do the same, there response would have been the same: “[it] must go. This is non-negotiable.”
Do you have a source for this claim? AFAIU the history of C, it was invented for entirely pragmatic reasons.
> This went out of the window with clang, and with gcc (and glibc!) being based on C++.
This is nonsensical, whichever way one interprets "being based on C++". The GCC and Clang compilers conform to the standard whichever version of C or C++ you ask them compile. (Anything else is a bug and should be reported. They may accept a few non-conforming programs, but that's allowable under either implementation-defined behavior or undefined behavior. Remember that UB can include 'doing what the programmer expects'.)
In practice, you need all, for various reasons.
> in terms of the correctness of the final product
What is there "correctness"? Ah, I see in your later post:
> I'm only referring to the implementation being true to the language specification
The answer is: even checking for that is not something which can put the humans out of the loop. But humans alone (as in "extremely thorough code reviews" or "very careful software engineers") are provably not enough as the "space" that has to be tested is too big to be useful without "a comprehensive test suite." Still I'm not aware that a complete "comprehensive test suite" which would be "enough" for any non-trivial language can exist.
In my experience, non-trivial goals are simply bigger than any "simple" or "single" "solution." So in that sense, any project eventually fails due to the human factor. It's the humans that will use anything else as the tools to achieve the goal. It's the quality of humans that are in charge that is the precondition for everything else.
So the answer is always: prioritize for having good people, anything else will be taken care of by them: they will maintain the tests, be careful, do the reviews etc. Don't expect you can "save" on something.
Asking humans to be careful at scale (whether authoring or reviewing) just won't work.
Keep in mind that you need loooooots of test suites, not just one.
Here's a test from 2007 that tests a similar error in the instruction selector:
https://github.com/llvm-project/llvm/blob/master/test/CodeGe...
Evidently the error that caused this test to be rewritten reordered the `ret` instruction before the `mov $1, $eax` instruction that sets the return value. This tests assures that the bug will never happen again.