Could you explain what you mean? What bug is there in sqlite? What blog post are you referring to?
Via someone else in the thread, this was reported as a bug in LLVM, which was confirmed and is now fixed: https://bugs.llvm.org/show_bug.cgi?id=46194. In light of this, I wonder why you claim that this was a bug in sqlite.
1. OSSFuzz reports a bug against SQLite.
2. SQLite dev tries to fix the reported problem but is unable to repro the bug on his desktop
3. SQLite dev replicates the OSSFuzz build environment which uses clang-11.0.0 and is then able to repro the reported bug
4. SQLite dev finds that the bug isn't in SQLite at all, but rather in clang-11.0.0
5. SQLite dev patches SQLite to work around the clang bug, posts a brief note about this on the SQLite forum, and goes to bed.
6. The post on the SQLite forum is picked up by HN while the SQLite dev is asleep. LLVM devs read HN, isolate and fix the clang problem, all before the SQLite dev wakes up.
Now compare this with what happens when you meet an MSSQL bug.
In 2014 the small company I worked for made a release with my work in it and I was spoken to the next day because it had crashed. The boss was not pleased as it had to be down to my work, nothing else had changed. In fact it was down to geographical data in the DB being changed as well. One of the geog functions (edit: think it was STIntersects, but the next bit applies to many such functions) claimed to return a 0 or a 1 - but in fact I found, separately documented, that it could also return NULL.
While spending a couple of hours digging this out, I found via a blog that geography data types used approximate algorithms (which is not surprising, and quite reasonable when you understand why, but it was not documented). These two problems caused us what could have been a serious production problem.
There's lots more. I've hit too many. Just a few weeks ago I did a straightforward recursive CTE on a small data set, which took a minute to run. Eh? It was an optimiser bug. Breaking off the base query into a temp table first and it dropped to a second.
I've too much of this shit, much more. MS doesn't care. BTW if you use the long-existing "update from ...join" syntax which MS released years ago as an extension, be aware it allows you to do stupid things that ansi-standard SQL does not, like repeated unpredictable updates to the same row, in the same single sql statement. I hit that last year in someone else's code. It's broken. Well, what's new.
There's parts of the 'What's new in C# 8.0' that aren't really there, and the provided code samples are entirely wrong (i.e. not just one issue, the language literally can't do the feature.)
See also the state of most documentation for ASPNETCORE.
> BTW if you use the long-existing "update from ...join" syntax which MS released years ago as an extension, be aware it allows you to do stupid things that ansi-standard SQL does not, like repeated unpredictable updates to the same row, in the same single sql statement. I hit that last year in someone else's code. It's broken. Well, what's new.
I'm curious if you have an example. I know it's not ANSI standard but I've yet to run into any issues.
MERGE on the other hand is a dumpster fire and everyone knows it.
As for update from, see here https://www.sqlservercentral.com/forums/topic/is-update-from...
IIRC If you join on unique columns you're safe, but like I said, I found 2 cases of this last year in some code where the cols joined were not were not unique. That meant the output was, quite simply, junk.
> select isnumeric('$+.') > 1
so it thinks it's a number but obviously it isn't
> select cast('$+.' as int) > Conversion failed when converting the varchar value '$+.' to data type int.
This kind of thing makes data cleaning even less fun than it is. They now have TryParse(), but the above (which I'm pretty sure used to allow an even wider range of crap) is just unnecessary, and is a sign of how they view their own flagship product.
> No, the point here is that the fuzzer "caught" a bug in sqlite.
Which is where the confusion stems from. Perhaps "'caught a bug' in sqlite" would have been a more clear way of writing what you intended?
> OSSFuzz has been reported bug 23003 against SQLite.
Just looking at these lines, I don’t understand the code, though. Apparently, sqlite3VdbeMemRelease can change pMem->flags, but that change can be ignored. I can’t think of a simple explanation for that.
Also, the function name sqlite3VdbeMemRelease seems to indicate memory gets released, but pMem->flags = MEM_Str|MEM_Term…, to me, indicates it contains a string.
No the changes _must_ be ignored. This is where the bug came from.
The code does store the flags, does an operation and then restores some flags while setting some new ones.
But after clang reorders the operations it will not restore the flags from _before_ the function calls but instead takes them from _after_ the function call.
I.e. any change `sqlite3VdbeMemRelease` does to `pMem` mutst be ignored but after reordering changes to `MEM_AffMask` `MEM_Subtype` are no longer ignored.
A project cannot burden its code trying to work on every single commit of every compiler out there.
clang fixed it in one day. for gcc I expect 5 years to fix their regressions.
Rust has Crater which allows new rustc builds to be run against the entire public crates.io dgraph.
Clang is lucky that OSSFuzz does testing of unreleased clang version on other OSS projects, like SQLite.
Does anyone have a link to the reported clang bug?
For example, it might well be that the function that clang believes does not modify pmem->flags "promises" that it does not modify the value. In which case it might be an SQLite bug, that clang just did not happen to "correctly optimize" before.
This can happen, e.g., if you mark your function with __attribute__((const)), for example, which promises that your function does not access memory via arguments that have pointers. It then becomes trivial for clang to make the transformation above.
The bug wouldn't even need to be on the function definition, it might just be on the declaration, or on a declaration somewhere, in which case you have an ODR violation as well.
All that seems very unlikely since it appears to be clear that the function does modify memory, but still, I'd just wait for the clang bug to be filled with a reproducer, and the issue to be tracked down.
It does strongly hint at a bug in clang, but from the linked page, I at least cannot conclude that yet.
http://lists.llvm.org/pipermail/llvm-bugs/2020-June/084028.h...
I have better things to do than reading the whole SQLite code base like, for example, fixing the bugs of the people that actually bothered to fill in a bug report.
Until this bug is confirmed, it can be anything, and unfortunately, bugs reported for C and C++ projects are quite often not bugs in the compiler, but UB in user code.
The fact that you think that just by looking at a single function signature one can deduce where this could be UB related to const depicts the problem. UB is unfortunately not that simple, and I can think of a couple of ways of const-related UB in which the declaration "looked fine".
Normally true, but not for SQLite. SQLite is probably the most thoroughly tested and correctness conscious piece of software anywhere in open source. I think with SQLite, the burden is on the compiler writer.
https://bugs.chromium.org/p/oss-fuzz/issues/list?can=1&q=Pro...
Given the amounts of bugs that are both filled and fixed on a daily basis, reality does not seem to reflect this claim.
If you expect compiler writers to go and explore your open source project of choice in a blind hunt for compiler bugs that they can fix, I hope you sit down while you wait and do not really need these kinds of bugs actually fixed.
The actual semantics of the language are effectively impossible for a non-expert to tell from the code as written and operating. And what you don't know can hurt you. That's like juggling caltrops with bare hands - you're guaranteed to get hurt.
Which attitude? The one that requests users to fill a bug report if they think there is an error in our compiler ?
Please let us all know which software you write and ship, and how you inspect all inputs that your users pass it in the search for bugs, instead of requiring users to fill a bug report if they think they encounter a bug.
Or maybe you are talking about the attitude of not assuming that all users know the C standard to the letter and never write code that has bugs ?
I suppose that for all the C and C++ code that you write, when the binary doesn't work as expected, you start by looking for bugs in compilers instead of proofing your own code right? Because it is almost impossible to write subtly buggy code in C and C++ right?
The attitude that undefined behavior is seen as an opportunity for optimization, and backwards compatibility with your installed base is not particularly important. Not even if you know that the change breaks code in the wild which was intended for security checks.
See https://gcc.gnu.org/bugzilla/show_bug.cgi?id=30475 for discussion around one such change.
I will elide the specific self-congratulatory attitudes that you claim to have, and your not so subtle insults to myself. With one exception.
Or maybe you are talking about the attitude of not assuming that all users know the C standard to the letter and never write code that has bugs ?
This is exactly the reverse of the attitude that I've observed.
Including in yourself 2 posts up when you described how hard it can be to decide whether something is undefined behavior, but if it is undefined behavior then it isn't a compiler bug. Which, as a developer, tells me that I definitely am not able to decide whether the code will do anything like what appears to be written, or I will get unexpected nasal demons, in some unrelated section of code.
When you choose to write C, you are accepting the contract that if your program has UB, there are no guarantees about what it does.
If you can't uphold your end of the contract, C is not for you.
There is a huge difference between complaining that this contract is too hard to uphold, which is a complain that you should take with the C standard, and complaining to the compiler that, on contract violation, your program misbehaves.
If you can't accept that, again, C is not for you.
For every novice that complains about their mistake misbehaving I have 100 expert C programmers that would complain about bad codegen for programs that do uphold the contract.
From a compiler writer perspective, it is impossible to argue that we should penalize those that take the time and experience to play by C rules in favor of those that write broken programs.
100% of the experts would agree that it is hard to write correct code, but "disabling compiler optimizations" is not the right place to fix that.
If you believe that's the case, you are free to compile all your code with -O0, use volatile everywhere, or just use Python or some other language that suits you best.
But the argument that you are making is ridiculous. You don't like the consequences of breaking the law, so your conclusion is that prisons should be 5 star hotels and all the law-abiding tax-payers should finance them so that you enjoy them at their cost. I wish you the best luck in life with that attitude mate.
When you choose to write C, you are accepting the contract that if your program has UB, there are no guarantees about what it does.
Back when I started programming, C was widely understood as having a very different contract. Namely, "It is a very simple language with no guard rails, but describes exactly what the computer will do." This was the language as describe in K&R.
The contract that you describe was something that compiler writers came up with long after the language was in wide adoption. Based on my anecdotal impression, it also wasn't the direction that most of said user base was asking for. And it is the opposite of the direction that most other languages have gone.
If you can't uphold your end of the contract, C is not for you.
Which is what I said from the start. My exact words were, "As a developer, this attitude on the part of C compiler writers is a reason to avoid C and C++."
I generally am fine in slower languages like Python, JavaScript, SQL, Go, etc. The next time that I need performance, I will look into Rust. Isn't this exactly what you think that I should do?
But the argument that you are making is ridiculous. You don't like the consequences of breaking the law, so your conclusion is that prisons should be 5 star hotels and all the law-abiding tax-payers should finance them so that you enjoy them at their cost. I wish you the best luck in life with that attitude mate.
You know, if you argue with what I said rather than what you think I said, you might make more sense. Again, here is what I said:
"The actual semantics of the language are effectively impossible for a non-expert to tell from the code as written and operating. And what you don't know can hurt you. That's like juggling caltrops with bare hands - you're guaranteed to get hurt."
Given the level of care and knowledge which is required to identify undefined behavior, the first sentence is true. Given how much compiler authors value squeezing more performance over backwards compatibility, so is the second. The third is the logical result of the first two.
Expecting a compiler (and the hardware it runs on) to be perfect is a mistake.
- Parsing, type-checking, and pre-simplifications
- Assembling and linking
And there have been bugs on edge cases there. Moreover, correctness is limited to what is specified. COQ does not guarantee that there are no missing specifications or theorems. The C standard is written in English, people had to manually encode the standard in COQ.
I build LLVM from source a lot, and master not even building because somebody made a change that doesn't even compiler is astonishing.
FWIW, LLVM has a pretty big test suite. But... to catch errors... you actually... have to... run it...
It is often the case that Rust, which does have a policy of "master always passes all tests" catches bugs in LLVM before LLVM does.
I'm glad sending patches or looking into bugs whenever I find them in OSS but having to jump through hoops just to get things compiled is a big no-no.
The bot will catch it, yes, but you are still filling the queue of maintainers.
For example, some people send doc fixes. Sometimes these doc fixes have "bugs" that actually make compilation fail. Like adding a line break and forgetting to mark the next line as a comment...
Compiling LLVM for such changes takes quite a bit of time, so of course people don't even do it.
So why would you try to compile your software to verify those instead of just, e.g., generating your documentation ?
Also, this was not specifically targeted at llvm but all open source software which may not even have a build-bot or published regression checks.
All I'm saying is: Please keep master clean so it always compiles so people can easily pick it up and contribute.
And obviously, under the hope that somebody notices the failure by then, tracks it down to the right commit, and fixes the commit, instead of adding a workaround that avoid having to find the root of the problem.
"master doesn't pass tests" isn't just about naming. Its about culture. Its about agreeing that there is some source of code that everyone agrees on should work, and having the culture of agreeing that all modifications to it should work as well, and that it is the responsibility of those doing the modifications to make sure that's the case.
Continuous integration for compilers is a well-understood problem. There are posts in this thread explaining how Rust handles it. I understand that it's a royal pain to set up and maintain. I am a compiler engineer and am happy to work on compilation stuff but wouldn't want to touch our CI system with a ten-foot pole. But LLVM is a project driven by Google and Apple, there are really no excuses for not finding the people willing to do this.
If the commits don't even build they might as well be thrown away. It's going to make regression tracking next to impossible otherwise. The only reason to keep a branch around is because it's a small ephemeral branch under active development, or because it's an eternal branch that you can use to track regressions (using git bisect, for example).
But how would you tell which one that is? I guess even within the current system one could set up a bot that monitors master, and whenever the latest commit builds and passes tests, it updates a "known working" branch. But this doesn't seem to be the case. And if such a system is set up, it would make a lot more sense for contributors to push to a "testing" branch instead of master and then only updating master for working builds, rather than the other way round.
Sure, one way is to have an automatic test suite and an integration system that moves commits that passed it from one branch to another. If you can actually make a test suite that has no false negatives and finds enough negatives to be useful, that's great.
In my (limited) experience test suites are not close to that ideal. I've worked with systems where there was an automatic build test at least, and sometimes that catches a bad commit. But I think for many fast-moving projects there are few useful errors these systems can find, and sometimes they have annoying false negatives. The thing that I can think of that would most improve my productivity right now is a check that nobody added tab characters in the source code.
(I realize my comment might not apply to a situation like LLVM, where there could be a large test suite of programs to compile and run for example).
Compilers are also often very modular. LLVM in particular makes it very easy to take a bitcode file, apply a certain well-defined set of transformations, dump the resulting bitcode file, and check whether it fulfills the expected properties, e.g., "this computation has been optimized to this other computation". LLVM has a tool called lit that does this kind of test, and there is an extensive test suite that uses it. Of course none of this guarantees that other programs will behave differently, but in my experience these test suites really are very good and useful.
The only missing piece in LLVM is consistent enforcement of a policy that tests must pass in order for the repository to reach certain well-defined states.
I mean something like setting up something building and running the tests for everything included in Debian, homebrew and vcpkg and looking for regressions compared to the last release of the compiler.