The unsafe version can be deprecated and eventually removed, but that is down the road.
The unsafe version can be deprecated and eventually removed, but that is down the road.
1. Have 3 versions of the thing to change: select_nospec, select_nospec_safe, select_nospec_unsafe.
select_nospec_unsafe is identical to select_nospec. Do the mass replace. All stay same.
2. Delete select_nospec
3. Start migrating to select_nospec_safe. You can still do this in small steps, because you old thing stay.
4. Where select_nospec_unsafe must be, keep it and maybe mark with a comment or similar to not delete it later
5. You finish and all is changed.
Maybe replace select_nospec_safe to select_nospec if wanna.
It's relatively easy to do the mass replace in master. You leave the burden of the replacement in the other branches to people who own those branches.
Set up an automated script, to keep the 'unreplaced' version out of master (and people can also use that script to check their own branches).
Not sure if that counts as non-trivial, yet?
There are probably some other trade-offs involved. The people running kernel development ain't idiots.
Assuming perfect tooling and infrastructure doesn't sound like it necessarily have to be a big deal.
But, how far into fantasy land are we? How much effort would be required to get there? Is this and all similar problems large enough to justify it?
Is adequately tooling and infrastructure even possible?
Does that make proposed changes, such as moving to C11, order of magnitudes more difficult?
The only major issues is that the change is let half-way. But that is only possible with MAJOR undisicpline.
---
The major point, to me, is that this EXCESSIVE fear of breakage is not good. Yes, making upgrades is pain, but you CAN make the pain tolerable with some planning.
Refactoring is like exercise for code: Everyone dislike the idea of excessive, but is GOOD.
And lets be honest, among all the things that could requiere a refactoring, this case is on the most simples of the simples scenario..
struct foo *iterator;
list_for_each_entry(iterator, &foo_list, list) {
do_something_with(iterator);
}
While the new version would be something like: list_for_each_entry_v2(iterator, &foo_list, list) {
do_something_with(iterator);
}
So the ’iterator’ variable may be removed, but as this is c, might also be reused multiple times, or come from some struct member, union, or whatever. So i presume something like this would have to be fixed case by case.Mind you, there are not a ton of people around who have the time to learn how to write custom clang-tidy rules. It might be done for 15,000 changes that potentially fix kernel vulnerabilities, though.
You can brows the clang-tidy source code yourself. There are existing checks specific to the Linux kernel in there already. Well, there’s one check. But if you use clang-tidy, you’ll discover that automatic refactoring of extremely large code bases is within reach.
https://github.com/llvm/llvm-project/tree/main/clang-tools-e...
The only question is whether you would want to spend the staff hours working on a clang-tidy check. Large code bases are exactly where the tradeoff makes sense.
http://bbannier.github.io/blog/2015/05/02/Writing-a-basic-cl...
Yes, it involves forking. Forking isn’t that big a deal. We even had checks that were specific to internal libraries.
Really… I have worked at two companies that had extended clang-tidy to add custom checks for their code base. The whole point of clang-tidy is to automate code fixes in other parts of your code base. When done well, use of clang-tide pays down technical debt faster rather than adding to it.
> tools for mass renaming that try to guarantee that there are no regressions
beyond sed or whatever a fancy IDE already does? What kind of voodoo magic are we talking about here? Did Google solve the halting problem and I just haven't noticed yet...
As for the actual renaming, yes, definitely beyond what sed does but probably around what the fanciest IDE does. Imagine a big global symbol dependency graph produced by the entire build toolchain all at once and cached, I think?
Also, this isn't the halting problem unless your codebase allows for fully dynamic invocation :)
Given that compilers can be smart-enough to detect "non-dynamic" function-pointer invocation (e.g. when an execution-trace proves that a function-pointer parameter always points to the same function address), it's not safe to say that "all" function pointer invocations are dynamic.
Another case to consider is when one implements (Smalltalk-style) OOP with message-passing: in many cases it's possible to build that without needing to use function-pointers at all.
So you can exclude many things that would in-principle be possible.
I'm not 100% sure, but i remember that i could have have both a two and three function argument be called with the same function pointer, its probably UB however i think that as long as any dereferenced functions adhered to the standard argument behavior of the platform specific calling convention it worked, and i dont think gcc, clang or msvc complained, but my memory might be wrong.
You want there to be two states: State A where the thing was called old_name and State B where the thing is called new_name. State AB where some code thinks it is called old_name but other code thinks it is called new_name is broken, and must not exist.
† This might seem outrageous to modern programmers, but CVS thinks in terms of files, so from its point of view it's fine if out of sixty files you tried to commit, 48 of them succeeded and 12 failed. Good luck fixing the resulting mess.
SVN, baz/bzr and friends are what you get if you add networking before atomic commits. Git is what you get if you start with the proper data model and only then add networking.
as you might imagine, it’s probably a pretty involved process to touch thousands of lines of code in a complicated system.
Bors works for git. Google has similar tooling for their system. (I think it's called the 'train'.)
1. https://en.wikipedia.org/wiki/Coccinelle_(software) 2. https://en.wikipedia.org/wiki/Sparse
Cool concept, but once you've dealt with any legacy codebase that has used it extensively, it feels like an anti-pattern/footgun. Ultimately you just want the language to do those things directly and for the compiler to be smart enough to optimize it later.
Essentially it is too clever for its own good.
#define BASE hey
#define HEADERS foo bar baz
#append HEADERS bad bal bah
#push OSSUFFIX
#ifdef WIN32
#define OSSUFFIX win32
#else
#define OSSUFFIX unknown
#endif
#foreach H HEADERS
#eval include "$(BASE)/$(OSSUFFIX)/$(HEADERS)"
#endfor
#pop OSSUFFIX
People go to great lengths making all sorts of weird structures from x-macros, repeated statements, defines that only exist to be used by other defines, etc all to work around existing preprocessor limitations - and many of them would simply become unnecessary if the preprocessor could do things like variable editing, loops and being able to eval its own commands.Even though some stuff can be done via language features, it is often necessary and more flexible to work with the source code itself.
When I first learned C++, using cfront in the late 80's, lots of the language was implemented as C pre-processor macros.
You can even check out how it was done: https://github.com/seyko2/cfront-3/blob/master/src/template....
Its also surreally poorly designed, encouraging worse habits than C itself. Avoiding definition collisions and lack of namespaces alone make it horrific for any moderate sized project and up. Combine that with the near complete lack of static analysis tools, and its a recipe for disaster.
I don't claim to be an expert in knowing what is/isn't in the various standards; I just look at the build errors and static analysis alarms. That said, I think the argument is that in the vast majority of cases it's not actually possible to change the type of a variable; off the top of my head, I can only think of pointer co-erosion. Otherwise, if you define an unsigned int, it stays an unsigned int.
Now strong/weak typing isn't necessarily the same thing as type safety. C has always seemed astonishingly bad on that front. It's like they try to trick novice developers into thinking they have a robust type system sometimes.
If you define functions implicitly they will link to any symbol, EVEN A VARIABLE... this is bananas. I think I sort of understand why compilers work this way, but it really feels like a bug in the language. Under some compilers you might not even get a warning for implicit function, either! TI had a compiler that hid them by default.
Enums are garbage in C. Again, you can misuse them, and may not even get a warning. You can pass an enum for color to a function that takes an enum for kittens and the compiler will be happy as a clam. The way they're often used can cause constant implicit conversions every time there's an assign/compare. May not be a problem for most positive values, but it's annoying if you're trying to develop standard compliant code. MISRA defines an "essential type system" and normal enums usage violates it.
I'm probably forgetting a TON of deficiencies. Please add them or correct me where I'm wrong.
It’s the C preprocessor that causes a mess. Tooling for C and C++, like automatic refactoring, lags behind tooling for languages like Java, C#, and Go. Refactoring tools have to deal with macros, conditionals inside #if/#else blocks, and header search paths.
In this case, the refactoring involves removing a variable from the enclosing scope of a macro invocation. The most likely way to automatically refactor it would be to write a custom check in clang-tidy.
Eg C tracks whether a variable is an int or a char. Haskell also tracks whether a function causes side effects or not. More sophisticated systems also track whether a function has to return eventually or could run forever.
Typed languages let you know whats broken at compile time and aides in refactoring code you can see. It doesn't help you refactor code you can't see.
The unison language offers an interesting take on the problem of code evolution backward compatibility. https://www.unisonweb.org/