Systematically removing code
thepugautomatic.com
thepugautomatic.com
This is especially problematic in Ruby, as the author mentions, as the language is so dynamic that you're stuck having to write 100% coverage to ensure that variables are even named correctly.
On one end you have completely dynamic languages like Smalltalk where basically everything is late-bound, and it's perhaps hard to imagine that much of anything can be determined statically. On the other end you have completely static languages like ML where the compiler knows all and it's very difficult to escape its all-seeing gaze.
And then there's this huge gulf in the middle where the language leans in one direction or the other, but programmer discipline is perhaps even more important. Python code that's scrupulous about type hinting and doesn't go hog-wild with the metaprogramming, for example, can be easier to refactor safely and with the support of tools than Java code that makes heavy use of reflection might be.
Source: I currently work with, and occasionally have to refactor, examples of both.
I think that sometimes maintainability comparisons fail to be entirely fair because they rely on cross-language comparisons of codebases of similar size, rather than of similar levels of functionality. Which is only natural, since substantive opportunities for the latter kind of comparison are rare in the wild, but still.
To take a concrete example: The original Smalltalk system was this enormously featureful piece of software, even by today's standards. And it was written in one of the most dynamic languages ever, but it was also famously maintainable. My impression (admittedly based only on poking around at the edges a bit and not actually living in Smalltalk in any serious way) is that the reason for this is that the whole thing was implemented with so little - and such simple - code that, in stark contrast to just about any other GUI environment, it was feasible for a single person to read and understand the whole thing in a reasonable amount of time.
* As the size of the codebase increases, dynamic languages fare worse. Refactoring 10K lines of Python in a sweeping fashion is much harder than 100
* How many of those 1000 lines of Java are boilerplate, and could be generated/hinted at by IDE's?
* The more you're abstracted away from what's actually going on (types, memory, etc), the harder it is to optimize for performance, or even let it become part of your initial design/thought process. This rears its head later in the process
* How much does static analysis prevent you from creating bugs in the first place? If it takes longer to write error-free code, then the line count comparisons aren't always that meaningful
In my experience, basically not at all. It might help you catch some bugs a bit sooner, but, outside of languages like C that have no guard rails to speak of, the bugs I've seen cause problems in production are rarely, if ever, the kinds of bugs that static analysis tools are even trying to look for. I am a big fan of linting, but it's primarily because I like anything that makes the code less surprising to future readers.
For a similar reason, I see the Java boilerplate as a maintenance liability because, while the IDE can write it for you, it's a readability burden. Future readers have to look at that boilerplate, recognize it as such, and then read more closely to make sure it is perfectly standard boilerplate before they actually know what the code is doing. And any small deviation creates uncertainty. Was that deviation intentional or was it a mistake? The thing is, boilerplate is boring, so nobody actually does that. In short, it's a classic example of the paradox of automation.
Honestly, I think this trumps everything else. If you don't care enough to write good code, you won't. Yes, some languages make it easier than others to write good software, but that won't matter if you're just banging out code without a care.
That said, I do miss the typiness of C/C++ when I'm writing Python. I'd be interested to know how one enforces things a bit more strictly, even if it's only by running a separate tool (so something to be run on every checkin via CI, much as I run static linters for C and C++ code).
That's how. There are so many Python linters out there.
Ah, okay. I've been using pychecker, pyflakes and pylint for quite some time, was just hoping there was something more that would help do stricter type checking. I'm aware of python-contract as an orthogonal way to enforce things, but was a little disappointed it's not included with standard python (at work I don't always have the option of installing external packages).
At work, we are limited (for the most part) to what is included in RHEL repositories, and even more by the fact we are still on 6.9 on some projects. Hence why I mentioned python-contract, as it's not in RHEL repos (that I could find). Thankfully pyflakes/checker/lint are. I will have to check for mypy.
There are some interesting experiments in the Smalltalk world right now with a kind of "gradual" typing. Because Smalltalk uses and image and VM, as you use the system you can see what "kinds" of objects are passed with specific messages. The system then builds up a knowledge base about what works and what doesn't, and you can make use of that.
It's pretty cool!
Personally, I don't think type specifications are incredibly onerous, and I'd rather err on the side of a compiler catching my dumb typo instantly, instead of waiting for a 15-minute CircleCI run to finish.
I'm not really aware of a type system that enforces code actually be called. I know lots of linters do check for unused code, but generally only at the local level (ex - you imported/included X but never called it)
I agree 100% that it saves time during refactoring, but it's not going to catch that your NicheHelper class that was only called in one place is not called anymore. Nor is it going to find and highlight css that can be ditched as well.
* Not marked public
* Not called from somewhere in the call graph of a function that's marked public, or your main function.
This isn't just simply public though, to be clear. It's "externally callable," that is, it can see transitively if something is just locally public or not.
You also can't dynamically generate functions at runtime, or do things like method_missing, so this ends up being much more complete than a similar analysis in a language with those sorts of features.
The issue with a lot of dynamic languages is things like monkey patching and reflection, which means that functions can be called at runtime but otherwise there is no way of checking beforehand.
Some people swear by this dynamic flexibility and the benefits it brings, but there are some serious downsides in terms of refactoring.
Dynamic language tooling is really like banging goddamn rocks together, and people take pride in that...
Your NiceHelper function gets one type as argument and returns the same type? Is it `void NiceHelper(...)`?
Severely restricting the number of `void X(...)` functions on your code helps a lot. Also `a someF(a)` is a type that is useful once in a while, but if it's useful all the time something may be broken.
You cannot turn a compiled language into a language that offers the same flexibility as a dynamic language.
I guess that's a strong argument for Tailwind and other CSS-in-JS frameworks, since the CSS disappears when the JS does.
An approach like BEM can help somewhat because you rely less on cascading.
But then again, because it's tediously explicit, one may generate BEM class names in Sass or with PostCSS in a way that makes those long, explicit names harder to grep for…
What would you pay to have a service that did this automatically on your codebase? You remove something in a PR and it adds a commit that cleans up everything else that can be.
I will require changes in code reviews if PhpStorm cannot analyze usages in your code.
If it's being used, it should be tested, and then your test will complain.
I agree that tests are super valuable for many things, but it isn't enough here – at least not with the way my team writes tests.
(Our) tests wouldn't catch unused CSS classes or i18n keys.
(Our) tests wouldn't catch an unused method that is unit-tested. The unit test would happily pass because the unused method is still around and does what it should… in a vacuum.
Tests are great for catching live code that was incorrectly removed, but I'm not sure how they would help catch dead code that was incorrectly left in – but would love to hear more about it :)