Whereas in huge JS codebases relying on framework magic you can really have hard time to understand what's going on without plugging a debugger (and even then it's still not easy). Trying to refactor legacy code that relies on `this` and prototypes is a nightmare.
I sometimes wonder if JS programmers know that with C#/VS you can just change the name of a class, and without doing anything else (not even replacing), all involved names just change, and everything works 100%, without writing any test.
And the same goes for moving a method to another class, or renaming member variables, modules, files... everything just works and is instantaneous.
To me, not being able to aggresively refactor withour worrying something will break somewhere is like backwards and old school (in the bad way)
When a project evolves, a class, a struct... grows and suddenly a name, perfectly fine at the beggining, does not make sense any longer. With C#/VS you just change and 2 seconds later you're still working.
The same can't be said for languages like JS.
If I get code and all of a sudden not even my IDE knows what's going on, it clues me in that there are layers of indirection in the codebase I'm going to have to waste time reading into, like magic methods or loose typing.
Both PHPStorm and IDEA are written in Java (they share a common code base). Don't confuse [Java eagerly allocating memory from the operating system but then not using it for anything (yet)] with java using a lot of memory. For speed, java allocates memory it anticipates using in the future. This lets it manage it's own memory without involving the operating system and thus context switches. If you start using a lot of other applications and available system memory starts running low; java will detect this and relinquish the pre-allocated memory it's not using.
edit: further learning those simple tools has a more generic application to other problems while the IDE magic will likely go until there's an issue to be understood and you're depriving yourself of learning generally applicable tooling to trade a slight inconvenience for magic that has the potential to cause great inconvenience and once that magic from the IDE bites you and you learn how it works - you've learned something with no general application to other problems.
edit2: since I can't reply in-line
jcelerier - I _could_ do that with CLI tooling and depending on the language it would be trivial. However, I would _never_ do that because changing the method/function name to suit my context would be very unlikely to avoid making other code less readable. That's a huge anti-pattern that your IDE is making easy. Further, consider the implications for language design when making language design decisions that require a specific type of IDE magic to be considered a reasonable language design decision. I have a hard time believing relying on the IDE lock-in is good for the community of that language.
philwelch - I genuinely don't know the answer, but what happens when you change a class name -> forget you do it before saving the file -> go to another file and try to change the class name there but have a typo -> go back to the original file and save the change? I would bet the typo'd change goes unchanged and now you're left scratching your head since that magic has always worked before.
Are you relying on a correct and comprehensive project/workspace setup within the IDE?
Of course. Not much point in using an IDE otherwise.
To quote myself, "fix a single file manually, instead of a dozen or a hundred or a thousand".
And what happens when you need to rename, say, a method name that's also used as a variable name in many places? xWhat happens if multiple classes define the same method name, but you only want to rename that method for a single class?
Yes, you can probably do it in sed. If you're used to it enough, you can probably do it pretty quick, with sufficient regexps that you only have to debug a few times.
Or you can right click or hit CMD-. or hit F2 or whatever with your cursor on a method name, type in the new name, wait 2 seconds, and you're done.
You generally get compilation/build errors highlighting the bits that were missed (for whatever reason)
That's not something I "rely" on, it's the first thing I make happen whenever I use an IDE.
Or you have a larger codebase split up into multiple .Net solutions and DLLs where code is referenced through the generated DLLs.
The refactor looks great, until it starts breaking things for the many consumers of your code.
In fact, the latter is a situation where simple text tools will make life a lot safer and easier than whatever the IDE does (which pretty much will be limited to the bounds of the .sln file, whereas a simple text tool operating over files in the entire codebase will alert you to the fact that this is being used elsewhere).
My argument isn't to claim that simple text tools are better at refactoring. They aren't. However, I've run across several situations where we've had issues in a massive codebase because developers believed that refactoring was as easy as hitting the Refactor button in Visual Studio, and everything will work magically.
The belief in programming 'magic' (whether through IDEs or frameworks that abstract a ton) without understanding what those tools actually do is what I'm worried about. Use IDEs to do what you need to by all means, but it's much better when you actually understand what it's doing before doing something with it.
Sure--and to be fair, I've never worked in .Net much. But if you're changing the package interface, you should arguably do so in a backwards-compatible fashion anyway, depending on how your stuff actually gets built and deployed. If everything's contained in the same repo and you can't refactor across multiple components in the same repo, that sounds like a problem.
> The belief in programming 'magic' (whether through IDEs or frameworks that abstract a ton) without understanding what those tools actually do is what I'm worried about. Use IDEs to do what you need to by all means, but it's much better when you actually understand what it's doing before doing something with it.
Agreed!
Yes, in general, the ability of IDE to make these kind of large-scale changes quickly, reliably and mostly bug-free is nice. But I rarely miss it in languages where it's impossible.
It may be that one of the ways your language shapes your code is the way you structure it. For instance, in my Lisp codebases, I can hardly think of any place where Java-style automated refactoring could be useful; in the current largish codebase I work on, it rarely is the case that constructs introduced in one file are used directly in more than one or two other files - which makes refactoring entirely doable with a dumb text editor, and Java-like automated refactoring wouldn't save all that much time over doing a regex search and manually visiting every line listed in results.
I'd really like you to show me how you can refactor e.g. a method called "write" in a 500+kloc codebase in less than 5 seconds with "simple tools". With any C++ ide from the last decade it's basically a keyboard shortcut + typing the new name + fixing the few remaining compilation errors that may have cropped up in generic code. With sed & al ? good luck, see you in a few weeks.
I'm not taking sides on the issue discussed here, but regarding this quote... This is one of the biggest reasons I don't do OOP. Using plain functions with slightly longer but globally unique names so many issues go away.
And the biggest issue is that I can't read a code base if so many function names are not unique and possibly not conveying enough information for local understanding. That style always requires the reader to carry so much context in their head, to be able to resolve the "polymorphism" to concrete implementations. (Yes, in a proper IDE we can jump to the implementation interactively with a few keypresses, but it's still much harder to read).
Only on monolith code bases without any signs of modularity.
In other words, don't make it look complicated before it actually is.
That was kind of the gist of my comment.
foo.log(xyz);
bar.log(42);
vs log_event(foo, xyz);
log_event(bar, 42);
Provided that both approaches do the same, which is more clear? It is the second one, because it clearly suggests that there is no polymorphic switching based on foo/bar.Wading through non-trivial projects, pervasiveness of the first approach makes me nervous, because I'm told at every occasion, "don't even try to understand what's going on in the system globally. We .log(), shouldn't that be enough for you to know?".
While the second approach is calming to me, "You see, we have this simple logging system, where everything ends up being written to. It's not that complex after all, and if you want to know what's going on you can just jump directly to the log_event()" function.
Yes, by asserting that foo and bar are of the same type, and asserting that the .log() method is not virtual, I can reach the same conclusion. But that's the point, it requires work that you can't practically do if each line of code contains a method call. There is no mincing words, you're using the wrong syntax...
First foo or bar can be function pointers.
Or maybe log_event is a macro whose behavior depends on the environment. And I have used some crazylogging ones back in the 2000's when coding C, that would even dump the current state of the process alongside its parameters.
Or then again, log_event() plugs into a configurable logging system and one cannot predict its behavior just from the call site, even when using the same types.
Finally if we move away from C semantics, maybe log_event is overloaded or is doing multiple dispatch, thus the two calls, depending on the argument types, are not going to land on the same implementation.
Then different uses of the same identifier still resolve to the same implementation, and make the IDE jump to the same location, etc.
(I acknowledge that everything can be something else. And in theory you can redefine macros to have uses of the same identifier resolve to different implementations. Etc, etc. I'm sure you know about the IOCCC. But here we talk about good programming practice. It's a good practice not to suggest something complicated while the reality is simple. There's no argument to win here.)
> Finally if we move away from C semantics, maybe log_event is overloaded or is doing multiple dispatch, thus the two calls, depending on the argument types, are not going to land on the same implementation.
Which is why I do not use C++, or at least avoid most of its features. For me, it's all about clarity and reducing number of moving parts. It's not a contest in making complicated things that pretend to be simple. And it's not a contest in making simple things that look complicated. And it's not a contest in reducing the number of characters in individual identifiers at all costs.
so, question : suppose you have your log_event, and now your bosses comes and tells you that the client needs to have two additional different logging mechanisms - say, to journald and to some websocket log server. The system can log to either of your three logging mechanisms at any point through changing a configuration somewhere in a GUI but you also need some core classes to always be logged through journald no matter the state of the rest of the system.
How do you update your design to reflect that ?
What is needed in the face of changing requirements is actually this: lean structure with few dependencies that can be adapted.
That wasn't hard.
Also you've made your logging library project-specific by making a log_core_event function. Thanks but I'll keep using spdlog in my projects, which do not need newcomers to learn about your log_core_event, what is a core module and what is not, just to be able to write the three functions they are called for in that project !
That all has nothing to do with virtual methods.
> thread initialization if you have multiple threads doing logging on startup, right ? :-) )
There is no technical difference between "global" and "local" variables. It's merely a syntactical difference.
If you run into concurrency problems with global variables, you'll run into problems just as well with objects. Unless you make multiple objects. In which case, you could also have done thread-local storage, no?
In the end, it's much clearer to me to use global variables (TLS or not) since this way I can actually see the data relations. If there is just one instance of a thing, why would I pass it around as objects? It's not going to help understandability.
> Also you've made your logging library project-specific by making a log_core_event function.
That doesn't make sense. Can't I just add another function without breaking everything? Note that I probably wouldn't even add this function to "the logging library" but to the core modules...
> what is a core module and what is not
If you prefer having everything switched behind your back, resulting in unreadable and unintelligible source code... you do you.
More technical notes...
If you don't like the presence of two or more functions for logging, then why not have an explicit context argument? You call log_event(MODULE_FOO, foo, xyz) instead, and the log_event function implements all the plumbing logic in a central place, easily understood (Most logging approaches do that, I think. They have "facility" and "severity" arguments). This approach seems much more preferable to me compared to an implementation that you can't see locally, and that can't be influenced locally. What would you do if we go for the OOP approach with a virtual method, and now call into a different module that would need a different logger?
If having a dynamically switched implementation was really what we want to do, there is no problem with doing that. My post was just about using method syntax for static things, which I think harms readability.
well, that cuts short to the discussion. There are plenty of places & coding standards where those are outright forbidden.
> If you prefer having everything switched behind your back, resulting in unreadable and unintelligible source code... you do you.
we have very different notions of readable. For me the less stuff I have to read and the more readable it is - I just want my code to be the pure domain problem and could not give a rat's ass about how logging or networking or whatever is implemented if those aren't my domain.
> Can't I just add another function without breaking everything?
that does not break the code but that breaks Joe intern's mental model and expectations of anyone who will think that log_* functions are from the same library
> then why not have an explicit context argument?
and here we are, reinventing OOP by hand in C. There is literally no difference between log_event(context, ...) and context.log_event(...) - but you get better tooling :
* completion -for the first I'd most certainly end up typing the full name if there are a lot of l-started functions while for the second I'd type c<shortcut>.l<shortcut><enter> and only see the few relevant functions in the autocompletion list) etc... in the second.
* namespacing - if I want to see all the logging functions available I can just click on "context" and my IDE will happily show me all the available methods in a panel and nothing else. With C functions, well, I have to scroll in a list of 150000 functions. And also wonder (& have to explain) why log_event() is in <util/logger.h> and log_event_core is in <core/core_logger.h>.
* split implementations - if at some point you want to port your software to, say, an arduino without network capabilities (talking literally forom past experience here), the monolithic log_event function now becomes an #ifdef HAS_NETWORK / #ifdef HAS_JOURNALD / ... mumbo-jumbo. While with multiple logger subtypes it's just a matter of changing your build system to not compile the unavailable implementations, without even needing to change the code (or change a list of types in a header if you use static polymorphism instead).
> What would you do if we go for the OOP approach with a virtual method, and now call into a different module that would need a different logger?
dependency injection ? modules can request either for a default logger which will be sourced from the gui configuration, or whatever specific logger they actually need due to functional requirements.
Besides, I still don't see the readability problem with virtuals. Again in my IDE if I want to see "what happens" I press F2 on context.log_event and it displays me the list of all the possible implementations if it isn't able to determine that one is used statically at that point. And I much prefer to read three five-lines email_logger::log, journald_logger::log, websocket_logger::log functions than a pile of if/else's.
But when I had a glance at the project you linked on github, that was basically the first thing I found there...
A suggestion that I have there would be to use the linker instead (make different source files based on implemntation / existance of an implementation). This way you can eliminate most of the #ifdefs.
Regarding dependency injection, I don't see a practical difference to just using global variables / global functions, apart from the fact that by convention DI gives you back objects with virtual implementations, which I don't like.
I won't go into the rest of the things you said because we pretty much discussed it all.
* jar-hell (and dll-hell) aside
That makes me assume two things: you work alone and your projects don't exceed 5.000 (fivethousand) lines of code. Without some namespacing and an expected method length less than 100, more likely 10 lines than 90 or more, you end up with roughly 160 names. An average human knows 150 people, maybe by their name. And while you have friends or at least parents and grandparents occupying some of the »address sace« few of your methods will not be »at hand« for your brain.
And in case you already prefix your methods with something like io_write(...), db_query(...) and you hand in context-specific references like io_write with a sort of file handle and db_query with a sort of database connection I transform all this 1:1 in an OO-style language and back.
For a clearer description of what I mean, see my other comment below where I write two lines of code in both ways.
The tools are at a point where I'm usually breaking my code when I'm doing manual refactoring, but never when I let the tools do the job.
> what happens when you change a class name -> forget you do it before saving the file -> go to another file and try to change the class name there but have a typo -> go back to the original file and save the change?
A good IDE does not care if you save the file or not. It is always up-to-date. Renaming a class in just one file is also impossible, since the IDE will do refactorings atomically over the whole project. But lets just assume you somehow managed to fuck it up that badly (things can always go wrong): You will still realize your error almost instantly because a statically typed language does not defer such checks to runtime. The program wouldn't even run anymore.
By far the biggest issue with automatic refactoring is accidentally renaming stuff in comments (or forgetting to do so) because I'm too lazy to look at every single place. But find&replace has the same problems.
pub fn add(a: i64, b: i64) i64 {
return a - b;
}https://tryidris.herokuapp.com/compile#YWRkIDogSW50IC0+IElud...
The above link refers to the following Idris code:
add : Int -> Int -> Int
add a b = a - b
addTest : add 4 5 = 9
addTest = ReflHow exactly does this knowledge help? I know this is possible and I do miss it. But it's extremely diffcult to support it in JS because of how the languages works. VSCode tries, but it's still limited. So should we all switch to writing C#? You can't avoid writing JS if you want to create a web app. There are tools to avoid it, but one has to know what's under the hood or they'll have a bad time.
TypeScript is doing great progress on this part, and it supports easy refactors, but it's still lacking in some parts, for example the type inference is still poor when it comes to complex cases and the type declarations depend on will of maintainers.
Another honorable mention is Spring Boot @Conditional for which adding a minor little diversion from defaults will turn a functioning happy project in a dysfunctional nightmare.
:'(
I have even come to question if we should make all our javascript apis async from the start to avoid this scenario though it seems a rather extreme guideline. We didn't really do that, but I still dread this type of refactoring.
(My dream code exploring tool would be a cross between an IDE and a tablet&pen friendly PDF reader. I should be able to navigate through files using all the semantics-aware IDE goodies, the tool should also understand version control, but it should also be able to generate and manipulate structural graphs (as in, boxes and arrows), as well as let me highlight, draw and freehand over the code.)
The overhead has significant cost for trivial gains. The costs result in overly complicated systems to solve simple problems. Most of the complexity in these systems is for dealing with complexity brought in by the ecosystem, not the problem domain.
Ruby Example:
Ruby projects tend to always have really well maintained test suites. The language makes it easy for you to introduce non-obvious bugs or write unreadable (yet efficient) meta code.
This has lead to 1) very mature testing frameworks (meh) and 2) high adoption and usage of these frameworks (actually impressive).
Funny enough, I think adoption of IDEs is extremely low in the Ruby world. 90% of the developers I interact with in my city are on Sublime/VSC/Vim. I've seen some pretty impressive usage of Rubymine that has made me curious, but never bothered to really spend time with it myself. I'd be really curious what the teams at GitHub/Stripe/[Insert big Ruby company] typically use editor-wise.
How does it compare to tmux/screen?
That’s an assumption which advantages Java by default. Other languages may be able avoid having a large code base entirely. I recall reading some stories of people rewriting giant Java code bases into very small Clojure programs that were more flexible, maintainable, and scalable.
I must also admit I woild rather ten lines of code and be clear and certain, so being verbose doesn't bother me much.
This isn't a secret, and in certain corners not even considered generally true -- backlash against microkernels and microservice architectures have intensified in the last half-decade, to give two examples.
I don't see a reason an eCommerce framework couldn't have a small footprint. The lower bound is the amount of actual, necessary complexity in the problem space. But that's usually orders of magnitude less than the size of the codebase, because of all the boilerplate and greenspunning that happens in Java, Enterprise Java in particular.
(Note that in commercial work, I moved from Java to Common Lisp, so I have a different view than typical about just how much boilerplate and repetition can be simplified and removed when your language lets you do it.)
I'd rather have 1M lines of Java than 100k lines of JavaScript, and if essential complexity took 1M lines of Java to flesh out, you aren't reaching parity with 100k lines of JS.
Why not? I can easily see reducing SLOC 10-50x by switching from Java to JavaScript.
> if essential complexity took 1M lines of Java to flesh out
I very much doubt that in a 1M SLOC Java codebase - at least in pre-Java 8 codebase - the essential complexity is any more than 1% of that. My work experience with Java suggests that the language and the surrounding ecosystem and practices promote extreme proliferation of accidental complexity.
I wouldn't mind seeing the other side of the fence though, I will admit I have worked in enterprise codebases for most of my career so I am definitely biased.
I recently rewrote a c# project that had not been completed in elixir. Hugely, anecdotal (as all of these types of stories are) but the project went from ~5kLoC to <800LoC when rewritten in elixir. That's after the elixir app was feature complete too. Now you might believe that's a developer experience gap, but the c# dev had way more experience with c# and in my experience this other dev is much better than me generally. Within a week of learning elixir he was catching my bugs all over the place.
In comparison, my current project in TypeScript is at this stage now and by using an IDE things are quite under control. Refactoring is easy and even major architectural changes are manageable (we did ~5 so far, depending on how you count).
Wasn't the case with Erlang. We only managed to do one such major change and it almost killed the project. It probably even actually killed the project as we were late to the party and failed afterwards.
I can rewrite any bad written code in less lines, same language same framework as initial code.
Granted it isn't 10s of thousands of lines, but it isn't small either. It does a lot of things ranging from webapis to data pipelining with ingestion and egestion of large datasets. One batch job can result in 10s of GBs in output files. It has to be performant too, 10s of millions of "operations" in one batch job is common.
Admittedly the process/threading model can be painful to work around at times, but totally doable, and is honestly somewhat friendly to distributed systems to an extent because it somewhat forces you into an actor model paradigm.
I'm not trying to belittle what you do, but if your good base is in the 4 figures in size then it's not large (in fact yes, it is small). When people talk about large code bases they are generally talking 100-1000x larger than yours, or more.
I just think it's foolish to think Java is the only language that can handle large codebases. I mean just look at the linux kernel.
EDIT: Removed comment about processing capabilities because it's irrelevant to the discussion. I left my original comments about lines and the IDE though.
Current job I've seen single files that are over 30k lines.
I disagree. Most people working on the Linux kernel (including Linus) or the major GNU projects (including RMS) are using vim or emacs with maybe some supplemental find/grep/sed/awk.
(Yes, I know you can do Java in VS Code but I wouldn't want to manage a large Java codebase with it.)
Why not?
“Code should be written for humans to read and only incidentally for machines to run” (paraphrasing a famous quote by Harold abelson)
You can find all the places that call a method and know they are really all. That sort of basic thing.
E.g. not for writing, but for analyzing what unknown code does and for refactoring that code. Comparing to javascript, in js you have to be much more careful when refactoring and changing things.
1. used when a person has something more to say, or is about to add a remark unconnected to the current subject; by the way.
2. in an incidental manner; as a chance occurrence.
That quote seems pretty misinformed, code has to be a lot more connected to the machine to be runnable that that quote would have us believe.
Otherwise this would be valid code:
Ask the user their name. After they enter the name, load a game. Use some cool graphics.
That being said, perhaps that will be possible in 100 years. (Yes the first line may be possible in 10 years)
In other words, don't doo premature optimization. Write the code so that it's readable, except on the chance occurrence where you need to eke out more performance.
So this is a cute statement but isn't it actually false? None of us would be sitting around code reviewing each other's code for fun if there wasn't a mountain of money / value / knowledge to be extracted from a computer running the code at the end of the day.
Actually this is easily falsifiable. As soon as quantum computing becomes possible the first languages will probably be quite horrible, with painful tooling and torturous debugging and difficult to read control flow compared to current modern languages. Nonetheless, a huge army of programmers will be trying to write code for these quantum computers, for the sake of the computer, and not for any human to read.
Of course its not a universal rule, its not a law of physics. Doesn't make it less useful as a guiding principle (similarly, design patterns in software architecture).
The core argument is that we're besides the point in conventional (2019) computing where we need to worry too much about optimizing code to run as efficiently as possible. The complexity and bottlenecks to delivering features is now in writing software that is meant to be worked on and maintained by a large number of humans. And for that purpose, it is better to optimize for clarity.
I think the conundrum of java is not quite the same thing as strictness.
* Eclipse has a crappy UI and can't handle Java9+ code very well, but can handle large code bases
* IntelliJ Ultimate has nice UI and assistance, but it totally fails with large code bases and still tricks you if you run code inside the IDE, especially when you want to use jigsaw modules and not classpath, as it uses by default
Unfortunately, using Java with other editors (vim, VS Codium, etc) is just difficult