Should you split that file?
pathsensitive.com
pathsensitive.com
1: file agnostic. While files provide a scope for imports, in general the compiler can’t tell the difference between the same set of classes split across multiple files or crammed into the same file.
2: unordered - files don’t import one another, or get processed in particular sequence. The compiler takes a set of files and turns that into a set of type definitions.
2: partial types, which lets you even split a class definition between multiple files
3: regions, to let you bracket of sections within a file and organize things how you want.
No ‘one class per file’ rules. No ‘references at the top of the file to every other file you need’. You can reorganize code freely, adding or combining files at will.
Used to have a setup where, during development, I write all my .js files as "script files" rather than modules, all placed inside a folder, and the "build-step" is simply concatenating all .js files in that folder. All the functions in each script file are globally visible from all the other script files and there's no need to import/export anything. No one should publish code in that form, but during development it was incredibly freeing and helps me concentrate on the actual programming tasks at hand rather than getting sidetracked by trying to figure out the right way to organize the code into modules, something that (imo) I'm in a much better position to do when the programming tasks have largely been solved.
It's like writing a book. The natural process is to first create substantive content in fragments and then gradually for a chapter or section structure to naturally emerge from these fragmentary bits of content. Focusing too much on 'modules' during early development is imo akin to being preoccupied with spending time deciding what chapters and sections you want your book to have before even having a substantial base of content.
Unfortunately the current tooling for JS doesn't really support this workflow as well as I'd like.
It’s only convenient as long as you stay in visual studio. Not that I would write C# any other way, but the comparison isn’t exactly fair. It’s very easy to create reusable libs in both TS and Python.
> It’s very easy to create reusable libs in both TS and Python.
I've obviously been doing something wrong then, because I've spent many cumulative hours trying to make these work and there's always compromises. Ever tried making an app with create-react-app that uses a shared lib, with reasonable support for HMR (or even just a working live reload), type hints, breakpoints, de-duplicated dependencies, and linting?
As for python, is it even possible to do with a standard python install? I can't remember all the issues I ran into, but the conclusion I came to is that doing it without something like poetry (at the very least) is so much effort that it's not worth it (and let's not even think about making it cross-platform). It doesn't help that every single source on the internet seems to recommend a different approach.
Compare this to C#, where there's exactly one way to add a reference to a shared library, it takes one line in a csproj file, and it works with just about every type of application out of the box.
Edit: To be clear, I'm talking about locally reusable libs, not published and versioned packages.
But if you are clicking around a project and scrolling with your mouse, instead of taking advantage of your IDE's ability to use keystrokes to find and open files, jump to references/definitions, find and replace text, etc., you are handicapping yourself. It's worth the investment to learn.
Really? I've observed the opposite.
The developers that I work with that make heavy use of things like fuzzy file matching, go to definition, etc tend to have a _much_ harder time understanding the codebase.
My theory as to why this is the case is that they're never building or taking advantage of any higher level context. Having to browser through the directory structure (assuming it roughly mirrors the code structure), having to look through classes, etc forces you to actually build some awareness of the greater context of that method you're looking at.
Using "go to definition", you may find yourself in a method, then later you hit "go to definition" and find yourself in another method, and unless you've made a conscious effort to observe so, you may very likely have never noticed they're even in the same class and presumably related.
Using "go to definition" you might end up in a "process" method. Knowing that's in the "Image" class in the "ImageResizer" module would probably add a lot of clarity to what that's doing before you even get started looking at code.
When I need to onboard in a new codebase, I very much _don't_ use those tools until I have a solid grasp on the big picture.
One trick I use[1] is adding certain characters before the section name. e.g:
¡¡ settings
This makes searching and jumping between sections a lot easier.[0]: https://en.wikipedia.org/wiki/Literate_programming
[1]: https://ricardoanderegg.com/posts/write-apps-in-single-file/
I've used this occasionally in Emacs Lisp code and it's pretty nice, but I've never done it in other languages because I suspect that non-Emacs-users won't be able to handle it :P
The Emacs Wiki has a whole page on this: https://www.emacswiki.org/emacs/PageBreaks
Splitting into files is an approach that is easier to lint against (e.g. max 500 lines - a best-effort mechanism to foster SRP).
And it's not really problematic after mastering a few key aspects: jump to definition, and more interestingly Peek Definition.
It's also important to not traverse files top-to-bottom (which yes, is very tempting).
Optimal code traversal is tree-like, you jump from definition to definition, skipping over the irrelevant, and gathering necessary info on demand.
The file/line location of a function becomes irrelevant.
A (successful!) ex-colleague of mine read thousands of lines long files top to bottom.
I don't know how, but somehow this worked for them.
You could as well try memorizing the entire $framework API, which most people would agree that isn't the best use of one's energies.
The minor rough point is that text editors probably don’t support mixing line folding techniques, so you probably won’t be able to fold sections and functions, which it would sometimes be nice to be able to do. Maybe at some point I’ll get round to making a hybrid fold expression for Vim to mix multiple fold methods, though honestly I find Vim’s folding pretty janky (my Rust foldexpr works fine when you load code, but start editing it and the fold tree rapidly falls apart in painful ways and it’s absolutely Vim’s fault) and would probably prefer to shift the entire thing to FFI for more flexibility.
Some languages/IDE ecosystems support other ways of doing this sectioning, like Visual Studio’s “regions”, which gets around the folding problem since it’s all in the language rather than being a mixture of language syntax and independently-chosen convention.
Things computers are good at: organizing files, topological sorting, ontologies.
The point is I wish we could just ask the computer to sort things out for us, show different views, etc. this would require a software system able to work with more granular code units, maybe something like [1].
I still have some ptsd from huge refactoring where a big part of the hassle was to manually create directories, move files around, fix imports, build files, cyclic dependency errors, sigh…
I suspect a good solution would be to have a language agnostic dsl to describe a program’s data flow, then use it to generate the boilerplate of functions involved. Then you could ask the system to show you only the pieces of the code strictly related to a certain feature, hopefully never again have to deal with file structure manually.
I agree. I'm usually opinionated about a related subject: code formatting. But I'm not too opinionated in one case, where we're using Prettier. The project is autoformatted and expects everyone else to do the same. It's so refreshing to let the computer take over sometimes, even if the end result isn't exactly what I would have typed out myself. If a piece of software could so something similar with the folders, files, and code that would be pretty great. Maybe. It seems like a much harder problem to solve than simple formatting, without introducing a bunch of bugs.
But you seem to be suggesting a different way of viewing the code, which I guess might be the way. But still underneath you'd have the spaghetti code, just sitting there not being dealt with, like a complicated Microsoft Word document.
I'm wondering if the unison devs already have plans on working on a protocol like that.
The problem I've seen with this idea is that in order for it to work, you have to be willing to add more metadata to your program, because programs themselves don't have the necessary data in them. You can make it mandatory to your language, but than just raises the friction for people trying to use that language.
But, if you're willing to add the metadata to a hypothetical new system for slicing and dicing code, why aren't you using the existing mechanisms in your system for organizing and documenting things?
If you are, then you make the delta between "it may not be perfect but here's some documented and organized code" and the perfect vision small enough to not be worth a lot of effort to chase.
If you aren't, then you aren't going to give the perfect all-singing slicer and dicer what it needs either.
Replace "you" with "coworker(s)" for the same dilemma, only sadder. And your generic strawman coworkers are going to be very upset when you try to get them to use this language... "it's so hard, it's always complaining about perfectly good code, I'm missing my deadlines because it's making me document things I'll never use, I'm just putting 'a' as the documentation for these things anyhow, oh look here's our boss ordering us to go back to our previous language because we're all spending too much time documenting rather than building new features".
Some things can't really be fixed by process.
I'm kind of stumped by this assertion. What data is missing?
If it is required or necessary, doesn't it make it just data, as in an integral part of the code?
I'm confused by what you mean. Can you be more specific as to what sort of metadata that has to be embedded in the code would be needed?
I've never used it, but watching this Smalltalk-80 IDE demonstration seems like it could be a workable idea.
https://youtu.be/JLPiMl8XUKU?si=Hly_8m4GiaDSQmpn&t=269
Edit: A more modern example using Pharo. You can see in the video that he's not accessing files, rather everything is sorted into packages and classes.
Even advanced type systems mainly focus on guiding how components connect, but they don't explain _why_ a program behaves a certain way or how the data is supposed to flow. Documentation is often unreliable as it becomes stale easily. As developers we spend and order of magnitude more time building a graph of data flow in our heads than actually writing code. Could we reify that graph?
IMO that crucial metadata would answer the questions we would ask the Senior developer in our team when we need to work in the codebase:
* When dealing with feature X, what are the involved software components Y and Z, how do we create instances of those?
* What are the actual features the system implements? Is this the correct behavior? Is this what the system was intended to do?
* If I point to a specific file and line in the project, what features does that code support? Why is that file there?
Expressing the answers in a machine-readable format, in a language agnostic way, could look something like:
System {
Dependencies
- Database depends on Logging
- UI depends on Database
- ...
Feature A
- Subtopic 1
- Subtopic 2
Feature B
- ...
}
Feature A {
requires database
queries X
if (database has data)
produces Y, Z
else
doTheThing(X, "blah")
}
This is a broad idea, and the devil is in the details... This would not replace type systems or traditional languages. I'm thinking this spec-lang would have the minimum amount of flowchart-like constructs, and everything even slightly more complicated would rely on an interface. The system would implement all boilerplate and layout the project entirely. It would be amazing if we could also express state machines or even statecharts.Teams could use off-the shelf generators, or create a custom one, to produce the interfaces and boilerplate implementation. Since the system has a lot of information we could query it to answer questions like "give me all code of feature A", or even ask it to create symlinks so we can focus our attention in a single folder of files. Also different generators could create the structure that is more appropriate for different programming languages. It could be a good tool to assist rewrites, if needed for whatever reason. Tests/specs could be generated automatically, to a certain extent.
I'm sure some people here will have PTSD about 4th gen languages and BPEL/UML nightmares :-). But even though BPEL is horrible I think the core of the idea is good and worth revisiting.
Based on what do you make this statement?
I don't feel like this is a problem. At least not for me :| Of course, size can make everything a problem, but for whatever projects I had to work with, organizing files was usually an entry-level problem that would normally sort itself out after a month or so.
In other words, I don't feel like this is a problem worth solving.
But the friction to get something else established is so huge - all tooling would need to change ... unless perhaps one develops a way to translate between the internal format and text.
When moving a function to a different module (e.g. moving bar() from foo::bar too baz::bar), the development environment should automatically change all references to foo::bar() to baz::bar(). And this is just the tip of the iceberg of what the development environment should do for you.
The problem is, developers have come up with all sorts of clever hacks to do this searching and refactoring, including IDEs which do much of the heavy lifting, but at the end of the day they still fall short of what should be a simple task for the developer.
As for additional metadata that would be required to make this work: objections that programmers would be reticent to provide this info are similar to objections to writing good documentation, and my response is: you reap what you sow. If you really want good, structured code, you have to write it, and you have to do so in an environment that accepts such structured data. You can't get blood from a stone.
In short, for large, complex projects, our current method of writing code as textfiles is unwieldy. Unfortunately developers' visions are clouded by their hard-earned ability to parse those textfiles, and don't see this as a problem.
Of course it is a goal, but not always possible. You sometimes have two (or more) conflicting "relations" in the codebase. E.g.: you are dealing with taxes in various countries. Do you group by country or by kind of taxes (on sales, profits, earnings, energy, real estate...) ?
The right answer really depends on how your team is organized and how you are making changes.
Seems like you would be dealing with one of 2 reasonably well defined problems in this case. How to group would follow logically. It's either:
1) "If i will be dealing with one specific type of tax at a time and how it is applied in different countries. (i.e. What are the sales taxes in UK, France, and Germany? )" Group by tax.
OR
2) "If i will be looking at one country at a time and their assorted taxes. (i.e. What are the taxes in the UK for income, sales and value added?)" Group by country.
Else "i have failed to properly define the problem." that's gonna be the problem.
You could make two tools, but now you're duplicating implementation. Factor out a shared library/service/modularization-approach-de-jour? Good idea, but we're back to the question of how to structure things.
You don't always get to enforce that "OR"
2. So you have a class that has a bunch of getters and setters. Let's just assume that "generate them automatically" is not an option. You want to make it really easy to see the part of the class which is getters, and the part of the class which is setters, and then skim past that. How do you do it?
3. So you have a file that defines 3 data structures. Each data structure has a definition, a bunch of functions for parsing it, and a bunch of functions for serializing it. The author suggests that you split the file into 3 sections for the types, with subsections each for the definition, parsing, and serializing. How would you do it? Let's say the language is Rust or Typescript.
Dear God it's a 3300 line file. Any way but one long 3300 line file is a significant improvement. I'm being hyperbolic but seriously, instead of button.less it should be deduplicated (less can be much less verbose, pun intended) and be a button directory with several subcomponents in it, like each top level heading. Less is a serious language that you should use the semantic features to organize your code instead of just comments.
> So you have a class that has a bunch of getters and setters. Let's just assume that "generate them automatically" is not an option. You want to make it really easy to see the part of the class which is getters, and the part of the class which is setters, and then skim past that. How do you do it?
You put it into a getters file and import it into the parent class. Almost every language has features for this. Or you put each attribute into its own file, if any of the getters and setters has smarter logic than just x = parameters[x]. Ideally you build classes that don't have so many attributes that it's difficult to scroll past the getters/setters in the first place - N > 8 is a significant warning sign the code needs to be split unless it's a configuration class or equivalent.
> So you have a file that defines 3 data structures. Each data structure has a definition, a bunch of functions for parsing it, and a bunch of functions for serializing it. The author suggests that you split the file into 3 sections for the types, with subsections each for the definition, parsing, and serializing. How would you do it? Let's say the language is Rust or Typescript.
It should definitely be in 3 files, possibly 3 folders, in Typescript: (/thing/index.ts, /thing/parsing.ts, /thing/serializing/json.ts) It's so marvelously easy to import things, you should be using modules amply to split up your code. Obviously the 10-lines-of-import-for-3-lines-of-code is too much, but seriously, imports are easy and cheap.
Wait, what?
So you can do this in C/C++ for sure. #include is not just for imports. But Java? C#? How!
Ruby has mixins and Python supports multiple inheritance. Go encourages composition over deep nested hierarchies.
Many main stream languages support this kind of functionality.
I happily hack every day on a project that is self contained in a single 12K line (and growing) file. For me, splitting into multiple files would have negative utility. Everything is essentially in one place. I can find anything I need for my project with '/' search very quickly in vim.
My style is obviously not for everyone but it works great for me. I programmed for decades with traditional file splits and only in the last few years have I switched to single file. I have little interest in going back. For me, it is liberating to stop thinking about directory layout entirely. It also helps me to use simpler tools (I only use vim with no plugins) in part because I don't need help managing multiple files.
Also I don't envy you reviewing the diffs when you add something that affects e.g. indentation on the file.
Should the OP have written a manuscript to make you feel better? In your mind, should a long, drawn out article only be refuted by something of similar length?
It's not about "feeling better"? It's that the comment is very dismissive and sums up to "just refactor". Advice of "just refactor" (or "just do x") is lazy and ignores all context.
Nope
> only be refuted by something of similar length?
No, but if you don't address the points that have been made, you're not refuting. The comment says "refactor your codebase to put the related code together", but the article already addresses some downsides to this approach.
Does the commenter believe that those downsides are more avoidable than the article states? Or maybe they believe the downsides are dwarfed by the upsides? Or maybe something else. We don't know, so we can't evaluate the position effectively.
> Refactor your codebase to put the related code together, and then this "problem" will disappear
But how do I do that exactly? What does this even mean? I could literally put the whole codebase into one file since it is all related somehow or another.
Yup. We all are. This is only a problem because editing raw plaintext code as single source of truth is a bad idea, and we're reaching its limits. "Split or don't split" is one of many holy wars that can't be solved, because they happen on the Pareto frontier. The only way to move forward is to accept that different coding tasks need different representations, and the computer should synthesize them for us, and we generally should not touch the underlying single source of truth, anymore than we manually poke in assembly files generated by our compilers.
The alternative to plain text is programming-language specific version control and source code formats, and managing codebases using some external database tool, or something similar. It also means you can't easily interoperate with existing plain text code.
I predict the only way people will switch is if there's some overwhelming competitive advantage in storing code as something other than text files, to the point where business competitors will go bankrupt from not upgrading technologies. And then the open source community will adopt it after the industry does.
It’s a shallow criticism by HN standards.
It's definitely worth trying to have your code neatly organized in a way that maps to some clean conceptual model, but doing that well is going to require using every tool at your disposal—which includes ways of grouping code that don't need the sort of crisp definition and conceptual coherence of an in-language abstraction.
Sections and sub-sections seem like a reasonable way to do that with our modern, painfully limited tools. If we weren't limited by needing everything to live in plain text, I'd also reach for something like a system of tags.
In those cases, and in the cases where a single large file just makes more sense or is unavoidable (IaC/build scripts, shell/sql/migration scripts, config files, etc.), the author's methods are absolutely valuable.
> or is unavoidable (IaC/build scripts, shell/sql/migration scripts, config files, etc.
This is a failure of the tooling in question. We shouldn't accept shoddily built tooling as an acceptable justification for unreadable codebases. Obviously, there are tradeoffs and sometimes you just need to get the tool out the door, but when a particular tool (like a professional IaC project) becomes some people's primary programming languages, then treating those languages like second class citizens who "unavoidably" are going to just be awful to work with is just limiting the growth of the tool as a whole. Terraform modules are a great example of this—all that ceremony and hassle just for what is effectively a single function call! If Terraform had single-file modules and a better import system than "relative file paths", I'd expect we'd see a lot more Terraform code bases with smaller root module files.
In a related idea, I wonder what organizing code imports by one or more "tags" instead of import paths would be like. Probably too confusing after a while (tag systems tend to get that way), but fun to daydream about
But of course there are lots of other factors constraining the way a file can/should be ordered, depending on language
In our (Swift) codebase, this typically translates into "publicly-consumable interface up top, private and internal functions down below".
That said, staying on topic with the article, I think the times where I have a difficult time following along are when it's just a monolithic function or class, and especially if it has implicit behaviors via decorators or mixins.
I like this approach to organisation because it helps avoid dependency cyles. https://fsharpforfunandprofit.com/posts/cyclic-dependencies/
I once had to try and understand the entirety of a subsection by hopping between files that mutually depend on each other, which was quite confusing. UML diagrams I drew didn't help my understanding but drawing a dependency tree (the functions called that depends on no more are leaves and functions that call other functions are nodes) before I knew of this technique did help.
This does have its disadvantages though so I think it's a good default in some cases but one should be given the choice to break it when needed.
This is why open-source code tends to be some of the best architected and documented code out there: it's pretty much the definition of "by committee" (in the best sense of that term) and meant to be inviting to anyone to contribute. So of course you want it to look nice.
At BigCo, these factors don't really come into play (even though it would probably be better for the company if they did). There's also a lot of differences between writing software and building robots. Systems engineers build and refine requirements documents, pages upon pages of "this tiny part should be built like this, should have these constraints, etc." with traceability, discussion, and clear deliverables. Software, on the other hand, is unfortunately mostly just "patch a few things together and get the login form working."
somefunc = ...
where
subfunc1 = ...
subfunc2 = ...
otherfunc = ...
where
subfunc1 = ...
where
subfuncA = ...
subfunc2 = ...
Then the reader can instantly tell top-level concerns from implementation details, and the amount of jumping-around is reduced since those subfunctions are limited in scope.When your cpp file is so large that you begin losing debug info in upstream tools like Sentry, you should probably split it.
Related, when your functions are so large that they fail to compile under higher debug modes for emscripten (too many locals!), you should probably split it.
(send help)
“;” cycles through headers.
Never understood that “more import lines than lines of code” style too.
Comment anchors and similar are close though.
https://marketplace.visualstudio.com/items?itemName=ExodiusS...
class Something
def execute
first_method
endprivate
def first_method
second_method + 1
enddef second_method
2
endend
This should have been just one method instead of thousands of tiny one line methods. I blame Uncle Bob and that whole group for brainwashing the Ruby community.
You technically can do this in plain Java, but it comes with many caveats and probably won't work with a default build setup. And no, nested classes are not the same thing.
They actually are if they're public static. Quite literally so, nested static classes are the exact same as "normal" classes but named `OuterName$InnerName.class`