Code colocation is king
koenvangilst.nl
koenvangilst.nl
So, really, the only real way to keep functions linked is to not be lazy, and try to have proper documentation locally in your function which includes a "see also" bit.
Agreed. I do that, but never got around to factoring out a one-liner that expresses it.
I also tend to have fairly deep directory trees that tend to reflect the code hierarchy/structure.
I will also factor out large chunks into standalone subprojects, and reinclude them as packages. This results in very high quality code, and also gives me dozens of libraries of tested, ship-quality, reusable code.
Another thing that I do, is have fairly large source files, that aggregate multiple classes/structs/enums/protocols that relate to each other. It drives me nuts to deal with the typical Java "Each class has its own file -no matter how small" thing.
I practice what might be called "eXtreme documentation. I rely on Jazzy (sort of like Doxygen), and now, DocC, so every element in my code has a headerdoc. I make heavy use of the "// MARK:" macro, as well.
I write about that, here: https://littlegreenviper.com/miscellany/leaving-a-legacy/ (I need to add a DocC update).
1) Pretty complex, and fairly insular; and/or
2) Possibly useful, elsewhere.
If that's the case, I will then stop work on the main project, and take some time to extract and "genericize" the subproject. I'll usually set it up as a standalone open-source project; complete with tests and documentation. As the commenter stated, I think that this results in some excellent code. I always clean the house before the guests arrive.
This may happen before I have completed the coding in the main project, or may happen as the result of a review, after the fact.
In some cases, I very clearly need to develop a subproject before starting on the main project, or before certain milestones within that project (for example, SDKs or drivers). In that case, the timelines are completely separate.
If you look at my GH repos, you'll see a whole bunch of these projects, including some rather strange ones, like an XML duration parser[0]. These are the types of projects that I extract.
In some cases, I end up not using the extracted project in my main project (happens to some of my UI widgets). In that case, even though I am not using it, I still have an excellent project for the future. Here's an example[1]. I have ended up not using the spinner in my own work, as it was too obtrusive a widget, but it's nice to have it available for future projects.
[0] https://github.com/RiftValleySoftware/RVS_ParseXMLDuration
Couple this with extensive use of inheritance and design patterns, and you have a recipe for awfulness. One of my previous teams had an implementation where rendering a piece of HTML would involve digging through about twenty different source files, each with maybe 3-4 lines of actual code other than the class definition boilerplate. One top-level line would send you down the inheritance hierarchy of ClassThatGetsData/ClassThatGetsDataPlusThisOneOtherThing/...PlusThisOtherOtherThing/... etc etc, then the same thing for ClassThatProcessesData/... and then ClassThatRendersData/...
Adding a single bit of data to rendered HTML (e.g. a tiny star for 'this product is well-reviewed') would involve altering every single source file in the multiple trees, and maybe adding some new specialisations to make the trees even deeper.
At the time, fresh out of Uni with a head full of design patterns, I thought this was just how enterprise code was meant to be structured!
Same thing for algorithms: BloomFilters are the classic example. You can tell when it has hit the font page of HackerNews recently because suddenly everyone starts looking for problems that it can solve (poorly).
Just don't get me started on blockchain... :-)
Ugh... and that's the flaw in teaching design patterns.
Don't worry, that happened to me too, though I didn't go to a university.
Design patterns are an advanced tool. If you're learning to code, design patterns are mostly non-applicable to the kinds of problems that newbs are solving. Design patterns are meant for specific problems, but somehow we end up believing they're the end-all-be-all and that we must be using some design pattern in our work. If you have design patterns pounded into your head, and you believe the only design patterns that exist are the one's someone has already named, then not using a particular design pattern can seem like chaos even when it's not.
I wish we'd stop teaching that shit. If you've been coding long enough and faced complex enough problems, you'll either come to embrace some design patterns or you won't. Otherwise they'll likely just be misapplied.
OOP fits this view of mine as well. In general, I don't think most programmers need to be aware off OOP principles because they will almost certainly misuse them. And to what end? Several single purpose classes and "tiny functions" scattered across multiple files that others are now forced to jump between.
In other languages or modern Java you can replace half of them with simple constructs. Factory pattern for example can be replaced with simple inline callbacks to anonymous functions. Same with many others.
There is value in keeping some source together, but I think it's more judgment based than you imply, it can also drive me nuts when I have to dig through a dozen 2k+ line files to find the thing I need to change, and it also reduces my confidence that I can change that thing without introducing unwanted side effects (unit tests help increase that confidence regardless of code structure).
As someone that regularly digs through 200-1.5K files (2K is a bit much, for me), I can report that good documentation makes a huge difference.
I tend to use a lot of test harnesses (which, IMNSHO, are much better application exemplars than most unit tests).
I also use a lot of headerdoc stuff. In Xcode, it allows your own code to show up in the QuickHelp navigator tab, and you can generate really good SDK docs.
It also stays fresh. Easy to maintain. Using "breakers" is a huge aspect of my code. Makes it much easier to scan, and liberal use of // MARK: is good.
In my experience, there’s no substitute for good, old-fashioned, disciplined “monkey-tests,” when it comes to UI software. Even test harnesses aren’t always relevant, and I end up running most of my tests on the production code, as the project matures.
Requires discipline. And, in my experience, a well-documented codebase (“Why,” as opposed to “what,” when documenting internals), pays off in spades. Headerdoc markup works wonders. It, quite literally, makes self-documenting code, and documenting the interfaces (as opposed to the internals) has a great shelf life.
Since I use Xcode, it "live-parses" my interface docs, and shows them in the QuickHelp panel, on the right side of the screen, so, when I select a function that I wrote, it displays the interface, just like it does for the Apple-authored stuff. Very cool. Since all my included packages (the ones I wrote, anyway -which is 99% of them), use it, I have great "live" documentation.
I used to argue with my Japanese peers about this stuff. They were not fans of automated testing. I was able to get them on board with unit-testing engine code, but they were absolutely correct about testing GUI code, and I had to cede the point to them.
They were very disciplined, when it came to “monkey testing,” code structure, and documentation. One reason, was because they tended to rotate engineers through projects all the time, and leaving a legacy was important. I write about what I have learned, here: https://littlegreenviper.com/miscellany/leaving-a-legacy/
If I've got a structural outline panel, like in IntelliJ-based editors, or the SuperCharger plugin in Visual Studio, then whatever, it's easy, because I have that table-of-contents to easily navigate with.
With less powerful tools, the "every type in its own named file" pattern is much more useful.
I use Xcode’s editor. It makes zipping around big text files, quite easy, and going through multiple text files, a pain.
This can go both ways - I've seen plenty of projects where you have to dig through 3k lines of unrelated spaghetti code just to find the bit you need. That turns into a special piece of hell when you have multiple team members working on it.
The original motivation for having more small files was driven by these "god objects" and the limitations of version control. CVS, Visual Source Safe and to a lesser extent Subversion were much easier to use together with smaller files.
If you're using modern version control (git) then that reason has gone.
My personal preference is to split horizontal concerns into their own modules and depend on them; they become the core / support libraries for my team/organisation. For services I try to keep the domain-specific parts close. Tests go in a different module which mirrors the structure. When units start to get big (or become cross-cutting) they get factored into smaller units or other modules.
In the end it's all a balance and unless you're under outside constraints (cough SonarCube box-ticking busywork idiocy cough) then you can generally keep it sane.
And it also seems notable to me that Rails (which I generally like; I am not a Rails hater this is not a Rails hate comment) -- often seems to be trying to do the opposite. Like separating a controller and it's view code -- things which in the typical Rails app are pretty tightly coupled -- into the higher level controllers and views folders, putting all controllers next to each other and all views next to each other, instead of a controller next to it's coupled view. (which I do find an inconvenience as a developer).
But the risingly popular view_component library for use with Rails makes the 'proximate' choice, putting the view template(s) next to the logic it's coupled to.
I think I find it easier to find common abstractions on the {models/views/controllers} level than the resource level and that frustrates me with Django app-based dev.
But controllers and view templates almost always do, because of the nature of the architecture.
So actually just using the `view_component` gem for all my views from now on (not just partials) probably satisfies me!
In fact, the idea of trying to model complex systems in a text format divided in files (most programming languages) doesn't quite hold... gracefully at least. For example, the frequent discussions about inheritance and generics are pretty revealing of the fact that we mix modelling and implementations in the same working space, when in fact in many cases it would be better to work on those at different layers.
So, in my opinion, to really make "code colocation" better you would kinda need to start modelling complex systems with richer toolsets that don't try to express them only with code files. You can't properly work with complex systems with a single view, no matter which one you pick.
But if you take a step back, they are actually not the right abstraction for most things. In this case, a tagging system would serve the purpose much better.
Unfortunately, tooling that deals with files&folders and text is very mature and it will be hard to extend or even replace it.
It is quite useful to find find problem areas and tacit knowledge such as “when you change this API these two client side adapters should probably also be updated”
There is a lot of separate innovation that could be combined; so I think.
While inspired from these various things I doubt it would just be "another smalltalk" or "just a nocode thing" which many modern versions of this end up looking like.
If you have a partial/catalog/product/buy.js file, have a partial/catalog/product/buy.scss file as well. Finding the CSS for the buy.js component inside add-to-cart.scss is very lame.
So true. Small things like this are why I'm so glad I had a software engineering job at a real company for a few years, even though I don't necessarily want to do that for life. I learned a lot from looking at codebase conventions and quick questions to the elders. Unfortunately, no one can help you with the other hard problem: naming.
It also means I don't have to worry about where Dev X put thing Y, because I know where to look for it already.
That's what software patterns are good for as well, a shared set of idioms so you don't have to invent new approaches. I think that's important in a professional environment but less so in a personal one. For personal projects I just do whatever is fun.
I have a guiding principle here: If, a few minutes after naming something, you want to use the thing and (without looking it up) your first attempt at writing the name is correct, it's a decent name.
If you name a function processFoos and then later in your code your first intuition for what it was called is createBars, then there is some disconnect (and possibly some missing conceptual clarity) in what this function is supposed to accomplish. Is it more about accepting foos or is it more about creating bars? I find it very productive to, in such a case, dive a bit deeper into these differences (which may often also reveal something about different perspectives between caller and callee).
So TL;DR: If you don't have to look up the name of what you wrote minutes ago, it's probably a good name.
One thing that gets lost is the application-specific knowledge, which is the whole reason you are doing things. Yes, your database is now near-optimally managed according to $vendor, except it's a poor fit for the application, as it received a one size fits all configuration.
Second problem is velocity. If 1 team can adapt the database structure, modify the backend code and services, you'l go a lot faster than if you have to book a databaser for a week, a developer for 2 weeks, etc...
Next problem is implicit waterfall: The developer will hide data in the wrong columns, because otherwise the databaser has to be called back, which causes rescheduling and a rebudget( i.e. management now hates you). It's only temporary, everybody tells themselves, until the right person revisits the application again next month.
And god help you if the architect did not deliver perfect work the first time around, because then everybody is creating things that don't fit together. The architect being of course the guy/girl who drew some boxes and arrows between 2 very important meetings, at the point where business requirements were not yet delivered and nobody knew what the application actually needed to do. Architects are expensive, so they're long gone before the first character of the code is ever written.
Final problem: Your application is now spread over 10s of silos, and nobody knows what connects to what. Any maintenance done on it starts with someone walking around between silos, asking everybody what their piece of the puzzel is.
So my opinion is to do the reverse: Let the dev team modify the database structure, even if the result is clearly suboptimal. They will feel the pain from their mistakes and have the ability to fix things. It might take a while, but a coherent team will figure things out. A loose bunch specialists, available part-time? Not a chance.
IMHO, over-emphasis on efficiency is a common trap; efficiency and agility are at the poles, and it's the latter that often matters more.
So, how to strike a pragmatic balance? I've been part of successful experiments with a small, free-roaming, multidisciplinary "red team" or "green team" that crossed org boundaries to solve gnarly problems, free up log jams, and facilitate step-function improvements in standard teams' capabilities.
> "Architects are expensive, so they're long gone before the first character of the code is ever written."
As a hands-on architect, I don't consider my work complete until there's been meaningful collaboration with dev leads and in-depth review of working code. It's a dynamic and iterative process. Immutable, ivory tower boxes-and-arrows -- divorced from the realities of actual software development at any kind of scale -- are insufficient. A prescribed, linear, waterfall process of biz req -> architecture -> implementation is bound to fail. Success requires embracing the rich interplay between various forms of software design (requirements, IX/UX, architecture, and implementation) and stepping in and out of them where and when appropriate.
The pragmatic approach I've taken to deal with this is make heavier use of custom context managers so creation and cleanup code must stay together.
But I'm also getting tired of whitespace sensitive languages.
The closest I can think of is:
if something:
...handle something..
elif something_else:
...handle something else..
if other_thing:
...handle other thing...
elif yet_another_thing:
...handle yet another thing...
else:
...for everything else...
If I had a frustration with python it would be the prevalence of "who needs types when we have dicts". This situation has been somewhat improved by dataclasses but lots of our code predates them by many years.1 - Last time I asked "who doesn't" on the internet, I found somebody who doesn't. But it's by far the most popular option.
The whitespace works differently between those languages. The Python syntax has limitations that aren't on the Haskell one.
But also, Haskell tends to place the kind of things the GP is talking about on the beginning of your expression (because of the "reverse" order on function composition), while Python tends to place them at the end of a block.
http://williamtpayne.blogspot.com/2012/07/structure.html
http://williamtpayne.blogspot.com/2012/05/development-concer...
Looking back on this now, it all feels a bit unconventional and strange, so I wouldn't make the same recommendations today. I think that people want to operate in an environment with fewer, and less rigid constraints. I do think some of the driving concerns are still valid though, and still worth considering.
A couple new problems we're dealing with now:
1. Finding older common things and deprecating them. When it's in a common space it feels like it's everyone's responsibility which means nobody ends up working on it. Maybe more narrow, clear ownership would solve that problem.
2. Someone finds the common thing that almost fits their need, if only it had one more little feature. The problem is when this happens many times and you end up with this complicated beast of an abstraction. This is probably solved by finding ways to decompose the abstraction and by following a principle of "do one thing well" or something about simplicity.
CSS-in-JS and Tailwind greatly helped here to encapsulate things.
Yes: ./User.ts
It's also worth noting that there's likely not one single "correct" organization/taxonomy, and certainly not one that's static and guaranteed to last forever. It truly is one of the most complex parts of programming because it's subjective, subject to change and is in a lot of ways right through the center of what it means to abstract something in the first place. I don't think there are any simple answers because it's equal parts philosophical and mundane.
Software mixes things that model the real world, convenient fictions, abstract truths and guilty hacks and I believe trying to make sense of it is at least as much an art as a science.
Random thought: Can the IDE recommend next file to work on - “people who worked with this file, also worked on”. The real problem is not where we keep it, it’s how quickly we can get to it.
But it's a nice idea to keep unit tests specifically alongside the "units" of code that they test.
I'd much rather see the 'mirrored' approach.
> Random thought: Can the IDE recommend next file to work on
I'm _sure_ I've seen something like this in IntelliJ.
EDIT: I did! It's in the changes tab > the eye icon > 'show files related to active changelist'. It then sometimes will make suggestions based on project history.
It is largely because of this reason that I disdain patterns which separate things out into top level folders like /views, /reducers, /actions, /utils. Most code should be organized at the feature level with separate modules for global utilities (which are always hard to organize).
I think this methodology works well because more files means more mental overhead and context switching, particularly for code reviews. The best code review given a reasonably sized change (not 20kloc) is a single file where the file contains little not relevant for the changed code. Not only is it easier to comprehend, it also avoids unnecessary merge conflicts.
No piece of code is so bland that it can’t be placed somewhere with a name that describes what it does or what category it belongs to.
The difference is stark when I have the lean, more-or-less functional core business logic library sitting beside this baroque stringly-typed reflected 7-bean salad of a web application shell.
Code can and should be aligned to the business (sub)domain(s). It is much harder to do on a frontend, but possible nonetheless.
if your codebase is 3 features each of which has 3 similar functions (listener, db model, middleware, let's say)
do you group by similar functions? (middlewares.py, models.py)? or do you group by feature? (feature1.py). depending on whether 'work' is a feature change or a refactor this week, the correlations will be different
If you package logic according to the process, it's a lot easier to control the scope of logic that is exclusive to a given process (some things are of course shared). Languages that can enforce package privacy, file privacy, etc. help a lot here.
In general, unlimited scope & dependencies is what strangles a dev team to death, so any option that restricts scope to what is necessary is arguably best.
I think you're right that that's one of the key ways that devs (esp noobs to the codebase) need to navigate, and must be supported. I think 'stack graphs' are supposed to do that[1], but have never tried them.
but there are other navigation modes we need to support as well -- like if you're deleting something, or upgrading something, 'find all references'. if you're inventorying use of a DB type, you might want to audit field types in all models.
I sometimes wonder if the future is tagging code as 'belonging to a feature', but they can live wherever -- because devs have different needs on different days.
The best option I've found for "find everything that references this database table and see if it will be affected by this thing I need to change" is to give database tables really unique names (like "tbl_user" instead of "user") and just grep the hell out of everything.
Also if I'm writing large programs I'm going to use a compiler. I mean if a linter can tell you "Hey this python won't work because you removed this function" then good for python, otherwise bad for python. And if I'm going to include 100 open-source libraries then god help me if my build system can't tell that one of them went missing or that the author removed a function I'm using in version 1.2.3.4.b.
But all of that is about dependencies, and dependencies kill teams. That's why I want the scope of any logic as tightly constrained and enforced as possible.
I'm sure you could design a multi-dimensional behavioral categorizer/IDE/language system that blows everybody's minds and wouldn't that be cool, but I doubt anyone would use it, because behavioral categorization really doesn't solve problems. It just looks nice.
But any rate, at the very minimum, back to this point: Don't try to compensate for the failure of behavioral categorization with microservices when you could have just packaged things along the same lines of division.
interested in 'library versions as types' and have been thinking about this problem on and off. version numbers are a proxy for the combination of: 1) call signature, and 2) internal semantics
linters / typesystems are good at (1). If we had a system that could do (2), version numbers would be less necessary and compatibility could be proof-based. (Well, to the extent that the semantic assertions are valid).
>Turns out that "where to put code" is one of the hard things in software engineering and there are no silver bullets. That's part of the reason why there are so few easy tutorials on this subject.
This is called modularization and there are lots of papers and books that discuss it in decent detail.
This is essentially "naming things" (from "There are only two hard things in Computer Science..."), applied to namespaces.
I don't know too much about compilers, but if you're improved proximity results in better proximity of the byte code then you're less likely to have a cache miss. That's a big deal.
If A calls B, then A folder should contain B folder.
:-D
cd a
ln -s . a
Problem solved; perfect organization.gcc:
a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/test.h:1:10: fatal error: a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/a/test.h: Too many levels of symbolic links
python: ModuleNotFoundError: No module named 'tmp.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a.a'