Why do we need modules at all? (2011)
erlang.org
erlang.org
1. Start by questioning a long held assumption ("Do we need modules? Why? Is it because we stored functions in files?" Etc...)
2. Consider alternatives / what ifs. ("What if we used a database instead of files?")
3. Think through implications of the alternatives ("How would we identify functions?")
IMHO, this is a powerful technique for generating new ideas. Step 4, not explicitly performed here, might be to take ideas generated from steps 2 & 3 and apply them to other domains. ("Ok, perhaps a database of functions wouldn't be great for code, but perhaps it (or something similar) would be great for some other domain...")
I think codeq 2 is being worked on
It also looks like this is only suitable for functional languages, and possibly just non-OO languages?
What about the type signatures of these functions? What do they look like? How well can different types be substituted? Do we have to version the types as well?
How do we handle testing in this scenario? Every version of every function? How about combinations of functions?
How about deprecation? This is hard enough to do with standard library functions; how long has gets() been deprecated, for example?
When integrating across maintenance boundaries, renaming isn't possible but a simple prefix is almost as good as a namespace.
I always wanted a code completion tool which I could ask: hey, I have these two values, say, one field and one pango layout, but forgot what to do next. And it would figure out that I want
pango_layout_set_text(
layout,
gtk_entry_get_text(entry),
-1);
among the others. Soo cool.If you look at the /usr/bin directory of a Linux system there aren't any namespaces. There's 1,000+ binaries thrown into 1 folder, each with a distinct name and things work nicely.
Imagine if you tried to group those binaries into 1 or more folder driven namespaces. Where would things go? The amount of bike shedding would be endless and I don't think it would be a better solution in the end even if you managed to pull it off.
All of Joe's points in that post are spot on.
That's the last example of a nice design I would bring, and the worst possibly argument in favor of no modules/namespacing.
A flat /usr/bin/... with 1000+ binaries in UNIX is a bad solution we ended up with because back in the day where UNIX was created there were just a few dozens of binaries, so it was "good enough".
What's actually wrong with it though? Why isn't it still "good enough"? It's hard to read the list, but Linux has plenty of tools to fix that particular problem. I don't really see the issue.
Name clauses.
Version clauses. "All binaries inside a folder" needs renames to allow for a different version (even if it's a different minor version).
It's just the binaries, which was OK back in the "program = single binary", before "man" even, days, but now it splits the program from different assets it uses (man pages, default configuration, image assets, etc). This in turn makes deleting/moving more difficult, and adds all kinds of baggage to package managers.
>It's hard to read the list, but Linux has plenty of tools to fix that particular problem.
That's a description of the problem to me. Needing "plenty of tools" to fix an initial bad commitment.
I see it as easier to discover.
Imagine if you had to goto /usr/bin/output/user to find the whoami binary.
If you did, who is responsible for making those directories and what happens when you have a binary that could really belong in 5 different places?
https://zolk3ri.name/cgit/zpkg/
It seems very minimal but supports my requirements. It makes sense to me, because everything gets installed (usually via make install) to ~/.local/pkg/foo-1.0/ and the like, and then symlinks are created to, say, ~/.local/{bin,etc,include,lib,man,share,var}. According to the source code, it is configurable via environment variables. It seems to share similarities with GNU's stow.
The other issue is that the approach is extremely inflexible, if you want to replace /usr/bin/program with a new version, everybody on the system is forced to use the new version, even if the new version might break stuff. Traditional `/usr/bin` doesn't provide a solution for this. You can of course work around that by using `/opt` or installing things in your home directory, but at that point you are just using the file system as namespace.
i think the solution to this problem (module or not) will be the same solution to the shared lib problem.
npm is a dumpster fire, but the larger projects are usually fine (the /bin equivalent). the worst problem is that every large project pull in different versions of whatever libs they wanted to use (which the singleton util lib proposed in the article would solve)
Anyone who has had to try to manage security permissions on a big flat list of files knows it's like holding jello with rubber bands.
As an another example people might be familiar with: having to manage a big PHP or Ruby codebase where everything is separated by the models, views, and controllers a hierarchical level _above_ the business function. It's a nightmare. Trying to pull out the useful business functionality that's interleaved everywhere makes it very hard to pull codebases apart to where different teams can work on them.
With respect to Joe's comment here:
> Bad: It's very difficult to decide which module to put an individual function in. Break encapsulation (see later)
As much as I like Joe's thought process, that "difficult" part is otherwise known as "the real work". All systems which scale in complexity--as opposed to the simple problem scaled in performance--need strong boundaries. They also all have a hierarchical component to them.
I still don't fully agree with it but it makes some things easier. The Windows way also made multiple version installs possible.
And command line tools have that folder added to the PATH, and Windows has a real and ongoing problem with PATH getting too long for the OS to handle when you have enough of them.
Windows is GUI optimized with shell as an afterthought, Linux is vice versa.
It isn't very uncommon to see binaries inside /usr/bin being symlinks to binaries in other directories, even if you stick with distribution-only packages.
NixOS and distri are notable package managers that do make installing different versions of packages pretty easy.
Also i will add that in fact the systen binaries are already split: at minimum /usr/bin and /usr/sbin (but also historically, /bin and /sbin although i think in recent linux system the last two are just symlinks to the formers, or the other way around)
Seriously though, considering the CLI software I use occasionally on my Linux machine, for any given program one - and only one - of the following will tell you its version: -v, -version, --version, version. I never can guess in advance which one is the right to use.
With functions it's worse. In a flat namespace you will inevitably run into a problem of two people coming up with the same name.
IMO python and python3 is a very welcome feature not a negative side effect of no namespaces. At call time it's super explicit so you know exactly what version you're using.
But in most module driven programming languages you might import the foo function at the top of your file and then use foo in 5 different places within that file. However, you spend 99% of your time working with the code in the file not glancing at imports, so now you're left wondering not only where foo is coming from, but who provided foo. Is it from the standard library, your own code base or a third party author? Suddenly you need to keep all of this in your head and it sucks.
Phoenix (a web framework for Elixir) has been taking steps to remove a lot of loose function imports so you know what module they are coming from at call time, but that's because Elixir as a language has modules. But even still, explicitness where you're using it is so much better in the long run for maintenance, even if it involves typing a few more characters.
As always with dynamic things and editing code for them, this will be more difficult to get right when imports are dynamic.
That's a versioning issue, that would happen with namespaces too. You'll have to disambiguate somehow and if a version number is the natural way, then that's the natural way regardless of if a single name or namespaced.
Keeping the next incompatible version as a separate thing (eg python vs python3, instead of continuing python but with semvers saying that 3.0 breaks 2.x compatibility) is actually a good idea. It leads to fragmentation, sure, but it makes dealing with the different versions easier IMHO and incompatible versions may as well be totally different things anyway.
> With functions it's worse. In a flat namespace you will inevitably run into a problem of two people coming up with the same name.
Yes, completely agree!
Microsoft COM would like a word. {20D04FE0-3AEA-1069-A2D8-08002B30309D}
I'm not even sure that reinventing COM would be a bad idea - you can definitely do some great things with it, and the specific implementation details of COM are crufty. But doing so deliberately rather than accidentally would be better, with a review of benefits and problems of COM approaches.
As for the community, maybe not on SV, but there are plenty of Microsoft communities and meeting groups.
https://developer.apple.com/documentation/corefoundation/cfp...
no they don't
I can see it for core libraries, but outside of that I think namespaces are necessary for the greater language ecosystem.
I think the main reason PHP got such a bad rap for flat functions is because they implemented it poorly. So many things were extremely inconsistent when it came to naming conventions and argument orders that it became near impossible to remember them in a systematic way. But couldn't that problem be addressed today since we're aware that naming consistency is an important programming language feature?
> I can see it for core libraries, but outside of that I think namespaces are necessary for the greater language ecosystem.
Yeah I could get on board with that. Docker sort of does this now with Docker images. At the UI level (the tooling we use in our day to day like the CLI) we don't need to access "elixir" or "python" images with a namespace (Docker adds it for us automatically), but we are free to create our own brightball/elixir or nickjj/elixir images to avoid conflicts.
But in a programming world, I would like to type brightball_elixir at call time in my code instead of importing elixir from brightball, because with the latter I have no idea if elixir is coming from you, me, or the standard library unless I move away from the code and look at the import.
That type of stuff. The core language was about the same regardless, but the ecosystem improved dramatically.
To be picky: module and namespace are different concepts.
C++ has namespace and (soon) will have namespaces AND modules.
Namespaces can be seen in C++ as a pure hardcoded prefix for your functions and are a perfectly valid concept with Joe's idea of a "giant KV store for functions".
When Joe Armstrong proposed this, I thought, "Hm, that's an interesting idea." But I think NPM has shown that it isn't such a great idea in practice.
I.e. the release would still contain a curated, officially supported set of function that a) work together and b) are reasonably sure to not execute miners, not display ads and not post your environment variables to pastebin.
Maybe it’s self evident to you that this is a horrible thing, but it’s not to me.
But if you care to share, I would be happy to learn about the horrible things that will befall me, before they happen!
I recently published a tiny library on NPM for changing the case of strings (snake_case, etc.). It felt so small and silly that I was tempted to add some other string manipulation functions, to make it a more "serious" library. But it struck me that this would just be silly. How often do you need a whole collection of string manipulation functions? Never! It would just add to the size of dist folders everywhere, with no gain whatsoever. I suspect this applies in many situations.
Always! The Javascript world lacks a good standard library or something like STL that covers most of what you need during normal development. Instead you have thousands of little packages of varying quality. I would much prefer a few big libraries with more power and consistency.
I mean that's where proper tree shaking would void this argument. When you include lodash (or lodash-es more precisely) for instance only the used functions are in the final bundle with a modern bundler.
All the time. It’s called a standard library and most sane languages have one.
Not initially, but as your code grows, you start to call out to more and more of those string manipulation functions... and if each of them is its own separate package, you get an inconsistent mess at the interface level - just like PHP used to be, except PHP didn't consist of individually downloadable functions with much more metadata than actual code.
He does discuss information hiding in an Erlang context, so he didn’t completely ignore your concern.
I’m not sure why you felt the need to (tonally) attack him.
> To avoid all mental anguish when I need a small function that should be somewhere else and isn't I stick it in a module elib1_misc.erl.
I can totally relate. It doesn't reduce collisions significantly if a function is named Foo::Bar() instead of Foo_Bar(). But the latter has the advantage that we are free to move the function (avoiding initial decision paralysis, and later it's easier to clean up) and being required to reference this function with fully qualified name leads to much more readable code (IMO).
What I've been doing lately is I just make descriptive function names. I don't even use a prefix, but may still include type names in the function name. Example, slightly long name: copy_from_StreamsChain_to_Textnode(). That's totally fine in terms of avoiding (the FUD concerning) collisions and it flows very naturally.
Just like with Objects/methods, not doing namespaces leads to a significant reduction in decision paralysis (time consuming taxonomic concerns - "where should I put...?") for me.
- If "where should I put" is unclear, "where should I find" is equally unclear.
- Code should be mostly "local" anyway (what's the word again?)
- Text search exists, and IDEs nowadays even do fuzzy search.
- I didn't say we should make a total mess. If there is an obvious organizational strategy, just go for it! But namespaces are not that important; you can still get most of the organizational advantages of namespaces by just grouping in files and maybe prefixing.
Obv. There are issues with this approach if fully automatic (how many sum functions take two ints and return int?) but there's no reason why we can't use a sig/return to filter on ie. The same function can be in 2 modules if filtering by sig or by return type.
With typing similar to Haskell we can be even more strict - the sig/return is a contract, and we can have multiple types like userInputInt, randomInt which can differentiate similar inputs.
With good typing and behavioural provability we may be able to move towards a call-by-meaning system
We use RDBMS to manage large quantities of physical objects, so why not also use them to manage large quantities of code snippets, such as event handlers. Early databases were hierarchical also, but over time hierarchies ran out of steam. TOP believes eventually the same lesson will be learned about code management.
Take MVC for example. Some code maintenance tasks are by entity and not just M or V or C. For instance, adding a new column to a given table. So if one could run a little query to list links to all modules/files/snippets for a given entity, then entity-based tasks are easier.
https://www.youtube.com/watch?v=kc9HwsxE1OY
f(x1,x2){} < x1.f(x2){x1} < f(x1,x2){x1,x2}
But modules are useful for organising code and choosing what to bother the compiler with.
Organization of code at a larger granularity than individual functions, is something that's really useful and important in most non-trivial applications.
Topicality and meaningfulness of organization is something that should be valued. Removing or skipping it because Joe can't be bothered separating his String, Text File and Process funcs would be foolishness. Resulting in only a new form of spaghetti.
And the (searchable) database is called npm.
The database / infrastructure could be permissionless and you could still define curated subsets of functions in addition to that.
But seriously, I know his description is a bit different, but I think NPM has shown that such a model is really difficult to maintain.