Why do we need modules at all? (2011)
erlang.org
erlang.org
Genuinely curious, thinking outside the box (writing a new programming language on top of Prolog?), and he treated people with utmost respect. I was a nobody lucky enough to escort him around Chicago one day when he was attending a conference, and we spent a couple of hours talking about art, Erlang, Riak, and man I wish I could remember what else.
Really? And every developer will create custom, broken, non-standard "namespace" system: admin_get_user vs. readers_get_user, or maybe admin.get_user vs. readers.get_user, or maybe get_user_admin, get_user_readers, etc. Surely, this will stir a lot of creativity, but I am not sure we need that.
https://news.ycombinator.com/item?id=8572600 - Nov 7, 2014 (76 comments)
https://news.ycombinator.com/item?id=10409507 - Oct 18, 2015 (46 comments)
https://news.ycombinator.com/item?id=20808000 - Aug 27, 2019 (103 comments)
EDIT: I also highly recommend reading the original mailing list thread. A lot of interesting discussion there that I don't recall having read before.
Purely functional on the other hand has none of these problems. This is the approach taken by the Unison Language people, which I think makes the right design decisions.
> Each Unison definition is identified by a hash of its syntax tree. Put another way, Unison code is content-addressed.
In either language class, you need to manage transactions, usually implicitly, if you want to do any meaningful work. This is a gap that either language class can easily solve and in many cases, there are working implementations of these ideas that do just that.
What problem are you actually trying to solve?
I think the idea is that if a function always gives the same result with the same inputs, it is considered “pure enough.”
myMethod :: IO ()
myMethod = putStrLn "Hello World!"It’s like a built in bikeshedding bait for the whole programming language.
Yes, this is a very strange belief
Non-blocking non-erroring send ! and optionally blocking/non-blocking receive, with pattern matching, are built in to the language.
Not a monad in sight.
That’s damning with faint praise.
Monads is a solution for dealing with IO in a strongly typed programming language, where we want all functions to be pure, and we don't have access to linear types. (almost) Every other programming language, when it had the chance at doing this, just threw the towel without a fight, and decided to allow IO everywhere. Haskell decided to stick to its principles instead, and produced something different, guided by a different set of constraints. Different means with different tradeoffs in different places. If Haskell instead decided to allow mutability everywhere, like most programming languages, we wouldn't be talking about it. Because it would be another language with weird syntax, and you can find a lot of them in the list of defunct programming languages at wikipedia.
And the juice is worth the squeeze if you want to see breakthrough things in programming, and not just another Algol with updated syntax.
For reference, other samples of groundbreaking (to some degree) stuff:
* Rust: mutability only when the borrow checker allows it
* Coroutines/continuations: functions that depart away from the entry/exit model
* Prolog: a programming language constructed around functions with 4 connection points: enter/exit/fail/redo
I think assuming that someone doesn't understand something is only rude if there are no indicators for a misunderstanding.
javcasas seemed to assume that TylerE doesn't understand Monads because TylerE stated that there is a Monad insanity in Haskell to replicate a print statement.
I would also assume javcasas' comment wasn't made in good faith, but I still don't see that an insult was voiced.
Not every rudeness can be interpreted as an insult I think
Truly insane.
Now do it inside a vanilla function that isn't implicitly run inside the IO monad.
What do you mean by run implicitly? I explicitly put IO expressions in the main function and they run. That’s the same as any imperative language.
I’m trying to figure out what you’re getting at. Are you saying it’s hard to do IO outside of the IO Monad?
> Look at all the monad insanity Haskell has to do to get the equivalent of a print statement
but it's easy to do that! If you meant something more complex and subtle perhaps you should have made a more complex and subtle claim.
> There is nothing that has to do with monads at all in printing a string. The idea that `putStrLn "hello world"` is monadic is as absurd as saying that `[1,2,3]` is monadic.
You haven't linked any exemplar tutorials, but I feel confident in saying the ones you're referring to teach the general concept of monads, and not how to do IO with monads.
Quite a number of JavaScript developers learned how to use `andThen` with Promises/A+ back in the day, and they didn't need to learn about monads either. (In fact, when it was raised that promises are really monads, there was serious drama and outrage [1] -- an existence proof if there ever was one, that you can use something productively without understanding it as a monad.)
Likewise, nobody would claim that you need to understand monoids before you can concatenate lists. Concatenating lists is as native to lists as sequencing commands is to IO; it's just part of how those types work. The monad interface abstracts over that idea, and learning that abstraction in its full generality is, typically, what people stumble on.
[0]: https://blog.jle.im/entry/io-monad-considered-harmful.html
[1]: https://github.com/promises-aplus/promises-spec/issues/94
Relative to programming in general, monads are easy to work with, easy enough to understand, but I’ll grant you often poorly explained.
Those people writing monad tutorials (those are written about as often as they are read) aren't trying to learn how to do I/O.
(And of course "it's just a monoid in the category of endofunctors".)
myMethod :: IO ()
myMethod = putStrLn "Hello World!"
Yes, so hard... "insane" /s f x y z = do
let
foo = x + y
bar = foo * z
putStrLn $ "Here's bar" <> (show bar)
pure bar
If all you want is to log to the console for debugging from pure code, it would look like f x y z =
let
foo = x + y
bar = foo * z
in
trace ("Here's bar" <> (show bar)) bar
Not recommended, but if you're dead set on commingling your effectful and pure code, you can freely mix them with the function unsafePerformIO.In my experience having to write large applications with significant test coverage, the time savings from not screwing around with stubs and mocks and dealing with the external effects of test runs -- the relative annoyance of using a couple "do"s and "pure"s and separating out data transformation from IO is a small price to pay.
> Look at all the monad insanity Haskell has to do to get the equivalent of a print statement
perhaps you really meant
> Look at all the monad insanity Haskell has to do to get the equivalent of a print statement within pure code
I still don't agree as such, but yes, there is indeed ceremony involved in turning something that was pure into something effectful. That is in fact the whole point!
main = putStrLn “foo” >> putStrLn “bar”
main = putStrLn "Hello, world"If you want to just write IO, you can just define a function with an IO () value and use it in any other function that resolves to IO (), or call other functions that live in IO *, or any pure functions, etc etc.
There is 5x as much ceremony for simple stuff in C or Java...
AFAIU the Haskell community recognized the issue as common enough to grant the existence of Debug.Trace.
That's a feature. It forces you to separate pure functions from I/O. If you spend time thinking about how you (re)factor your work you can end up with most of the interesting logic being in pure functions that are then very easy to write tests for (because you don't have to mock network services, clocks, etc.), and all the interesting I/O logic gets segregated and made [hopefully] small.
Yes, indeed you do. But it's a bit strange that that is a complaint. It's the whole point of fine grained effect tracking. If you change what effects a function does then you have to acknowledge that by changing the code that use that function (directly or indirectly)!
Yes, but that's not a "nightmare" (or whatever the parent called it), it's the whole value proposition.
Declaring a function as non-IO is a contract to your callers that you don't do IO. You don't need to do it. You can write everything in IO if you choose. You can also call into the wonderful ecosystem of libraries, because IO functions can call other IO functions, as well as non-IO functions.
There is only ever friction if you declare your function to be IO-free. If you declare your function to be IO-free, but call IO from inside it, it's a compile error because of course it is.
So why bother declaring anything IO-free? If you do arbitrary IO in a Parser, you can't back-track. If you do arbitrary IO in Parallel code, you invite race conditions. Transactions is my favourite example, though:
.NET [1]
> Disillusionment Part I: the I/O Problem
> It wasn’t long before we realized another sizeable, and more fundamental, challenge with unbounded transactions [...] What do we do with atomic blocks that do not simply consist of pure memory reads and writes? (In other words, the majority of blocks of code written today.) This was not just a pesky question of how to compile a piece of code, but rather struck right at the heart of the TM model.
Scala: [2]
> ScalaSTM does not have the goal of running arbitrary existing code, which is where most of their problems arose.
Java/Akka: [3]
> STM is considered as a failed experiment
Clojure: [4]
> Very simply the side-effects will happen again. In the above case this probably doesn’t matter, the log will be inconsistent but the real source of data (the ref) will be correct.
Sorry for the rant, but I really needed to highlight the fact that nonIO-calling-IO is not some language design flaw created by out-of-touch academics. It's a fundamental problem.[1] https://joeduffyblog.com/2010/01/03/a-brief-retrospective-on...
[2] https://nbronson.github.io/scala-stm/faq.html
[3] https://groups.google.com/g/akka-user/c/3JWz-X5dbe8/m/YiV4WF...
[4] https://sw1nn.com/blog/2012/04/11/clojure-stm-what-why-how/
printGreeting = print "Hello, World!"
main = printGreeting
works every bit as much as main = print "Hello, World"
The only difference for main is that it's run because it's the entry point, but that's true of most languages and almost certainly not what you're complaining about I think?Point being, some ways that programming languages use the concept of a module are deep and let you do things you couldn’t otherwise do, some are more dispensable.
Then there’s Smalltalk where I believe programs aren’t collections of text files, you actually browse your code in a sort of code browser and it’s stored as runtime objects in a virtual machine image, basically!
Then there’s the matter of how code is released and distributed… packaging and libraries. (In practice, the words “module” and “package” are used in overlapping ways by different languages.) Are the single-function modules/packages of NPM a good thing? In practice it doesn’t seem that way.
It’s sort of analogous to trying to “giggify” all the jobs. Every individual function outsourced.
For programming languages, even with a flat module namespace, you get a land rush where good names get taken early by packages that might end up unpopular or abandoned.
Leveraging DNS seems like the answer. Java did it badly, but Go's approach seems fine, perhaps because it leverages GitHub's namespace too (for most modules).
O the contrary, Go's approach will lead to lots of problems while Java's is actually simpler and safer. The fact that you depend on the current status of DNS every time you build your code, for each and every one of your dependencies, including transitive ones, is completely nuts. Even Google realized this, and they solved it Google style: they added another automated system on top to try to add some stability (the Go proxy). And then they had to add holes to that system, because it turns out not all dependencies are public and so they can't solve this from on high (GOPRIVATE).
And still, if one of your dependencies decides to switch hosting provider for their source code, or loses their domain name, you have to make (small) changes in every code file that referenced that dependency.
Maven's solution is much simpler for everyone involved: DNS is only involved in registering a new module, it only serves as a form of authentication. After the initial registration, the module name is allocated to your Maven Central account, and it won't be revoked if you later lose that domain. If someone gets access to your domain, they don't also automatically become able to push malware to people who used your module for years, neither retroactively (which Go also handles) nor when they next upgrade (which Go will happily allow).
Maven and Gradle are part of the reason I don't use Java anymore. Java seems to have gone through multiple unfortunate build systems without settling on a good one.
What is it about Maven that you think makes it not a good build system? To me, I moved from Java to Go and I still miss Maven to this day.
/Wheel of Time
/TV Series/Wheel of Time
/Video Game/Wheel of Time
and /Wheel of Time
/Wheel of Time (TV Series)
/Wheel of Time (Video Game)/Francis Bacon
/Francis Bacon (artist)
without having to answer anguish-inducing questions that would be raised by
/Francis Bacon
/artist/Francis Bacon
e.g. questions such as:
1. "Maybe we should put Lord Bacon under a category also, maybe `philosopher/Francis Bacon`"
2. "But he's also a statesman ... what about `statesman/Francis Bacon`? Is he more of a philosopher or a statesman"?
3. What about ordinary people (who happen to be involved in historical events) that have wikipedia entries, like George Floyd? Should he be assigned something like `person/George Floyd`? If so, should the two Francis Bacons be assigned `person/philosopher/Francis Bacon` and `person/artist/Francis Bacon` instead?
And so on
It’s a fundamental issue:
Hierarchical naming and categorization sucks.
In 95% of the cases, you want sets or even graphs rather than trees.
You _might_ be getting away with trees, because you didn’t encounter a case where the hierarchy breaks. But you’re being fooled by happenstance.
In the 5% of cases where trees are fine, you could also use sets or graphs.
This issue comes up all the time with all sorts of things, including urls and file structure.
The trick is to have a general “/type/name” format and to be as liberal and general as possible with what “type” means.
In programming, it’s useful to have namespaces for overall organization and conflict resolution. But it’s a trade off: you are nudged into a hierarchical ontology, with all the implied issues.
Wikipedia titles are free text (more or less). So they can afford not to introduce hierarchical naming and still have nice, easily addressable names without conflicts.
This avoids all sorts of problems. Most things simply can’t be categorized in a strict, hierarchical manner.
I realise this is somewhat tangential to your point, but shows how easy it is for innocent looking choices to end up creating annoyances.
More directly in terms of Joe's brainstorming:
- I don't see how the versioning matter is simplified by a flat space of functions. Before you had Nm modules to track and now you have Nf functions to track, with Nf >> Nm. Aggregating functions in library/modules to version is actually helping with versioning effort, not hindering it. More generally, the versioning of multi-component systems are complex affairs that can only be addressed by constraints - general engineering systems have standards + catalogs as the means of addressing this general engineering issue.
- Broadly I disagree with conflating modules and libraries. They are distinct conceptually. Modules could have state (and meta-data state), conceptually. Modules potentially could also have active elements internally. Modules can have life-cycles. To sum: modules conceputally are not just collections of (related) functions.
So the general question is 'can we live with just libraries of functions?'
I think PLT excitement here is not 'a k/v bag of richly annotated functions' -- the guaranteed end result of that approach is n variants of elaborate 'structure' encoded into the metadata Joe is talking about -- but rather pushing modules to extreme to make the distinction from libraries crystal clear.
The question (then) remains: are modules really just libraries? Was it always just about coexistence of related functions?
> Each Unison definition is identified by a hash of its syntax tree.
> Put another way, Unison code is content-addressed.Are there meaningful contributions to open source that are the size of a single function? Isn't this what leads to the terrible failure modes of NPM (hello left-pad)? If you want to contribute a single function, what's wrong with a blog post?
Also, has the browser not solved this with ESM imports from URLs? I can write a web page containing
import {someFunction} from "https://your-website.com/your-script.js";
This also solves the namespacing issues - at least to the extent that the domain name system solves namespacing issues.(Deno, of course, inherits this ability to import code from any URL.)
I'm being provocative. I've heard people make similar arguments over time, but never understood why anyone would want open source contributions to be a single function.
https://immerjs.github.io/immer/
https://immerjs.github.io/immer/api
The API page lists numerous functions, but the bottom of the page says "In most cases, the only thing you need to import from Immer is produce".
My thought is, immer does something novel, but the vast majority of functionality can be covered in one function.
I think another problem with the "left-pad" situation is the lacking "JavaScript standard library", but that situation is improving over time, especially now that IE11 is deprecated so it's a reasonable expectation to develop against Chrome / Firefox / Safari which are actively maintained and continue to implement new JavaScript features.
(Three functions rather than one, but it’s one main entrypoint and a couple of minor variants.)
I think this is a really well-designed API, and greatly preferable to a more fine-grained OO approach. It does a lot of work under the hood but keeps it carefully contained so you don’t get a bunch of dependency sprawl.
EDIT: it feels to me like the CRC32 implementation contained in that library is a prime example of what the original article had in mind. There should be one global CRC32 function which every person could pull and reuse, instead of bundling it with their own package!
Here’s an interesting discussion on the topic, btw: https://github.com/sindresorhus/ama/issues/10
And having thousands of one-line dependencies is actually much worse than having one thousand-lines dependency, since you now depend on the whims of each of those creators not to stop distributing their work, not to change their license, not to inject malicious code, to address zero-days etc.. Ultimately external dependencies are a form of collaboration, and it's very hard work to efficiently collaborate with thousands of people.
The example in the article is the Fibonacci algorithm and general helper methods. A “remove whitespace” method could easily be set and forgotten (esp with testing). A 1000 line library probably needs maintenance because there is more moving parts and if there’s a bug, you’ll have to test/update a lot more things.
I think the solution to the problems raised are not small libraries but a large standard library with good composable tools. I don’t know the specifics of erlang but in many programming languages things like remove a string whitespace is easy to do inline with the STDLIB instead of downloading a helper method, either in a bigger context or alone.
So you don’t like single line node modules. I don’t think I’ve ever used one. Maybe as a transient dependency, though.
But a function could be written in multiple ways, and I’m not sure we actually mean “single line” here. You’re referring to something of arbitrarily small complexity.
But, getting back to the original comment, what’s to say that one exported function isn’t supported by 100 more non-exported functions? Or, worse case, a 1000 line function that could be decomposed into smaller blocks.
Now, I’m sure I’ve leaned on these sorts of dependencies. Especially ones that have a peer dependency on a broken release, where they monkey-patch the problem. Why would I invest in maintaining short lived code like that myself? My version manager would certainly tell me when things break.
The left-pad debacle proved that much of the ecosystem actually depends on such packages, at least indirectly.
> But, getting back to the original comment, what’s to say that one exported function isn’t supported by 100 more non-exported functions? Or, worse case, a 1000 line function that could be decomposed into smaller blocks.
I'm not sure what the point is here. Of course a huge module might only expose one public function. That's perfectly OK.
The context we were discussing was the question of whether an open-source project should be happy to accept a PR that only adds one (small) function to the project.
And the same question then extends to publishing that single small function as a separate module that others should add dependencies on, since that is essentially what TFA is arguing for.
But either way, the discussion was about modules that expose a small amount of (simple) functionality. And you're right, I'm saying "one line" as a proxy for that.
> Why would I invest in maintaining short lived code like that myself? My version manager would certainly tell me when things break.
I don't understand this part at all.
Perhaps what wasn't clear was what "ideal" I am proposing instead. My ideal is having a dependency tree that is both small (few direct dependencies) and flat (my direct dependencies should also have a small and flat dependency tree of their own). That necessarily gravitates towards fat modules that include lots of functionality under the maintenance of a single team. Of course, this also presupposes a package manager and build system that can efficiently ignore parts of those modules that are not being used in a particular program.
Let's assume the alternative to "small NPM packages" is not "large NPM packages", but "small modules, vendored into my own repository". In this comparison, the benefits of small NPM packages become:
- Transient dependencies are automatically managed for you
- You can pull updates from the repository to get new features and bug fixes, and when you do this, as above, changes in transient dependencies are automatically managed for you
- The package gets a vanity download count on npmjs.com
There are meaningful tradeoffs here, but they're different to trying to rehash the benefits of modular code. Of course modular code is good. But packaging and distributing code is a whole different story.
EDIT: I'm having a bit of a reaction against the excesses of the NPM ecosystem at the moment. As, I think, are a lot of developers. I hope the pendulum swings from its current state more towards "fewer, better dependencies". Not towards "dependencies the size of a single function".
| - do away with modules
| - all functions have unique distinct names
| - all functions have (lots of) meta data
| - all functions go into a global (searchable) Key-value database
So like MP3 tags for functions, and then you get to import by theme/genre/whatever. The system can automatically generate any dependency information not fully made explicit in the give metadata.It's... tempting. I could see this for jq.
No, just the anguish of navigating 1000 non-namespaced functions for the one you need. We might not need modules, but it's good to have namespaces - and just having metadata that are not part of the name doesn't cover that aspect.
Still, it's an interesting idea.
Although given potential problems with function deletion, maybe there's a case to be made for careful curation of these functions in a single place, we could call it the standard library.
Well, it becomes someone else's problem... which is about as far as a technical solution to this problem can go: the namespace problem is actually a social problem.
> but how do we discover > the initial name www.a.b? - there are two answers - a) we are given the name > (ie we click on a link) - we do not know the name but we search fo it ))
The third way is that you look through the list of everything that "www.a" provides on a hunch there should be something like "b" in it. Usually this list is short enough that you can do this and also recognize that "ph" is a synonym so you can use it.
Sure, renaming-at-import would help with globally unique names being unwieldy to use but we already can do this with module imports in pretty much any languages, even in C (preprocessor macros or, when you ran into by identical symbol names at the ABI level, symbol renaming).
So I think modules is one of those "worse is better" things: when you imagine how different and wonderful the world could be, it kinda obvious that modules are pretty mediocre. But they work, right here and now, and simple things are possible and complex things require hacks on top but still are mostly possible too.
You already have this global key value of names to functions. The name just happens to be module+function name
Am interesting theory.
https://github.com/joearms/elib1/blob/master/lib/src/elib1_m...
One could imagine replacing proof of work or whatever the "confirmation" process is with some sort of testing or peer review consensus protocol that would append reviewed code to the chain.
A blockchain contains a full history of all modificiations, which sounds a bit like version control if you look at it funny.
In this case, the proof-of-work would be peer-review instead of computation, and the "blocks" would be as frequent as "minor releases". The need for computation would be pretty light.
Not that I'd actually support it in practice lol
"A blockchain is a distributed ledger with growing lists of records (blocks) that are securely linked together via cryptographic hashes.[1][2][3][4] Each block contains a cryptographic hash of the previous block, a timestamp, and transaction data (generally represented as a Merkle tree, where data nodes are represented by leaves)"
...how is this relevant to the linked mailing-list thread? They're talking about compilation units.
- all functions go into a global (searchable) Key-value database
...
- there are no "open source projects" - only "the open source
Key-Value database of all functions"