The end of "Useless Ruby sugar": On intuitions and evolutions
zverok.substack.com
zverok.substack.com
Incidentally, this is actually the definition of syntactic sugar, from a language design point of view. People often use it derogatorily to refer to bits of syntax that they simply don't like, but as a term of art it refers to syntactic elements that are introduced in an effort to provide new ergonomics—new ways of working in a language—without altering the language's fundamental primitives. As the article notes, in theory this method of language extension should actually keep the language simpler than if these elements were introduced as new semantic constructs.
There are a few places where this goes wrong in practice.
The first is that you end up accumulating a lot more syntax than you would if each syntactic extension to the language required a new semantic basis. It's easy to support syntactic sugar on a compiler or interpreter backend, so a language that is built on this small-core-with-sugar model doesn't have any natural forces urging it towards simplicity. To an experienced developer who watched all the layers of abstraction accumulate this doesn't pose much of a problem, but it's not always obvious to a newcomer what all this syntax desugars into, so for them it may as well be a new semantic element!
The other major place where I've seen sugar go wrong is that it rarely gets extra tooling to support it. A brand new semantic element will naturally need to be supported at every level of the language's stack, but the entire point of adding sugar is to avoid that extra work, so the language developers often just assume that the existing tooling will suffice. This leads to rough edges in things like stack traces and the debugger, because the code that actually gets executed only bears a loose resemblance to the code that the developer wrote.
It should be possible to avoid these pitfalls and have a tight language with a small basis and solid tooling, but it requires a lot more discipline.
Shoutout to the folks behind Sorbet for making an awesome LSP for VSCode [1]. It comes close to the tooling expected nowadays, bit is still very far behind even popular dynamic languages like JS or Python.
In a context where I'm working on a few languages regularly and among a larger set of languages that coworkers are in charge of, I can see how Ruby has solved a ton of problems that other languages are still struggling with.
- package management: bundler/rubygems has been a sa-holved problem for the longest time whereas npm/yarn is an unholy mess, python is only beginning to see the light after refusing for so long, go mod is kind of a step in the right direction but still falls short in others, cargo is an outright copy but tries too hard to also be rake, a distinct problem making it a kintchensink... Gemfiles being descriptive, one can generate lockfiles from their platform including information for a foreign one and it's going to resolve consistently, including corner cases like this gem has a transitive dependency on a ruby extension and this platform has a binary gem at that version but that platform can build from source at that other version so I'm going to pick the one that's consistent and respects the constraints that have been described, or bail out and yell at you. This makes using semver-like squiggly a breeze and a non-event.
- build tools: scons, cmake, ninja, doit & al. feel both incredibly complex and limited compared to rake. Every single time build files end up becoming obscure behemoths. When they don't they're incredibly focused on one thing but to make them to anything else they become a hodgepodge of additional external scripts and have to rely on so many tricks and hacks, which is why I used to stick to make. People seem to also have forgotten that rake can perfectly replace make as well, but rake can do so much more and still present a consistent interface and task implementation.
- testing: minitest/test, minitest/spec, spec are running circles around any other test framework. The composability of reusable contexts, nested cases, and shared examples allows one to cover a ton of behaviours with very little test code that is eminently descriptive. Just the other day I did a deep reflector of some part of the code, and I only had to change a bunch of helpers as to how things should operate, and the entirety of my functional test suite was literally a zero-line diff.
- talking to browsers: capybara is still my go-to to automate browsers, whether in tests or not, headless or not. It's just way too easy.
- linting: Rubocop has a metric ton of tunable checks to enforce your preferred style and optionally auto fixes a lot of violations, standardrb takes a tab from gofmt by giving you some sane community defaults for Rubocop and thus autoformats by default. Rubocop is easily extensible, either you pull things and can lint rakefiles, Rails, spec and whatnot, or if you have special cases you can write your own cops.
- typing and LSP: forget Sorbet, RBS+Steep is the way. RBS can generate typed skeletons either statically or dynamically using typeprof; ideally if you have good coverage you could run tests using typeprof and have automatic typing generation that can then be used entirely statically. Then the `rbs` command then allows you to statically ask a ton of questions about types but it's not a linter by itself, that's where Steep comes in. Steep comes with a LSP and it Just Works: autocompletion, inline type information, and so on. Type signatures can be packaged within gems, but if they're not there's rbs_collection, which also integrates with bundler.
So I'm not too sure what you're referring to as "tooling" that would be primitive when compared to other languages when in my mind it's the other languages that constantly play catch-up ;)
Sounds like I should give !rspec a go... Since the rest of your comment was super informative, I'm interested in your take on RSpec and Cucumber?
> RBS+Steep
Hot damn I am out of date and this sounds amazing. Can't wait to try it out... Thank you for posting!
> not too sure what you're referring to
Best guess is it's like with patterns: Ruby doesn't have that many "patterns" (particularly the canonical kind) because they're just plain unnecessary. Similarly, thinking over other languages... there's just not much that's needed, which might then look like "not a lot".
In every project that I've seen using it, it was a tremendous unnecessary waste of time in comparison to writing feature/acceptance tests just in Ruby. Gherkin is just such a poor language in terms of expressiveness compared to Ruby. Not to mention things like poor REPL experience when writing the scenarios.
I found that Cucumber misses a bit on its promise: in theory things are perfectly described by high level behaviours but in practice it's a bit leaky and I regularly had to work around the tool, which was of disservice to readability. After all if you are faced with a high level descriptive test but have to know about some internals as to why this or that case has to be tested or why it has to be written this or that way (sometimes tacky ways) then it's probably that the tool is not quite fit.
Rspec (or minitest/spec) eschews that: essentially it's just Ruby so you write what you need. The other side of the coin is that things you write might not be as "high level" as you initially want: you can have something working but very raw and possibly a tangled hard to maintain mess when it should be "higher level" (or several levels composed together). IOW the big advantage of Rspec is that it scales across from "lowest level" unit testing to "highest level" behaviour so it's always fit, but then it's on you to pick the "proper" level that fits and design your specs accordingly.
An example (not perfect by any means but gives a good idea of composition), here we extracted a bunch of functional ("integrated") behaviours into separate examples:
https://github.com/DataDog/dd-trace-rb/blob/57976d52bde573f9...
Note that some are nested, and with a flick they test behaviour when e.g `appsec` is enabled or `tracing` is disabled. This creates a ton of combinations for only a few tests and coverage goes through the roof.
Now here's how we use that:
https://github.com/DataDog/dd-trace-rb/blob/57976d52bde573f9...
Notice that magically every case is going to be tested and so we're absolutely sure that it works whatever a user toggles in their configuration.
Notice that we create a Rack app, and then the tests proceed, just picking whatever app is in that `let :app`, which leads me to something we have yet to do:
https://github.com/DataDog/dd-trace-rb/blob/57976d52bde573f9...
https://github.com/DataDog/dd-trace-rb/blob/57976d52bde573f9...
Notice how essentially just the `let :app` setup differs, but tests are entirely similar to the Rack ones (IIRC save for a couple that don't apply to Rack) so actually the overall behaviour can be entirely extracted and shared too! The result would be a per-framework test suite would be a) set up :app for that specific framework and b) a single `it_behaves_like "my super high level behaviour"` line, shared across all frameworks.
Bonus (not sure you're familiar with it so here goes): this is tested using Rack::Test, which thanks to Rack's design provides a bunch of request methods (`get` and so on) that merely call the Rack stack as if it were a web server. IOW the whole app code is called just as if there were say Puma or Webrick in front but without ever creating a web server or a socket or whatnot, so it's stupidly fast, and since we're straight in the Ruby code we can mock or stub or allow/expect methods to be received or whatever is needed to our heart's content.
Every bit conspires to produce a simple, readable, compact test suite that combinatorially covers every case and runs in a stupidly short amount of time.
It does seem like this can cover a bunch of what I usually turn to Cucumber for - the higher-level descriptions - although I think it reinforces how I generally like to use it: for the multi-step full user flow integration "tests" (leaning heavily towards the "this is a product description" and away from "these should be comprehensive"). That is: use something like RSpec to get your ~100% test coverage via unit tests, and then use Cucumber to both (1) ease communication with what a feature does, and (2) exercise all the code involved in a main user path.
TL;DR - Seems like the happiest path is something like (1) unit tests that barely use contexts, (2) integration (request, controller) tests that heavily use contexts (your examples), (2) feature (cucumber) tests for multi-step feature flows. Unit tests should aim for 100% coverage (but it's OK to miss); integration tests should aim to cover 100% of the different kinds of situations that are encountered (; feature tests should aim to communicate what users actually do.
PS - haha, if you're still at Datadog, talking with a colleague of yours on the sales side about switching from New Relic to y'all. Knowing this quality of code is under there is a a definite point in your favor :)
> if you're still at Datadog
As a matter of fact I am:
curl -s https://github.com/DataDog/dd-trace-rb/commit/176c642ca73679cabc5fa1a113bc9b600aa04dcd.patch | grep '^From:'
Shoot me an email if you feel like chatting.You mentioned build tools. Ruby's build tools are still inferior to Java's Maven. Granted, few build tools are as good as Maven. This is coupled with something you haven't mentioned, which is the native libraries that Ruby also depends on, plus the management of Ruby versions (e.g., rbenv, rvm), something that Java, for example, seldom has to deal with, i.e., Java dependencies are mostly pure Java, and the latest Java can still work with JARs compiled for Java 1.x. This backwards compatibility, which makes the latest version good enough for most projects, is also shared by Node. The latest Node LTS is usually just fine. Ruby's situation is not as bad as Python's in this regard, but many times you still get the sense that, in order to get a reproducible build, it's wise to just build inside a Docker image.
The problem I see with Rake, or Bundler/gemfiles is that the syntax is just Ruby code. It's a DSL described by Ruby, the language being good for it, but it's Ruby code nonetheless, which makes tooling more complicated than it needs be. Rake is cool and all, but `make` is available everywhere, and that basically kills any potential that Rake might have for non-Ruby projects.
I'm glad to see that RuboCop exists, or the LSP situation improving. The problem that Ruby will always have is that it's a dynamic language that's not JavaScript, hence anything related to static analysis is subpar. This is something you may be able to live with, due to the very playful way people tend to work with in Ruby (or other dynamic languages), but tooling will always be subpar compared to static languages.
And granted, many other languages have bad tooling, too.
I see this as a plus. A tool for ruby, using ruby to configure it. Yea you won't use it outside ruby projects, but that's totally fine.
...and then call them from Make ;)
Take "build tools". Java needs build tools because it's got to be compiled. Ruby isn't compiled, so you don't need build tools - you only need dependency management, and Bundler is a best-in-class for that. Trying to say "Ruby's build tools are inferior to Maven" is like saying ""fish bicycles aren't as good as human bicycles".
> Docker
Eh. RVM was good enough it got cloned (NVM); between that and `bundler exec`, you didn't really need anything else; Docker is nice because those two don't cover everything, but again - I think you're looking for solutions to problems Ruby largely doesn't have, and then judging the tooling for not existing.
> Rake
Uhhh hate to tell you this but 1) you don't need to use Rake for things you can do in Make and 2) you can't use Make for things you can do in Rake. Rake is mostly used for things that need the rest of your ruby code; aka, Rails. Sure, you can write a Ruby script that does all that and invoke it from Make... and then you've just recreated Rake, only without the bells and whistles. (What you do is just invoke the Rake tasks from Make)
> Rubocop... LSP... static analysis
Sure. There's some pretty cool new / on the horizon stuff around this (RBS and related, plus a new unified parser - https://rubyconf-2023.sessionize.com/session/531482)... but, and I realize this, in particular, is super subjective: you just don't have the same problems. Ruby programming tends to only use a few types (basic scalars, basic collections, and ORM models), and then testing. There just isn't the same need for static analysis, and it's because of everything from overall simpler code & parameters to tests covering the same need.
TL;DR - Yeah, if you try to program Ruby like it's Java you're gonna feel the "missing"tools, but largely (and subjectively) if you probram Ruby like it's Ruby you just don't have the problems that those tools exist to solve. (This is also why something like 75% of the Gang of Four patterns just don't show up in Ruby - they exist to solve issues that only occur in Java / languages like Java)
Edit: One place I do miss "better tooling" - sometimes! - is the debugger. Pry is amazing and gets you 90% of the way, but it's not quite as nice as an IDE-integrated breakpoint and the like. But, again, it's just not as needed - not when I can pop into the console and invoke the code manually.
Edit 2: Oh also autocomplete. Definitely miss type-aware autocomplete... but it sounds like RBS is making headway there.
Check out the latest IRB on 3.3, it finally gives Pry a run for its money, especially when jumping into the debugger!
https://github.com/ruby/irb#debugging-with-irb
$ ruby test.rb
From: test.rb @ line 7 :
1: def foo
2: 10.times do |i|
3: p i
4: end
5: end
6:
=> 7: binding.irb
8:
9: foo
irb(main):001> break 9
(rdbg:irb) break 9
#0 BP - Line /Users/loic.nageleisen/test.rb:9 (line)
irb:rdbg(main):002> break 3
#1 BP - Line /Users/loic.nageleisen/test.rb:3 (line)
irb:rdbg(main):003> continue
[4, 9] in test.rb
4| end
5| end
6|
7| binding.irb
8|
=> 9| foo
=>#0 <main> at test.rb:9
Stop by #0 BP - Line /Users/loic.nageleisen/test.rb:9 (line)
irb:rdbg(main):004> continue
[1, 9] in test.rb
1| def foo
2| 10.times do |i|
=> 3| p i
4| end
5| end
6|
7| binding.irb
8|
9| foo
=>#0 block {|i=0|} in foo at test.rb:3
#1 Integer#times at <internal:numeric>:237
# and 2 frames (use `bt' command for all frames)
Stop by #1 BP - Line /Users/loic.nageleisen/test.rb:3 (line)
irb:rdbg(main):005> bt
=>#0 block {|i=0|} in foo at test.rb:3
#1 Integer#times at <internal:numeric>:237
#2 Object#foo at test.rb:2
#3 <main> at test.rb:9
irb:rdbg(main):006>
See also the debug gem by itself:- https://github.com/ruby/debug#debug-command-on-the-debug-con...
- https://github.com/ruby/debug#remote-debugging
- https://marketplace.visualstudio.com/items?itemName=KoichiSa...
The problem is not Ruby's dynamic typing; table stakes here is the good-not-great tooling support for other dynamic languages like Python and JavaScript, not the great support for more typeful languages.
I think Ruby plays really poorly with editor/static analysis tooling because of the culture and conventions around blocks.
The fact that so many libraries/frameworks in Ruby use the very common pattern of "provide variables/context/symbols to blocks via scope (rather than block arguments)" massively increases the difficulty for tools to provide useful type info, completions, and validity analysis. As a pattern, "scope injection" is just harder to statically analyze from the outside, and, while not ubiquitous, it's alarmingly normalized in the Ruby community.
I'm not sure if it's the chicken or the egg for many Ruby programmers' heavy reliance on the REPL for autocompletion and probes of local/instance/class/global variables. Regardless, editors' and linters' usefulness is reduced in proportion to the dominance of this pattern across the Ruby ecosystem. Heck, even when heavily block-oriented libraries are well supported by IDEs (rspec, rails) the support turns out to be based on brittle pattern matching and hardcoded guesswork in the tools.
There are other cultural (e.g. tendency to overuse stringy method calls and method_missing, though this thankfully seems to be less common these days) and technical (e.g. autoloading, open classes) things that increase the difficulty of writing good Ruby tooling, but I suspect that they're much less responsible for the anemic tooling landscape than block-oriented programming.
I really do think it's largely cultural. You can do block/scope-injection -based programming (and eval, and method_missing via getattribute, and so on) in e.g. Python without too much trouble, but it's not nearly as much of a community norm as it is in Ruby, so people do it less.
And that's OK! Many prolific Ruby programmers have found ways around the tooling deficit and are productive with the language, and Ruby remains a go-to for people writing "DSL"s (though I find when people refer to DSLs in ruby they really mean the combined effect of parens-optional plus method_missing plus block scope injection), so this seems like far from a fatal flaw.
Well, I'm fairly experienced because I've been using Ruby since 2005 almost only with Rails plus some small scripts that I used to write in Perl, and yet I'm not up to date with all the new syntax and sugar that accumulated in the last years.
There are at least two reasons for that.
1. Working with Rails means that we tend to use only the features that make sense to use inside a Rails app. An extreme example: I experimented with Ractors but if they have a place in Rails app it's somewhere inside rack or db drivers, or maybe something like Sidekiq. They have no place in code that runs in the short lifetime of a web request.
2. Each customer use a different version of the language and of Rails, also among different apps of the same customer. They stick to those versions for years, usually until they approach end of life. That makes hard to remember which feature you can use in which project, so one tends to use the oldest common feature set. Example: I never saw anybody using pattern matching, which by the way changed between 2.7 and 3.x.
This irritated me a lot when I was a ruby programmer. Ruby's require is analogous to C's #include: modules are just files that get evaluated in order to add stuff to a global interpreter state.
It got to the point I went around asking people on IRC what they thought the perfect module system looked like. Studied long forgotten languages to see how they did it. To this day I'm obsessed with this. I feel like it's one of those things you just have to get right since day one because it's gonna be impossible to change it later.
So yeah, I'd agree with saying the migration is partial at best.
What's the theoretical alternative? Adding stuff to a global interpreter state is the goal.
The weird part of #include, to me, is not that it ends up modifying the global interpreter state. That's the point of calling it! The weird part, what C libraries tie themselves into endless knots to work around, is that when you #include something in more than one place, it will execute more than once.
* Lexical scoped imports a-la python
* Pass state to modules explicitly, which I think is what newspeak was doing (i.e there's no global state)
Either way you'd lose ruby's "let's add methods to Kernel and have them pop up everywhere" so it's a trade off.
Worth nothing: there are a few proposals around to have some sort of namespaces in Ruby, but it's unclear they'll ever happen.
E.g.
How is this an alternative to modifying the global interpreter state?
Before I do any imports, I can type this:
os.getcwd()
and the interpreter will crash, telling me that "os" doesn't mean anything.But if I do this:
import os
then, later, I can type exactly the same thing: os.getcwd()
and the interpreter will call the function and return a string.In order for the interpreter's behavior to be different when presented with exactly the same input, it's necessary that the state of the interpreter has changed. What's the alternative?
And since this is the entire purpose of a module system -- causing the global interpreter to recognize names that it otherwise wouldn't have -- it follows immediately that all module systems will operate by making changes to a global interpreter state.
The reason Python's module system is better than C's has nothing to do with whether or not either system modifies a global interpreter state, since that doesn't differ between the systems. The reason is that Python allows you to specify what you want those modifications to be. If I have two Python libraries named os, I can import them like so:
from lib1 import os
from lib2 import os as os2
and now they don't conflict. As far as I know, C has no such facility.Python's module system results in no implicit modifications to Python-level state between modules — at least, for a well-developed module. Since Python is so dynamic, you can of course jerry-rig all sorts of global consequences if you really want to. But without dynamic magic, all state is namespaced to the imported module, and must be explicitly given names within the importer, with no global consequences for parallel modules.
In your os.getcwd() example, the state is scoped to the importer module, so it's not an example of global interpreter state.
I don't know what Ruby's module system looks like, so I don't know for sure if that is the point the GP is making. It seems to me that you are in agreement, just disagree with the meaning of "global" in this context.
The reason they do that, of course, is that they want to be able to access #included code from the highest scope in their file. But once that's the requirement, we're back to saying that this is the goal of all module systems, and all of them will modify the global namespace in this way.
> all state is namespaced to the imported module, and must be explicitly given names within the importer, with no global consequences for parallel modules.
This is just false; state appears within the interpreter in the form of registering the imported module under the provided name. You are free to register a name that's already in use, in which case you will lose the ability to refer to the module that was registered with the same name earlier. Becoming unreachable is ordinarily considered a "consequence"; you can find people complaining that this has happened to them.
Why does the problem arise? Because importing means making changes to the state of the interpreter.
No, it doesn't?
def print_sysinfo():
import sys, os
print (f'{sys.platform} {sys.version}, current dir is {os.getcwd()}'
os.getcwd() # throws NameError, because os is not a thing there.
Now, you may argue that the 'os' module is still loaded after a call to print_sysinfo() and while it's true, it didn't have to be so: Python could as well had have semantics of unloading modules that have zero references to them, and re-loading them anew every time they are imported again.A language can have locally-scoped module semantics, it just that most of the time the static/global scope is much more convenient.
---main.c----
void f(void);
int main(void) {
int a = 0;
f();
return a;
}
void f(void) {
#include aux.c
}
---aux.c-----
int a = 1;
By your argument, #include doesn't make changes to global interpreter state either. Is there a sense in which the #include preprocessor directive works by making changes to the state of the global interpreter, but Python's import mechanics don't work by doing the same thing?There's only one interpreter and your goal is to make a change to its state. Hence, "global interpreter state". The change you make can affect behavior in, in principle, any context in which the interpreter is used. Could be inside the definition of `f()`. Could be outside the definition of `f()`. There's no difference in the mechanics.
What matters is language (not interpreter) semantics, here it's about variable scoping. In C, if I include "a.h" and then "b.h", "b.h" will have access to definitions from "a.h", even though lexically they were never brought into scope in that file. This is an actual problem that C/C++ developers have to deal with, since it makes it hard to create isolated headers.
In Python, the code in each module is evaluated in an isolated context, and returned as an object, so this isn't a problem. Edge cases like monkey patching don't really matter for >99% of usecases.
Why is that a change to the interpreter state? Consider some more C:
{
int a = 1;
a = 2;
a = 3;
}
The line declaring `a` makes a change to the interpreter state; `a` becomes a reference to something. Without line 1, lines 2 and 3 are invalid.The two lines after that, manipulating the value of `a`, don't make any such change. If you delete line 2, line 3 will be handled in exactly the same way as otherwise.
The reassignments to `a` make changes to the program's runtime state, but they have no effect on the interpreter. The braces, which I guess I have to call lines 0 and 4, are the opposite - they're very relevant to the interpreter, but it's hard to give them a runtime interpretation.
The particular task of resolving references isn't necessarily handled by anything you mentioned - it is the job of the parser to determine that a reference is being made, but not whether that reference goes anywhere. Similarly, this information is not included in the abstract syntax tree, which is fundamentally the same thing as the parser. It's perfectly possible to form a syntax tree that calls for the invocation of a function that doesn't exist, or assignment to a variable that doesn't exist.
In C writing an invocation of a function that doesn't exist will give you a compiler error, but in Python it won't; you'll get working code that trusts that whenever it happens to be invoked, the function will exist. And if, at invocation time, it doesn't, you'll get a runtime error instead.
Since this is the problem being addressed by the "interpreter", it appears to be the case that the interpreter in C exists at a different point of the compilation-execution pipeline than the interpreter in Python. The logical steps of determining what a program does need not coincide with the physical steps that are taken to convert the program into an executable artifact - or more fastidiously, they don't need to occur in the same order from program to program, and might be bundled with different other steps.
But what I've claimed is that the import facility, whatever it might be, is responsible for making changes to the state of the interpreter, wherever the interpreter might be active. It makes the difference between "the programmer has specified doing something meaningless" and "the programmer has specified doing something meaningful", and the meaning that it provides gets loaded into the state of the interpreter. It is not strictly necessary that this occur at the top level of the program, the global scope, but it is almost always the case because that is the scope where programmers almost always want imported code to be available to them.
> In C, if I include "a.h" and then "b.h", "b.h" will have access to definitions from "a.h", even though lexically they were never brought into scope in that file.
This actually appears to be the opposite of the normal concern with lexical scope. Lexical scope allows names defined at a higher scope to be accessed from a lower scope, but that's precisely what you're trying to prevent here.
It's worth noting that b.c is compiled without any knowledge of a.h or a.c (assuming it doesn't itself #include a.h, which would make the problem you mention something intentional instead of a pitfall). "libb" doesn't suffer any problems from being #included after "liba". The problem occurs in your code, where you have access to definitions from both header files, which is what you wanted when you #included both of them.
There is only one memory space in the process hence no variables are local. Well, this is true, in a sense, but I don't think it's a particularly useful perspective. Also, see the sibling comment.
I guess that scoped changes is an option, that is you ask to alter the interpreter behavior only in a lexical scoped.
Ruby has refinement in this spirit, though it’s not exactly operating at the same level as a require.
Here is a related discussion that arised in Crystal community a few years ago: https://github.com/crystal-lang/crystal/issues/4439
When a module is imported, the file is evaluated once, then stored in memory as a stateful object. Further imports of the same module, even from other files, reference the object in memory instead of evaluating the module again.
The dict used for this is accessible in sys.modules
I meant to allude to the fact Ruby has a single global namespace that all Ruby code adds bindings to. In Ruby we have to manually namespace stuff just like we add prefixes to functions in C. Python on the other hand creates isolated namespaces for each module which makes it impossible for them to clobber each other. Import creates links between those isolated namespaces.
This is the right way to do it but it suffers from a lack of correct tooling. The Self[0] programming environment assigns a "module" to each "slot" (class methods in Ruby land); an object may be composed from multiple modules at once. If Ruby had this kind of granularity in its inclusion system I believe that a lot of the confusion around where a particular method comes from would disappear.
What if I want to write some code that will use the module I've loaded?
Correct. We must create such mechanisms. The aptly named import mechanism seems to be the most common solution to that problem.
There's a heap with objects floating freely around, and two tables of symbols and the objects they point to. The sets of symbols in each table may intersect freely: both tables might have an "X" symbol without clashing with one another. On the other hand, there's no intersection between the sets of objects they point to at first. They start out fully isolated and their values are unreachable from the outside.
As you noted, this is of limited usefulness to programmers. There's no point in collecting symbols into a module if the values those symbols point to are unreachable. We need access to foreign symbols in order to make use of abstraction. To that end, the implementation may provide a symbol importing mechanism. It could be as simple as setting a symbol on the current table to the object pointed to by a symbol of the other table. In fact that's exactly how I implemented it in my language.
So this "imports" the symbol from another table to the current one. Both tables were completely isolated at first but now there are explicit links from one table to the other: there's two symbols from two separate modules pointing to the same object in memory, linking one module to the other. That object is most likely a function and importing it allows the programmer to call it.
> On the other hand, why bother giving it a symbol table at all?
Because symbol tables are how variables are implemented. They're literal hash tables mapping symbols to arbitrary values, symbols are just unique strings. If you don't give modules a symbol table, you are unable to set variables at the module level.
> So this "imports" the symbol from another table to the current one.
So in this terminology, we can "load" a module and this has absolutely no benefits, except that we can subsequently "import" the same module and importing gets us what we want, the ability to refer to the imported code.
We accomplish the import by making a change to the global symbol table, inserting references to objects within the module.
Is this different from what I've been saying above? The entire purpose of loading a module is to add things to the global interpreter state. Why is this an "alternative" to adding things to the global state?
I meant to allude to the fact Ruby has a single global namespace that all Ruby code adds bindings to. In Ruby we have to manually namespace stuff just like we add prefixes to functions in C. Python on the other hand creates isolated namespaces for each module which makes it impossible for them to clobber each other. Importing creates explicit links between those isolated namespaces.
> So in this terminology, we can "load" a module and this has absolutely no benefits, except that we can subsequently "import" the same module and importing gets us what we want, the ability to refer to the imported code.
Yes. We load modules, and we import symbols from already loaded modules. In most if not all languages I've learned, importing symbols loads the referenced module implicitly but it most definitely is a distinct operation.
Loading means reading the module's code from somewhere, usually the file system, and placing it into memory, then evaluating the code to obtain the resulting bindings, code like:
module xf
x = 10
f = (y) { y + 10 }
z = x + f(80)
The result of loading a module is just a hash table binding variables to their values, just like Javascript's require. modules["xf"] = {
name: "xf",
bindings: {
x: 10,
z: 100,
f: {
type: function,
arguments: ["y"],
code: "y + 10"
}
}
}
This is also true of native modules like ELF libraries, they contain the exact same table of symbols, the only difference is this evaluation is done ahead of time by the compiler. This hash table is almost always cached by the implementation so that it does not need to be loaded again.Importing essentially does this:
xf = modules["xf"]
bindings = xf["bindings"]
x = bindings["x"]
f = bindings["f"]
z = bindings["z"]
Naturally, the modules object must contain a "xf" key pointing to a module object containing the bindings for "x" and "f" and "z". Loading a module puts the module object in the modules cache.So loading a module without importing their symbols could present benefits. In most languages, they are lazily loaded, work is done only when they are referenced. You might prefer to load everything up front while the program is starting up though. I went through quite the adventure to embed lisp code into an ELF binary in such a way that Linux would automatically map the code in before the program even started executing. Having my interpreter import the symbols from the embedded modules which are already in memory was relatively trivial.
> The entire purpose of loading a module is to add things to the global interpreter state. Why is this an "alternative" to adding things to the global state?
Yes. The interpreter could maintain one namespace and make it so that all code modifies that namespace. It could also maintain a separate namespace for each logical module. In the latter case, changes are local rather than global, even though they're all part of the global state of an interpreter program.
It's analogous to git repositories being local and distributed even though everyone is using the same git implementation to operate on them.
Most people I talked with seemed used to and very happy with the Python model: modules are distinct namespaces for symbols and their referenced objects. Pretty much everyone I talked to had no difficulty understanding modules, it's always the package management that makes things complex no matter which language it is.
Something I spent quite a bit of time thinking about was whether it was worth reifying that model into the language as first class objects. In other words, should import be a special keyword that makes the interpreter magically bind symbols or a function which returns a regular everyday normal object? At the end of the day, Python's modules are just dictionaries, just like Javascript modules. Should I make that fact apparent or hide it behind language keywords? Interestingly, both languages went in the opposite directions: Python went from special import syntax to allowing you to access modules as a dictionary, while Javascript went from a require function that returns an object to an import statement. I ultimately chose a somewhat weird mix of both, powered by lisp's flexibility.
I also tried to figure out how compiled languages approached modules. So I dug up literature on Modula, Modula-2 and Oberon and tried to figure out how they represented modules. This ended up having a significant influence in my design in the form of isolated modules with export control. Each module contains its own table of symbols and their references. Importing is just setting a local symbol to the value of the other module's symbol, and only symbols in the exports list can be imported. I also liked the qualified and unqualified names: added the option to prefix the module name to the local symbol.
I also thought a lot about how to map modules to the file system. The idea of program folders from Windows has been an inspiration for a long time now. The idea is if the module itself is reachable then all of its submodules are also reachable by the loader. I also think it's important that no file escapes the package directory. Python and Ruby frequently have thing.{rb,py} scripts and thing/ directories side-by-side, I sought to eliminate that for the module's root directory only which results in a main file like thing/thing.{rb,py}.
To enable module-oriented development, I've found the most important feature of the modules and packaging system is editable libraries. Like pip's editable installs and npm link. When I develop a project, I often end up with several supporting libraries. Languages should support linking these local versions to the main project so they can be developed simultaneously.
Another interesting concept I ran into while researching modules is parameterized modules. Essentially, modules that take their dependencies as arguments. The modular equivalent to dependency injection I suppose. Instead of a module importing by symbol a specific library as a dependency and some package manager resolving it to actual files in the load path later on, the programmer explicitly loads the library and passes it to the module as an argument.
Instead of the symbolic imports we're all used to:
(import lib); lib imports lib2 internally
Module importing becomes analogous to function calls which construct an instance of the module given its dependencies: (import (lib2))
(import (lib lib2))
This is really elegant and more or less reifies package management into the language. However, it presents serious ergonomics issues because it forces the programmer to deal with all these package management and library loading details. The truth is we want to sweep all that ugly stuff under the rug, not deal with it every single time we import a module.The main benefit, the loose coupling that stems from the ability to substitute dependencies without having to change the importing module, can be accomplished in a declarative manner via package managers. Arch Linux packages for example may have a "provides" variable which allows multiple packages to implement an interface of sorts and be used interchangeably to satisfy dependencies. So I think parameterized modules imposed significant costs for little benefit.
One thought regarding the quantification element. I’ve grown to appreciate the uniformity of Go’s solution, where you always import an entire namespace, optionally aliasing it to avoid conflicts.
This small restriction makes some naming decisions simpler (for example sticking to config.Parse, and not the stuttering config.ParseConfig). Additionally it sprinkles the namespaces over the code so you get at least _some_ feel for them.
On another note, on some occasions I have thought to myself that it’s easy for the import section to get to little scrutiny on reviews, and wondered if there’s something that could make us wonder „should x depend on y?” more often.
That's a very good solution in general which makes everything uniform and consistent. It's only due to my personal tastes that I didn't implement it that way.
I'm obsessed with symbol management. I'm so obsessed with this I wrote my language in freestanding C just so there would be no libc and compiler cruft in the resulting ELF binary. I'd rather deal with complexity than see weird doubly underscored stuff in readelf output.
Names are everything in computer science. I have some kind of psychological need to have clean names. For that I need clean namespaces that I can shape to my will. So I absolutely wanted the ability to import only the symbols I needed in order to minimize the pollution of the namespaces. I support just importing everything as a convenience but I personally never use that feature. I also made sure I had the ability to rename imported symbols to anything I wanted just in case other programmers aren't as obsessed with names as I am.
I went so far with this I implemented basic control flow as a library of lisp macros. There are no reserved keywords or special cases, they're just normal functions that get imported like all the others. Thus they can be renamed or avoided entirely. In fact I made it so only two symbols are present in every namespace: import and export. And those can be overridden too after the programmer is done with them.
I probably have more than few screws loose or something. Go's approach is totally reasonable. Instead of importing N symbols from a module, it imports 1 symbol and nests all N symbols under it, and if there's a clash you only need to rename the module prefix. It's nice and creates single points of truth.
... Now that I think about it, the only reason I didn't implement it the Go way is I didn't want to add special syntax for nested symbols to my language. In other words, in my language "config.Parse" is a single symbol instead of a "config" + "Parse" pair. I feel like it just wouldn't be lisp anymore if I added syntax to decompose the former into the latter.
> sticking to config.Parse, and not the stuttering config.ParseConfig
Yes. I find the stuttering repetition in your latter example to be profoundly irritating and a symptom of a bad modules system. In my language I tried to prevent that by making it easy to add or remove module prefixes to the imported symbols. I also made it easy to rename symbols for good measure just in case people did it anyway.
There are a bunch of issues about this upstream about this, e.g: https://bugs.ruby-lang.org/issues/14982
My exploratory take on it: https://github.com/lloeki/package-ruby
I'm currently drafting a proposal and planning to open another one, more focused on a specific aspect and a different angle, and thus with a couple of comparative PoCs that'll be a bit different from my above exploratory attempt.
† ... via `class` and `module` keywords, as one could always include/extend/prepend or use `module_eval`/`class_eval`/`instance_eval`.
Ruby's classes are reified as normal objects. A module system manages names, symbols, references to those objects. To modify a class we just need its object, which name it comes from doesn't really matter. We'd be able to modify classes even if Ruby had a real modules system. All we need to do is get a reference to that class object, probably by importing it from somewhere into the current namespace.
Ruby does exactly that already. Every class just happens to get stored as a constant of the Object class. So creating a class A just stores the new Class instance in ::Object::A which is totally weird when I think about it. And when programmers write A, Ruby looks up the constant in current scopes and eventually finds ::Object::A which contains the Class instance. A proper modules system would allow programmers to define where that object ends up or gets pulled from.
And actually it can, which is part of my initial proposal with the linked repo, and goes much more in depth in my pending, non-public draft.
And what is your hot take on what Python does?
Javascript's system is essentially the same as Python's but with require being a function that returns normal objects instead of an import keyword which magically sets the symbols. At least it was back then, new versions of the language transitioned to a dedicated import keyword. I found the reification of modules into first class objects to be an elegant design, not sure if they discovered problems with it or if they were severe enough to justify that transition.
I think a mixture of both plus export control would be best. Syntax is friendly and good for most uses but a function that returns a reified module object is also important since it allows programmatic access to modules, allows creating plugin architectures and much more.
It's not only used in a derogatory way. It's just what it is. Syntax that doesn't change the power of the language particularly; just lets you save some keystrokes and write things more neatly.
I'm hoping for a λ keyword in Python as a synonym for lambda. Don't need it; would love it.
Ruby already allows utf-8 identifiers, so you can do π=3.14 for example, and `alias λ lambda`, but I agree it would be nice to get a built in alias for lambda.
“Just a spoonful of (syntactic) sugar makes the…”
Better to just have simple syntax to begin with, like Forth or some such.
If it is less characters to type out, and has a very specific intended purpose, there will always be an engineer on your team who tries to use it just one more time because they were lazy and didn't want to type the extra 5 characters, and then one more time, and then one more time, and then you're debugging code that is unreadable.
This is not an argument for/against any specific syntactic sugars, this is just an observation that I've seen.