I think the size of the standard library is just one of possibly many contributing factors that leads to a large number of dependencies. I think a part of it is culture, but another part of it is that the tooling _enables_ it. It's so incredibly easy to write some code, push it to crates.io and let everyone else use it. That's generally a good thing, but it winds up creating this spiral where there's almost no backpressure _against_ including a dependency in a project. This means there's very little standing in the way of letting the fullest expression of DRY run wild. There are some notable examples in the NPM ecosystem where it reaches ridiculous levels. But putting the extremes aside, there's a ton of grey area and it can be pretty difficult to convince someone to write a bit more code when something else might work off the shelf. (And I mean this in the most charitable way possible. I find myself in that situation.)
I do hope we can turn the Rust ecosystem around and stop regularly having dependency trees with hundreds of crates, but it's going to be a long and difficult road. For example, not everyone even agrees with my perspective that this is actually a bad thing.
Thus a lot of things get done in Python and Java, because of the ample standard libraries. I was able, after a year of lobbying and procedures and approvals, to get a Rust compiler, but there is zero chance of me ever getting anything off crates.io.
.NET is an interesting case, because of WinForms, which is a de facto part of the standard library. I think few C# developers think of WinForms as bloat; it's very convenient for making simple UIs. Yet putting, say, GTK+, in the Go standard library would doubtless be considered bloat. I don't think there are easy answers to these questions.
It's not that convenient on Linux though...
I have more thoughts, but they are very hand wavy and ill-formed, so please take them with a grain of salt. One of my theories for why the Go standard library has had as much success as it has, is that it doesn't necessarily provide implementations that go as fast as reasonably possible, and that tends to give more flexibility for exposing simpler APIs. A good microcosm of this idea is JSON (de)serialization. Without even blinking, I can think of three reasonably popular third party JSON (de)serialization libraries in the Go ecosystem. The one provided by the standard library is pretty slow compared to some of them. It's not clear to me that it can be fixed without changing the API. But it's a good example where the standard library has provided something, but it isn't good enough in a lot of cases, so folks wind up bringing in a third party dependency for it anyway.
But even that alone isn't necessarily a bad thing. encoding/json is likely good enough for a really large number of use cases. On top of that, it's very convenient to use. (I'd still take serde in Rust over Go's system any day, but that's a different conversation.) And this kind of fits within Go norms pretty well. Go was never built to be the fastest, so the fact that some of its standard library has perhaps sacrificed performance for some API simplicity is totally consistent with that norm. And I don't think that norm is a bad thing.
There are some other examples where Go's standard library is slower than what it could be, for example, CSV parsing and walking a directory hierarchy.
Overall, I think the balance struck by Go's standard library was very nicely done. However, I'm not convinced it could have been replicated by Rust. The reasons for that are just guesses, and wander too far into musings about how the language is itself developed. But even putting that aside, Rust is going to have stricter requirements, because people tend to gravitate toward Rust when performance is important. So if std doesn't provide the fastest possible thing, then it's going to be a bigger deal than if Go does the same thing.
Again, above is super hand wavy and just a bunch of opinions from my own personal perspective.
It feels to me more like code incidental to other purposes than a library-as-library. Some folks like that and extracting from a concrete use is often instructive and efficient.
On the other hand, there is a truly maddening inconsistency in whether you get an interface or a physical struct.
One of these lends itself to easy replacement, injection and mocking.
The other lends itself to writing the nth-tillion interface wrapper for the parts of the standard library which use physical structs. Which are, naturally, slightly different from and therefore incompatible everyone else's bangzillionth interface wrapper for the parts of the standard library which use physical structs.
Then there's errors. But that's another day's rant.
That’s an issue, though it’s also one with most Go libraries too.
The thing is though, unnecessary interfaces also kind of suck. It makes code harder to follow, when the concrete type is hidden behind an interface. Also, one of the things that’s common in Go is testing with real implementations, rather than mocking - and I used to do just that. Even with Redis, I had a tiny shim server that implements the Redis protocol, that I would use in tests. Not absolutely everything can be done efficiently this way, but the virtues of testing with real clients are hard to ignore. Mocks and stubs can hide a lot of bugs that you would need to hope are caught by slower integration or e2e tests... And by virtue of being slower, they generally would cover less branches, too.
> Then there's errors. But that's another day's rant.
Have you kept up with the latest? I think they’re headed in the right direction with errors. Specifically with Is and As, along with the %w directive. The %w directive is a bit weird, but honestly, it’s a clever solution, and it seems like it would work.
Are you suggesting including the core `rand` crate with some traits and a basic pseudo-random generator in the stdlib, and allowing other crates to use those traits to implement the many other[1] RNGs?
I'm honestly not sure this is a good idea, it seems the rand crate has undergone a few iterations, crystalizing this inside the stdlib might lead to problems?
What is far better and faster than SHA3?
And while SHA2 wasn't designed for that use, it's easy to make simple and provably correct constructs that turn a secure hash into a secure RNG.
I would definitely not want my std to be designed this way.
2. You definitely don't want your std to be designed with a simple API to a fast, secure, and popular crypto primitive? Pretend I named your favorite one, to avoid bikeshedding issues.
I get that security people don't want non-secure RNGs to exist out of fear that someone might try implement security related functions out of them but why should I care if I just want to choose between 3 types of enemies this wave?
[0]: https://doc.rust-lang.org/std/collections/hash_map/struct.Ra...
[1]: https://docs.rs/rand/0.7.0/rand/rngs/struct.ThreadRng.html
Javascript's Math.random and Crypto.getRandomValues works this way.
> Be aware that the underlying Rust language will continue to evolve. Syn is able to accommodate most kinds of Rust grammar changes via the nonexhaustive enums and Verbatim variants in the syntax tree, but we will plan to put out new major versions on a 12 to 24 month cadence to incorporate ongoing language changes as needed.
With std I am bounded by the current language version.
Having a defacto standard library has more issues IMO than a real standard library.
With that said, if `syn` finds itself in a spot where it needs to do a breaking change release for a new language feature, then that would be tricky. It sounds like syn's architecture is pretty flexible (non-exhaustive enums), but it's not clear to me that it could support all possible language additions without breaking changes.
Trustworthy is another thing altogether though. How do you handle this process in C and C++? Both of those languages have fairly spartan standard libraries as well.
I just did our license audit and 100% of our shipped deps were Apache | MIT. If you can clear those two licenses with legal, you should be good to go for virtually any crate.
Might be worth releasing my one liner for this if anyone else finds it useful.
And you usually see them cursing the platforms that don't care about POSIX support.
And ISO C++ has repented themselves from following C's footsteps and have been improving the standard library since C++11, mostly by integrating boost libraries.
Also, while some of your historical context is interesting, I don't really appreciate your editorialization, which I often find is off the mark personally.
So in reality POSIX complements libc as C's "runtime platform", even though ISO C never considered to make libc that big.
My historical context is how I experienced how things went, throughout the media we had available at the time.
I surely welcome factual corrections when I fail off the mark.
Everyone benefits from learning proper history facts.
I wasn't. Even with POSIX, it is spartan by today's standards. Look at the standard libraries of Python and Go. POSIX doesn't have JSON (de)serialization, HTTP servers, XML parsing and a whole boatload of other crap. So I don't think there is anything wrong with my characterization.
That is the ultimate problem of batteries includes vs a small standard library distilled into a specific example.
People used to tout Python's "batteries included" line as positive almost universally a decade and more ago. Then better replacements were developed, and those batteries started looking less appealing.
There's really no escaping that while including any base functionality. Either you include it and portions will be stale later as better interfaces and paradigms are developed, or you don't, and you risk slower adoption and harder usability as people need to figure out solutions for common tasks, even if that solution is as simple as find the crate that provides it. Because eventually that becomes a problem not of finding the crate, but finding the best crate out of the multiple that exist, which itself causes fragmentation of the common developer experience (and thus makes it harder to share knowledge and have a good community).
A case in point is how the ORM Entity Framework that comes with .NET has made the older NHibernate (a separate package) obsolete.
.NET developers love using the standard libs, but OTOH Microsoft has a lot of resources to create very complete libraries.
In theory they can move into the standard library, but none ever have. The process exists though.
Another example was CompletableFutures which were inspired by ListenableFutures from Guava.
I can now use these with guarantees that they will be stable as Java has strong commitments to backwards compatibility.
I may not be a typical .Net developer, but my personal feeling is that this is too general and somewhat glib.
If you're inexperienced you will (and should!) choose the default option, if there's one available, and the most commonly used option if there isn't a default. So, if you wanted an ORM, then before EF there was NHibernate. But after a while you become able (from painfully gained experience) to determine what you want from a tool and what trade-offs you want to make, so you might use something like Dapper or no ORM at all.
Personally, I have found some of the MS provided implementations - shall we say - less than optimal. The Unity Framework. EF (which I dearly wish I have never had to suffer using). MSTest. Enterprise Library (oh god ugh).
So I make other - informed - choices about what to use. Some things I keep using because they are 'good enough' and I have a library of utilities and a mental map of how they work - log4net, NUnit - and some things I find I just don't need any more (mocking libraries, for one).
MS still support their provided implementations - as they should - but even their own projects can and do use third-party frameworks rather than the MS provided implementation (for example, Bot Framework used AutoFac (back when I was using it, anyway) [1]) - because their developers have been released from the requirement to exclusively use MS provided frameworks, and are making their own choices about what to use in their projects, and consequently what their users should use when using those tools.
Eventually, of course, if a tool is so crucial that it becomes part of .Net itself - the best example I can think of is dependency injection in .Net Core, which is in Microsoft.Extensions.DependencyInjection - you'd have to be mighty stubborn to use anything else.
One other point: I wouldn't describe NHibernate as 'obsolete', but given their historic and chronic inability to keep their documentation sites live and working ([2] referenced from [3]) it's easy to get that impression. But people are certainly still using it [4]. Just not as many as there used to be.
[1] https://github.com/microsoft/botframework-sdk/issues/938 [2] http://www.nhforge.org/doc/nh/en/index.html [3] https://ayende.com/blog/4139/nhibernate-documentation [4] https://stackoverflow.com/questions/tagged/nhibernate
This did happen over time, but NHibernate was still really popular for a long time after EF came out, because of limitations it had.
I also don't think it was entirely because EF existed - over time, EF implemented more and more features that NHibernate had, yet at the same time it seemed like the NHibernate team had given up - there were no updates to it for a long time. It was the lack of updates that moved me to EF, but I always preferred NHibernate.
I won’t add a library to our stack if we have an existing solution.
Most of our apps are written in C# and use Entity Framework with either ASP.NET MVC or Windows Forms.
I would need an excellent reason to use something else.
We don't have to just guess or make biased claims about what would be useful to move to the standard library. We can look at the numbers and see what people are actually using.
These metapackage communities can then focus on interoperability of constituent packages w/o overburdening the standard libraries, and core language can remain lean and concise.
On point of the post, Tidyverse is a great deal of bloat by loading unnecessary packages instead of specifics and creating namespace issues where two functions have the same name. It's generally fine for interactive work, but causes so many issues in development as you can't pick and choose what's loaded into the environment.
In line with what you're saying, meta-libraries make sense in terms of developing a line of packages towards a singular vision, but only make sense when the core libraries aren't expanded regularly. Maybe in these cases more of push needs to be made to supporting the existing tools?
R is an odd example though, as the standard libraries are loaded by default and I think only really comprise of basic stats/graphics/data.frame tools. I don't really believe a great deal of extra tools have been added to base R in the last few years, just improved in terms of speed and memory use (e.g 3.5.0's change to compiled packages)
The advantage of a standard library is that you only need to learn one API instead of a dozen different APIs for doing the same thing, which means you can develop a degree of mastery over it. It also reduces the friction for using better abstractions. e.g. Every professional Python programmer knows defaultdict, whereas I rarely see that data structure used in other programming languages, it's too much of a leap to install a dependency to save a few if statements, but it all adds up.
The rust ecosystem has done well to converge on certain crates as sort of replacement for missing std features.
In practice (at least in the rust ecosystem), I only need to learn one interface for:
* regex (regex)
* serialization (serde)
* network requests (request)
There are de-facto base crates in the ecosystem.
Because that is the biggest asset from stuff being in the standard library.
To clarify, there are different levels of support on different platforms.
The standard library and 3rd party crates generally have excellent compatibility across mainstream platforms.
libc is barely a standard library, it's almost nothing.
C++ surely does include regex, with serialisation and http scheduled for C++23, or with luck with a TR as intermediate delivery.
Although serialisation depends on static reflection being finalized as well.
I think the definition is meaningless. Rust spends huge amount of resources to test itself and its surrounding ecosystem on tier 1 platforms.
I'm fairly certain regex from c++ std will run like crap, if at all on something with 4MB of RAM.
I am fairly certain that without profiling and defining a test configuration for a set of specific C++ compilers / standard C++ library I won't assert anything about std::regex performance with 4 MB of RAM.
In other words, not all functions in the stdlib of Ada, Pascal, C and C++ can be used in all possible target environments? Sounds like a failure to quality gate those standard libraries.
There aren't an endless number of profiles.
Each platform has different levels of support. Primary being Windows, Mac and Linux, where every pure Rust crate runs.
Std lib makes certain reasonable assumptions, for which it works e.g. malloc exists and panic! is implenented.
Looking at crates.io, regex looks pretty safe, as it’s authored by “The Rust Project Developers” and includes explicit future compatibility policies. Unfortunately, I can’t find an index of only the crates maintained by the Rust team.
Serde is obviously popular, but at first glance is a giant Swiss Army knife that will likely have lots of updates to keep track of that are completely unrelated to my project (whatever it is). If I search for JSON, I get an exact match result of the json crate, followed by a bunch of serde-adjacent crates, but not serde itself.
Request hasn’t been updated in 4 years, and has a total of less than 7000 downloads.
Not everyone or every project will have the same desires, though. Sometimes, a fast-moving experimental library is the right choice. The trouble is figuring out which I’m looking at.
If you'll want to update to always be on the latest version of each crate, well that discomfort about them potentially not working is part of the price.
Not everyone does, and that’s fine. I just want to know what a library developer’s stance on it is before I try to use their library.
I agree that this should be better documented and probably more integrated with crates.io somehow.
All these libraries are very well known within the community and are what I would come up with as a complete outsider (I don't think I've written more than a hundred lines of Rust code to this date).
You can also find some pointers here:
When talking of the standard library, static typing doesn't save you when breakage happens. It's better of course, at least the compiler protects you from obvious errors (although that doesn't work for transitive dependencies, with the dynamic linking to binaries that Java / the JVM does ;-))
The problem is when a piece of code that was compiling fine a year ago, fails to compile on a newer version of the standard library, due to breaking changes, that's going to take time and effort to fix.
And this gets worse when the breakage happens in dependencies and those dependencies are no longer maintained. This can always happen of course, not just due to the standard library, but due to transitive dependencies too. But still, breakage in the standard library, or in libraries that people depend on, is a bad thing. And consider that as the number of dependencies grows, so does the probability for having dependencies that are incompatible with one another (compiled against different versions of the same dependencies).
And semantic versioning doesn't work. Breaking compatibility will inflict pain on your downstream users, no matter how many processes you have in place for communicating it. And this is especially painful when you're talking about the standard library.
If the standard library introduces breaking changes, regardless if the language is static or dynamic, then it's not a standard library that you can trust. Period.
Also — when should you break compatibility, in the standard library or in any other library?
The answer should be never!. When we want to change things, we should change the namespace and thus publish an entirely new library that can be used alongside the old one. Unfortunately this isn't a widely held viewed, but I wish it was.
---
Going back to the batteries included aspect of some standard libraries, like that of Python, there's one effect that I don't like and that's not very visible in Python since the bar is pretty low there.
The standard library actively discourages alternatives.
When a piece of functionality from the standard library is good enough, it's going to discourage alternatives from the ecosystem that could be much better.
Some pieces of functionality definitely deserve to be "standard". Collections for example, yes, should be standard, because libraries communicate between themselves via collections. And that's what the primary purpose of a standard library is ... interoperability. Anything else is a liability.
https://docs.oracle.com/javase/8/docs/api/java/util/Map.html...
I also remember from years ago that Bundler in Ruby allowed version mismatches to coexist and would install both deps in the same tree, punting any runtime issues (same as NPM).
Any Gentoo or NIX users might have something to add here, as I can remember being regaled at conferences by them about this topic.
All that said, it would seem that Rust could have some kind of super-intelligence about dependencies due to all the static goodness. So my questions are:
a) who cares how much is userland and how much is std lib if everything is equally safe and documentable?
b) if "too much" becomes std lib, could any feelings of overwhelmingness not be mitigated with more namespacing?
But in the modern world, I agree with the small language people: the logistical problem is dead. Almost every modern language has push-button dependency management, so the difficulty overhead of using a third-party library is basically zero. And, as you pointed out, this means that the stdlib now competes with third-party libraries, and the third-party libraries proliferate wildly. So there's still a problem, it's just not logistical anymore.
The other problem a stdlib solves is making decisions. This is the npm nightmare: which of these fifty libraries do I use? They're easy to install, but hard to evaluate. This compounds because every library is also making those decisions with their dependencies and so on, so you can end up with 8 different implementations of basically the same thing. You don't have this problem with a stdlib because if there's a std::string, everyone expects your code to work with std::string.
So perhaps the modern stdlib is better off as a standards library: no code, just a mapping from name (ie "option_parser") to package@version (ie "clap@1.1.0"). The work of the stdlib would then be to curate this mapping in a cohesive way, so the packages are high quality individually, but also work well together and reflect the general direction of the language and its community. Whether or not these packages are actually bundled with the compiler is really just an implementation detail. The library isn't code; it's decisions.
Updates would be pinned to release versions, so backwards-incompatible changes coincide with a major version/edition/etc of the language. This would be an expected process, because, as with urllib{n+1}, the first answer isn't always the best. Third-party packages are just as available as ever, of course, and a destandardised package is only a (trivially automated) rename away.
Aside from providing a layer of simplicity and stability over the third-party ecosystem, the standardised libraries would serve as reference implementations for a common interface, like Node's express/connect, or Rust's tokio/futures. In other words, it could help concentrate community effort around emerging standards.
I grant that this is a very, uh, political approach to what has historically been a code problem, but if anything I think the current situation owes itself largely to treating community problems ("how do we agree on a foundational layer of library code?") with technical solutions (package registry + sort by stars).
For a language like Rust this looks different..
That beeing said, I think the whole Rust userbase would benefit from having some sort of collection of well tested crates and strictly divide between private, work in progress and production crates.
What they should've done is allow uploads only to "<username>/<cratename>". So if I decide to make a regex crate, it's called "majewsky/regex" at first and Alice can make "alice/regex" and Bob can make "bob/regex". Then at some point, the community (through some open process) decides that Alice's crate has the best API, so "alice/regex" gets aliased to just "regex".
That way, everything in the main namespace adheres to some sort of quality standard and has community support behind it. Because there's some explicit process gate to getting stuff into the main namespace, you could attach any number of beneficial requirements to it, e.g.:
- test coverage
- documentation coverage
- at least 3 people having committer rights to the repo, at least 2 of which must not be affiliated with the same company
But the underlying idea is sound: namespace and federation. Rust should ideally follow Java's solution. And it's completely forward portable: when namespacing is introduced, the crates.io/regex crates become io.crates.regex. Then you can have your com.github.majewsky.regex.
Rust does need a better package curation story, a Rust "expanded universe" of recommended packages whose provenance is more carefully tracked than the average package and whose APIs are stable. The Rust community needs to encourage other packages to depend on stable versions in that recommended set to reduce duplication.
But going further and actually making that set to be part of the "standard library", and thus released at the cadence of the Rust language itself, and managed under the same umbrella, would be harmful.
I could pin every dependency, and monitor CVEs for all my dependencies, and monitor the ownership/code changes for all my dependencies, and hope for updates in a timely manner should there be CVEs... or I can use the standard library and move on with my life.
Personal opinion of course. Some people enjoy monitoring CVE lists. :)
For instance, http client and server libraries are often in this gray area of uncertainty about whether they should be in the stdlib or not. Is this something language designers or compiler implementors have a lot of experience in? I would say not; sending and serving http requests are not something compilers need to do. Or take GUI libraries, what do language designers know about that?
I also know of many bad third party libraries, but there are tons of examples of really awful parts of standard libraries. The original date/time APIs in Java are a mess (they finally fixed this by in essence bringing in a third party API). The ssl bindings in the Ruby stdlib were a common source of bugs back when I was paying attention to this (maybe they've fixed it), same thing for the built in http stuff. Someone else mentioned the similar weakness of python's built in http, such that most people use a third party library instead. Even the java collections APIs are pretty poor such that people often augment them with things like guava or apache libraries.
My point is just that developing good libraries is a hard thing and I don't see any reason to think language designers or compiler maintainers are any better (or worse!) at it than other people. There isn't really a shortcut, you can't just cede authority to the powers that be on the core language teams, you just have to evaluate the quality of libraries for your use case yourself.
I propose that the dictionary entry for 'faint praise' be updated to include this example.
We used to run a whole gamestate of a shipped title on PSP in a 400kb block allocation. I've yet to see that in any other dynamic language of consequence.
One of the ideas I wanted very much to explore is scaling the API, both up and down. For building something akin to PyPy, you might want a 'kernel', a small set of libraries that were available everywhere, and several other levels that include more or different things.
Mobile takes the base and adds a few things suitable for mobile (storage, UI, broader networking). Desktop has a real UI, and then there's the kitchen sink like .Net and Java have.
But you have the same problems you always have with decomposition - if you didn't guess the right boundaries when you built the thing then removing or rearranging bits is a serious PITA. Sometimes I think the best we can hope for is to leave clear messages for the next language so that it doesn't organize things the way we did.
I think this was only possible because Lua itself is such a simple language. It says a lot that LuaJIT's implementation of this simple language is extremely complex in comparison to vanilla 5.1. Imagine the added complexity for something like Ruby that offers more than one associative data structure.
With LOVE you get a complete game engine runtime with all the boring OS abstractions taken care of and you can just start working.
I tried a few game development bouts with Rust since it seemed like a natural step up from C++ but the compile cycle kills a lot of my creative drive.
I think this is just a consequence of the language being compiled, as with C++. With Rust having a wide variety of language features the compile time is increased also. At least we're lucky to have languages today that offer far shorter turnaround times for building prototypes.
Completely agree. I hold firm to the believe that a strong standard library that considers modern software development goals will drive the success of a language.
Look at Go. Which is usually my example since not many do what Go does. You can do a whole web application in Go with minimal use (if any) of external libraries.
Back to Rust and their defense: they intend to adopt popular / quality packages from Cargo into the standard library or fill the gap between packages.
I secretly wish D had similar things to Go in the standard library. HTTP and SMTP etc which Python seems to have at least HTTP and sadly Go ditched their SMTP package for whatever reason making it awkward.
I hate having to learn a new package manager per new language. I love programming so I try a lot of languages out of love. Package managers and build systems are horrible UX in every language. I prefer to not rely on third party oh look now deprecated packages and start out with whats out of the box.
Libraries can be implemented by 3rd parties and maybe later adopted as standard libraries or become de-facto standards. That's how it has been done for most popular languages out there, and I think it's a smart decision.
Rust has a much broader aim, so a stdlib accommodating all of its usecases would be as comically huge as Python's, with all the problems that come from that.
"a stdlib accommodating all of its usecases would be as comically huge as Python's, with all the problems that come from that."
That's a straw man. I don't want a stdlib like Python's, I want one like Go's.