Tech Debt: My Rust Library Is Now a CDO
lucumr.pocoo.org
lucumr.pocoo.org
It's not exactly a left-pad situation, because the package remains fully functional (and crates.io doesn't allow deletions), but the package is used in 4000 other crates. Audit and auto-update tools will be complaining about use of an unmaintained crate.
FTFY. And vendoring a dependency is really just a hidden fork.
It's not a fork until you start diverging from upstream with local modifications.
I mean, implementing your fixes in the first place means you are already inherently forking it.
1. You find a bug, open an issue. 2. If bug is a minor issue, wait. If I want a fix really bad, offer to fix it. 3. If I still want a fix really bad and author did not respond, or does not want a fix, then fork the project for my own consumption. Using alternative dependencies are also an option.
Imo be ready to fork all the things if need be. Get comfortable with it before you're forced to.
cargo update
without reviewing the code changes in each dep (and their deps) wrt how they impact our compiled artifact. Ultimately we're each individually responsible for what we ship, and there's no amount of bureaucracy that will absolve us of the responsibility.For example I can’t imagine Ford checks literally every single individual component that gets put in their cars, right? I assume they have contracts and standards with manufacturers that make the parts, probably including random spot checks, but at some point to handle complexity it is necessary to trust/offload responsibility on some suppliers.
I bet you could get those sort of guarantees from, like, Intel or IBM if people generally were willing to pay them enough (can you now? Not sure, probably depends on the specific package). But nobody would make that sort of promise for free. The transparency of the open source ecosystem is an alternative to that sort of thing, no promises but you can check everything.
Just to be explicit (since I’m sure this looks like needless pedantry), I think it is important to remember that the open source and free software ecosystems are an active rejection of that sort of liability management framework, because that’s the only way to get contributions from a community of volunteers.
As a software assembly designer, I don't really have any way to say "this subcomponent remains intact" when I integrate it into my system. The compiler will chew things up and rearrange them at link time. This is a good thing! It makes the resulting binary go brr real good. But it's ultimately my responsibility to make sure the thing I ship--the compiled artifact--is correct. It's not really a valid move to blame one of my dependencies. I was supposed to review those dependencies and make sure they work in the context of the thing I shipped. Wouldn't it be a bit like the mechanical assembly designer blaming the bolt manufacturer for a failure of a bolt they modified while building the assembly (EDIT: or blaming the supplier when the assembler used the bolt in a way it wasn't designed for)? What guarantees can the bolt manufacturer reasonably give in that case?
EDIT: another thing is that source code isn't really analogous to a mechanical component. It's more like the design specification or blueprint for... something that converts electrical potential into heat and also does other stuff.
So yes, I `cargo update`.
The release notes for the last release explain: https://github.com/dtolnay/serde-yaml/releases/tag/0.9.34
This is an unrealistic and frankly a bit entitled of a criticism.
This is before we even get to stuff like “it’s not like yaml is a fast moving target, what continued maintenance is really even needed” parts of the equation, or the comment above that said he’s tried already.
We should remember that an unmaintained dependency isn’t the worst thing that can happen to a supply chain. There are far worse nightmare scenarios that involve exfiltrated keys, ransomewared files and such.
I’ll bet that if someone with a track record of contributing to the Rust community steps up, he’ll happily transfer control of the crate. But he’s not just going to assign it to some random internet user and put the whole ecosystem at risk.
I think he was generous enough with his time for creating the library in the first place and posting this update. It's not like he's owing anything to anybody. I don't understand how you can demand him to do further UNPAID work to recruit maintainers who would also do further unpaid work.
Why can't this "someone else" just fork the repo themselves?
Are we collectively having Subversion phantom pains or something? Just take it.
You can't. Ever. This is unironically why languages too old for package repositories like C and C++ lend themselves to stabler, higher-quality software.
On top of that is also the entitlement to feel like people providing you software for free MUST ALSO be responsible for polishing it forever. But, hey, I guess that attitude comes naturally with the package repository mindset. ;D People need to do all my work for me!
> stabler, higher-quality software.
Lol I’ll have some too buddy
https://github.com/dtolnay/serde-yaml/releases/tag/0.9.34
> As of this release, I am not planning to publish further versions of serde_yaml as none of my projects have been using YAML for a long time, so I have archived the GitHub repo and marked the crate deprecated in the version number.
I don't think anyone is claiming he is legally required to do it. Just that it would be nice. But in this case I think he did try to look for someone anyway.
Maybe this description is wrong but I can’t tell.
stuff -> insta
learned-rust-this-way -> yaml-rust
to default/defaulted -> serde-yaml
As I read it: insta was depending on yaml-rust. RUSTSEC issued an advisory against yaml-rust. Some shade is thrown on yaml-rust for being unmaintained and basically a student project of someone learning the language for the first time, adding to the irony/brittleness of its unmaintained (though still working) state and now the RUSTSEC callout. In insta trying to switch to some alternative to yaml-rust the next most common library serde-yaml preemptively announced its unmaintained status. So insta just vendored yaml-rust and called it a day, which currently works to remove the RUSTSEC complaints but doesn't fix the long term maintenance problem.
Not the case here but sometimes a library (with no dependencies) is really feature complete with no real need for changes other than security fixes.
I wish audit tools and corporate policies are smart enough to recognize this instead of nagging about "unmaintained" versions of projects with no recent commits.
The ball is in the court of the enterprise purchaser mandating the auditing company and they don’t give a ** if your dev team has to hot swap a yaml parsing library
So you have security theater and TPS reports consuming the time of some of your most highly compensated human resources.
Agreed but the gotcha is when a new issue comes out and there is no security fix available. It would be enough to say 'we are comfortable patching this if a CVE is found', but at that point they might as well start to maintain it.
The issue is that the context is always changing. It is almost like suggesting there is an evolutionarily perfect organism. At best, one suited perfectly well to the current context can exist, but that perfect suitability falls off the second the environment changes. And the environment always changes.
I totally get the sentiment, but I think the idea that any code is ever complete is unfortunately just not a realistic one to apply outside of select and narrow contexts.
It's not like it's desirable for standard formats to one day start getting interpreted in a different way.
You depend on calling a standard library function that was found deficient (maybe it's impossible to make it secure) and deprecated. Now there is a new function you should call. Your software doesn't work anymore when the deprecated function is removed.
Sure, you can say your software is feature complete but you have to specify the external environment it's supposed to run on. And THAT is always changing.
You're both right but looking at different timelines.
Relevant: https://www.oreilly.com/library/view/software-engineering-at...
That's a biiiiig if
There are Clojure packages that go untouched for years... not because they're abandoned, but because they are stable and good enough.
You seem to be misunderstanding the comic you link.
With him that was the situation.
On crates.io, serde_yaml has 56m lifetime downloads and 8.5m recent downloads, and from his own announcement David hasn't used yaml in any of his project in a while, yet has kept plugging at it until now.
And yaml-rust has 53.5m lifetime downloads and 5m recent downloads, and has been unmaintained for 4 years.
Why would he?
https://en.wikipedia.org/wiki/Collateralized_debt_obligation
My first guess was "chief data officer".
Specifically, it was a piece of partial ownership on a bunch of mortgages in 2008. The issue here is that when the junk mortgages defaulted, the CDO also went sideways [1]
In any case, what the author is thinking about is more "systemic risk" (think leftpad or cascading DNS failures) rather than debt-based analogies. One big reason 2008 became such a big mess was the systemic issues, not building products on debt [2]
1. Fun fact - before 2008, they had started making CDOs out of CDOs because of of the insatiable appetite for mortgage-backed financial products. When mortgages are underwritten for people withouth a job, and packaged in a product with a AAA credit rating, it turns out it's a profitable thing to sell.
2. Also the fact that Lehman and Bear Sterns had leverage ratios in the 30x-40x range, which is insane.
Whereas in the 2018 crisis, junk debt was rolled into CDOs but there were no 'knobs' the bankers could turn to actually improve the fundamentals of the junk debt. By repackaging the debt they did not get any right to fix the bad actors that made it bad. Things were actually bad, and when it defaulted there was insufficient firewall to prevent the contagion from affecting the encapsulating derivative.
Of course it is up to the maintainer but by absorbing an abandoned project, they are actually reducing the amount of unmaintained code, versus simply hiding it.
Probably the more accurate analogy to 2018 would be if a private equity firm started a company you could pay to upgrade the rating of an unmaintained Rust package, but no actual change happened and all they did was swap labels on things until the problematic rating disappeared.
Another advantage of vendoring like that is that you do have some tools for attacking the code now. If you have solid test coverage on your own library, you can now run your library and run code coverage tools on the newly-vendored library. Fixing the library may be a challenge, but you might be able to just go hacking into it and deleting anything your code doesn't touch relatively easily. Depends on the structure of the vendored library, of course. If you end up deleting all the vulnerable code, you do have a clear win. If you discover you were hitting some of it after all you may end up with an even more clear win.
Or, to put it another way, forking a library in general for public maintenance is a big commitment. Vendoring it and hacking it to just the bits you need is much smaller. Even if you don't actually perform this pruning, just the act of making that pruning easier is itself progress.
There are disadvantages too, certainly, but it's not all bad.
The fact the dependency was marked as abandoned and started getting flagged in security reports is what you want. That lets you make an informed decision about whether to vendor it (which removes the "bad actor sneaks something in" risk) or do something else. It's great that the Rust build system allows that so easily.
Given that the description of the problem includes generally not being affected by its issue I read that to be the rather common case in which one is just using a relatively small fraction of the library. To be concrete, let's say it's an image library, and it turns out there's a PNG parsing vulnerability, but your code only ever uses it for GIFs. If you vendor it you also have the opportunity to go in and just yank PNG, JPEG, TIFF, etc. support, and depending on the situation, plugin support (maybe you can trivially tweak your code to go directly to the JPEG support without passing through anything else).
I'm not saying this is a perfect solution with no downsides, but it's not all downside either.
Of course, you need the aforementioned solid test coverage. Without that you're just flying blind. That's another thing about the debt metaphor... it can compound, but unlike most monetary debt, code debt can compound on you unexpectedly.
I get the sense (only a sense because the author is a bit coy, but sibling comment by him seems to confirm) that this was more to appease these tools that look for “bad dependencies”.
This has the effect of keeping things relatively sane.
But this enforced discipline can work well at an org like Google that has plenty of time and $$ and isn't dedicated to a "move fast and break things" philosophy. Not sure how well it would work for a startup.
For the why, could be that the build would be dog slow if a network requests was made for every one of a bajillion dependencies. Or to avoid breaking the build if something is janked externally. Could be because C++ has no modules/libraries. Could be for the lawyer's sake.
Is any of that applicable? There are also other ways of vendoring Rust crates, e.g. mirroring them to an internal registry. That can have fewer downsides.
There are many many reasons but overall sanity and security concerns are paramount.
But also, read the original article referred to here... and the pain points seen there. Be careful about what you depend on.
I mean, most Linux distributions do the same thing - anything that's "vendored" as part of some project is supposed to be patched out from the source and pointed to the single system-wide version of that dependency. And Debian is packaging lots of Rust crates and Golang modules as build-time dependencies within their archive.
Could be an interesting test strategy, in some sense it’s directed fuzz testing. (With other libraries providing the random-ish invocations.)
Npm audit predominantly cries wolf on security issues and rolling things in, license permitting, is one of the most reliable ways of not getting inundated with bogus issues from users (who either don’t understand the context or who don’t care because their employers’ policies around these things was created in a space divorced from reality.)
(No, that regex being used as a part of your build pipeline’s transitive dependencies’ codebase is not actually something which can be exploited for a DoS attack.)
The ”issues” for deeply transitive dependency can be especially annoying to work around. Trouble is that the way things are structured you often can’t technically establish ”no, we will never enter that codepath so aren’t affected by flaws in it” or ”the only way we enter that codepath is with trusted input in an offline setting”, etc.
And even then, you could enter the problematic paths in future releases as your dependencies update.
Yeah, issues in dev dependencies are such a headache in the JS world. There exist scenarios in which these matter (a compromised build tool injecting malicious code into the lib you’re building) but they are vanishingly rare and hidden under the tidal wave of DOSable regexes that don’t matter at all when you just call them in your build. Couple that with typical build tools having a transitive tree of 50 gabillion dependencies and it’s a real slog. I think the tooling for reporting on these issues needs to differentiate “exploitable if you redistribute” vs “exploitable if you use in a build pipeline”.
I think by "suddenly", the author is saying it doesn't make sense for the same code to have a better debt rating when vendored than it did before. But that's only looking at the value of the code itself, which misses the most important part of the total value proposition.
When a maintainer vendors code in, the maintainer owns that code now. If you're an active maintainer vendoring in code from a dead project, you are increasing the value of that code because now there is an active human who may respond to issues, review pull requests, and fix bugs.
By another analogy, giving a neglected pet to a new owner increases the value of the pet because it will be better taken care of, healthier, live longer, etc.
If a package is deemed important enough for the larger community, there needs to be a way to rip the package registry entry out of the original owner's hands and to make it point to a fork. Natually, such a move will often involve a lot of drama, but it can actively protect downstream users.
This seems to be a non-issue. The beauty of open source is that you can fork it. The actual issue is that it requires time and effort to maintain an open source project and it's not easy to find someone else who has some time to spare.
This chaos can be reduced with a single "blessed" fork that actively accepts reasonable maintainance patches. There are fewer incentives for competing forks and dependencies (direct and transitive) of the original package can automatically benefit, too, because releases are made under the original package name. This is good because the neglect problem can be several levels deep.
There is the issue of manpower. There are two arguments here: one is that the bar for package takeover needs to be set such that community interest in a fork is high enough and the fork needs to be managed by the package registry owner for several reasons. One is that maintainers need to be appointed through the registry provider to keep the nodel functional. Maintainer appointments can be passed on after extended inactivity. The other argument is that a package registry provider puts themselves into a trusted position by default and has to actively maintain that trust. Thus, the burden of managing forks and fork maintainers shouldn't necessarily be seen as inappropriate.
Though an argument could be made that github is a host of free software dedicated to the copyright infringement of said software, at least it doesn't currently try to usurp the nominal ownership of the packages.
The package registry entry isn't the package. At no time is the original creator deprived of their rights to their package. Any fork will have to respect whatever license the original creator assigned. Open source licenses are allow forking by design for a reason.
Also, this whole thing needs to rest on a transparent process that doesn't make takeovers easy. There need to be clear criteria when takeovers can even be initiated, like extended inactivity (think 6 months or more) or unresponsive maintainers when serious issues are encountered. Also, there is not much harm in a mechanism for reclamation by the original package owner.
https://wiki.debian.org/NonMaintainerUpload
also, names in a registry should be "leased" not forever eternal perpetual immutable assigned to some guy
It may be done at package level at best but not at registry level. Since we are talking community it is not a customer/vendor relationship, most packages are just important enough to be available for free. It may also be at some humiliating amount like "here is 10 dollar/month now maintain 100K packages for us.
There’s lots of source code on the Internet that’s not maintained; it’s just as-is, one-and-done. Nobody is guaranteed updates for free code.
Ultimately, the problem remains - we're dependent on the work of others who are happy to provide us with free labour for now, but may not be in perpetuity. I don't know that there's a way around that other than compensating them for their time and effort and hope they continue their good work.
yaml-rust was a pure-rust implementation, that's literally in the tagline:
> A pure rust YAML implementation.
serde_yaml was arguably less pure-rust, as it relied on unsafe-libyaml, which is a c2rust translation of libyaml.
Any project with a bus factor of 1 is inherently dangerous and risky. On a long enough timescale, the probability that an open source maintainer will quit their project is 1. The only way around that is to make solo-maintained projects taboo.
For me, the go-to solution for this problem is to avoid third-party dependencies as much as possible. This is probably not an option for rust project given how anemic the authors of Rust's standard library want their library to be, but it is definitely an option when using other languages with fully featured libraries.
I rarely use external libraries (other that database drivers) because I deliberately work with languages that have all the things I need built in.
Code doesn't go bad on its own. I've borrowed or integrated code from decades ago just fine many times.
Personally I'd just ignore any and all complaints about that library and move on.
Isn't that what the author did by vendoring the library within his own project? It amounts to taking over maintenance for the bits of code you use. Sometimes this is bad (copy-and-paste coding, possibly duplicated effort) but when the third-party component is totally unmaintained it becomes quite excusable.
Having said that, I've not worked in languages where "batteries are not included"
cargo vendor --aggressive
which might prune all the code from your deps that is "dead" to your crate? Would it make the "review your deps" problem more tractable? This is a tangent to the article, but related in that ultimately we're responsible for our choice of dependencies and all that entails. It seems like there might be room for tooling to help us better take responsibility for what actually gets compiled into our crates.Basically I want something which eliminates all the dead code in my dependency graph, and makes me review the diff in the undead code every time the dependency graph changes, my code changes, or both. Otherwise I'm very unlikely to review my deps and their usage. If it doesn't appear as an actual diff in code review it's not getting reviewed, probably.
- Anyone know this guy?
- Has he logged on in the last year?
- Can someone contact his aunt in Xinjiang? Maybe she knows something
I experienced the same thing with a repository that wasn’t active. I mean a repository might be feature complete so why should it be active? Well other people wanted more from the repository. But not the author apparently. That’s fine. Well turns out that there was an active fork which was vastly more “featureful” by design. And a PR to that repository took less than three hours to be merged. How did I find that fork? Honestly I think it was mostly by coincidence since I saw the fork owner posting on some issue in the original repository.
Spacemacs was experiencing a reviewer bottleneck for a while (I don’t know what the status is now) in part because the repository owner was not very active. That got partly solved by more people getting commit rights apparently. But why don’t forks pick up the slack instead of waiting around for commit rights? Hijack the PRs and review them in your own forks. Merge them if they are good. Then come back to the upstream and ask the main guy to merge them. If he trusts you and would have given you commit rights anyway then it shouldn’t be a problem.
GitHub should prominently tell you about other forks if they are notable in some way:
- More activity (most important)
- Maybe more stars
Because then you don’t have to hope to God that the repository owner logs on every two weeks and keeps his thirty repositories updated about their maintenance status. (Honestly who doesn’t have dusty old code that works for them which you haven’t even thought about for months or years?)
Why is this so difficult? All repositories are peers after all.
Armin's answer here, of taking on explicit ownership and responsibility for the third-party code they use, seems like a reasonable and socially desirable response. Yeah, I get that it's annoying, but this is (IMO) pretty much how open source is supposed to work.
Nonetheless, I think RustSec's feed is working as intended, and I'd put the expectation on the consuming tools (including `cargo-audit`, which I believe is largely maintained by the same people that maintain the RustSec vulnerability database) to change.
It seems like he vendored it in order to avoid the useless alarm bells going off. Don’t we have a `Cargo.lock`? Just keeping the status quo seems fine.
But other people didn’t like it because of these automated alarm bells, hence it becomes “socially desirable” to do something that doesn’t materially make a difference.
Separately, I do hope VEX becomes more real / takes off to help address issues of non-applicability for vulnerabilities. That said, I don’t think this is a non-applicability scenario.
I’m not sold on the argument in the other comment. It’s just a may. Something is not worth more just because there is a potential for something.
Is he fixing the bugs reported against the defunct upstream? If not, what is the improvement? And if yes, how can will the rest of the library users benefit from those fixes?
Here's a quote from the project's license:
> 7. Disclaimer of Warranty. Unless required by applicable law or agreed to in writing, Licensor provides the Work (and each Contributor provides its Contributions) on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied, including, without limitation, any warranties or conditions of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A PARTICULAR PURPOSE. You are solely responsible for determining the appropriateness of using or redistributing the Work and assume any risks associated with Your exercise of permissions under this License.
Make enough changes and it’s not vendoring anymore; you made the code your own.
Or am I spoiled by .net, which has an excellent bcl?
That said, you’re right that “to keep binary size down” isn’t why the standard library is small. Heck, that you don’t compile from source means sometimes binaries are larger than they have to be, even with LTO! It is largely because anything that’s there has to be supported exactly as is forever, interface wise, and for many things, we did not have the confidence on what the right interface would be. Furthermore, it’s harder to maintain the standard library than it is a package out of tree for a variety of reasons.
~> rustup component list | rg src
rust-src (installed)
But this is so that that workflow works well, not because the compiler actually uses it during the build.> the compiler is smart enough to not be over-compiling unused symbols.
This is true even without the source, mind you, but the devil is just in the details. Stuff like the panic and/or formatting machinery is known to creep in even if you don't think you're using it, because some edge case might. This only truly matters for embedded and similar, but it's still a thing.
I think in Rust the choice is more one that originally prioritized viability of long-term maintenance of the project by keeping the volume and surface of the standard library small. For all more recent inclusion the inclusion criteria seem to be either "small ecosystem functionality that everybody is using and has proven to be valuable" (e.g. once_cell) or "what is close-to-optimal solution for critical parts of the ecosystem that we have to include to minimize ecosystem splintering" (e.g. the work around std::future). I wouldn't be surprised if there are actual written down criteria for std inclusion somewhere.
I’m not aware of the current state, but in the old days, it was roughly:
1. Is this useful in virtually every Rust program?
2. Is this something (data structures, mostly) that requires a lot of unsafe to implement?
3. Is this a vocabulary type?
Any of these is a slam dunk, anything else is more subjective.
That's too strict a criterion for a standard library!
You can tell that this isn’t literally true just by looking at the contents of the standard library.
MS has massive resources and can align the many many modules of .NET into major releases. The Rust project does not.
My knowledge about Rust's library system isn't great so there are probably several gaps in my argument.
There are crate "recommendation lists" like this: https://blessed.rs/crates
While these help, Rust's small standard library has slowed Rust adoption where I work because our IT security team must approve all third-party packages before we can use them. (Who blindly downloads code from the Internet and incorporates it in their software?)
This means that unless there is a compelling need to use Rust, most teams will use C# because .NET is approved and has almost everything. I prefer Rust to C#, but we are not using it that widely yet.
It's not a fair comparison, but it is how things currently are.
Do they also review the original tooling? Why would one single out third-party packages?
Everything used for software development is explicitly approved. This includes programming languages, compilers, debuggers, IDEs, etc. We are primarily a Microsoft shop, so the majority of our development tools follow that direction.
For FLOSS libraries, the approval process covers both IT security review, a source code scan / static analysis, as well as a legal review of the package's license to ensure it is not on the prohibited list.
This makes management happy since it prevents potential security and legal issues. It keeps our customers happy since they get quality software made from fully traceable components.
When a CVE is announced, we know immediately if we are impacted and what will need to be fixed.
Some places have no idea what their dependencies are. I am sure there are lots of log4j horror stories from Java shops that were not so careful.
- WinForms, WPF, UWP, WinUI 3.0/WinAppSDK, MAUI
- WebForms, MVC Framework style, MVC Core style, Razor Pages, Blazor, Hybrid Blazor (tryig to fit everywhere there is a Webwidget from the previous line)
- Latest design decisions being driven by Azure and ASP.NET requirements
Stuff like minimal APIs driving C# features, e.g. lambda improvements required for replacing controllers with minimal APIs.
There are hardly any desktop or COM improvements that are taken as seriously as the way the ASP.NET is.
EF Core works best with Microsoft offerings, SQL Server, CosmoDB,....
The whole aspect oriented framework now in preview to enable AOT for ASP.NET workloads.
Again: it is not a part of the language or runtime. If you do, however, have links to code changes that prove otherwise - please do post them.
Meanwhile, COM interests only a limited part of community, and there also exists actively maintained https://github.com/microsoft/CsWin32
A popular (if not the most, depending on location) choice to use EF Core with is PostgreSQL.
Meanwhile, ASP.NET Core simply continues to evolve just like other libraries, whether they are a part of dotnet/* or not.
GUI continues to get better with development of AvaloniaUI, Uno and MAUI.
.NET as a whole is way bigger than some of it's "first-party" solutions that you don't even have to use, that have little impact on language, runtime and compiler themselves aside from "this is a good demonstration where we can improve the general case". There also exists Unity, many companies using completely non-standard stacks on top of the language, etc. But you deliberately choose to ignore this information for unknown reasons.
It is genuinely bewildering to see this kind of flawed and strange impression of the ecosystem not rooted in reality. I suppose addressing this involves reading documentation and source material, which perhaps proves too challenging.
The difference between us, is that my professional existence is not tied to praise .NET folks for everything they do.
Everything that Microsoft ships on Visual Studio is part of .NET, without yes and buts.
The biggest problem is the exploding transitive dependency graph, including multiple versions of the same basic utilities.
Even small projects can suffer from this. It's too easy to reach into the grab-bag at crates.io and pull a little utility out and use it. I've personally built very small things only to find multiple versions or implementations of e.g. sha256 or rand or whatever in my own dep graph, and it gets really hard to prune them out.
At work we're fairly diligent about this but still it's exploding. And when you consider the matrix of minimum-supported toolchain versions for different projects, and if you're stuck on an older Rust toolchain, it's even more pain.
Rust projects starting from scratch could consider vendoring their dependencies rather than pulling from crates.io. But support for vendored dependency in Cargo is weak/bad. You run into all sorts of nightmares especially when using workspaces.
It's heresy and I'll get downvoted for it, but: More and more I lean towards using bazel or gn/ninja for fresh Rust projects, and walk away from cargo & crates entirely.
Cargo (+ Crates.io) was delicious to me at first, but it's a meal that tastes worse the more I have of it.
EDIT: it strikes me that I didn't say why exploding dependencies bugs me so much. There are logistical reasons for sure (competing APIs, compile fails, binary bloat, etc.) but the primary reason is: exploding complexity. Every additional dependency is now a potential problem for future me/future team.
This is the first step of your journey.
Next step is realizing the same about the language itself.
However, I also think the vast majority of Rust projects out there right now are using the wrong tool. They're using Rust to write web services and web sites, and the developers are coming from places like NodeJS, TypeScript, etc. It'd be fine, but it's I believe pushing the emphasis in the Rust language community in that direction.
I worked for years in the relatively sane subset of C++ that we used at Google. It's fine, I have to do C++ here and there for work, and I could even go back to it if I had to. But I mostly like Rust, and would miss it. Zig made some decent choices, but others I don't necessarily agree with, but I'd consider it for embedded or low-level dev of various types
Of course they are, same as what had happened with (to) C++. And the people that use Rust (or C++) because it's "fast" are mostly the same that write code which would be faster if it would have been written in Go, Java (Kotlin?), or C#. I guess every generation of programmers must make the same errors.
I prefer not to consider other people as ignorant or stupid. They just have different motivations or reasons.
But back to the topic: I do think Cargo should have started from looking at Maven as an engineering foundation rather than NPM. Maven shipped from the start with solutions to many of the problems that Cargo is sick with because they were not considered.
Oh, I didn't want to imply people are stupid, after all, I (and many other people) made the same error 25 years ago.
But talking about Cargo (the last time I used Java there had been no Maven yet), the packages really need at least some sort of namespaces. But I don't think that (newer) languages without a really "battery" standard library can evade the problems of "too many transitive dependencies".
And don't get me started on the async story. Yikes. Rust is the language that Tokio ate.
In the spirit of being charitable, I'd like to hear more about your journey here.
At the same time, to some people the comment above might come across as snide and/or unjustified. I can't get a confident read on the author's intended tone.
I would suggest Buck2 (which is written in Rust, so you can use the existing tooling to build it) https://buck2.build/ https://github.com/facebook/buck2
Rust support is also rather good ;) https://github.com/facebookincubator/reindeer But you need "fixups" to get many cargo packages to work with Buck 2: https://github.com/facebook/buck2/tree/main/shim/third-party...
One of the specific weaknesses I found with Bazel for embedded development was support for multiple toolchains was really not great. I'm sure it's improved, but I found the gn/ninja combination much better.
However, I'm procrastinating right now by reading HN instead of fixing our multi-compiler requirement toolchain setup for Bazel, and I am contemplating resigning from my job and IT altogether. I hate Bazel at this point.
If you find a good replacement for SW eng as a career, let me know.
Could you point us to some relevant discussions (GitHub, etc) so some of us can weigh in and possibly contribute to this pain point?
But basically: vendored (git submodule in this case) project that was workspaced, added as a dep inside a workspace... Blew up badly with general workspace complaints. And doesn't help that at work we have to accommodate working with older toolchain versions because we ship a hardware product with that requirement.
[1]: https://github.com/rust-lang/cargo/issues/7520
[2]: https://github.com/rust-lang/cargo/issues/7520#issuecomment-...
>is asked to clarify
>"I'm too busy!"
But sure, flame away, you're adding a lot to the conversation /s
Could you be clearer about this? What's wrong with cargo vendor?
(It's possible it's been fixed in recent Cargo versions, but that's not available to me.)
But in the end, cargo is a tool primarily designed for building crates that depend on crates in crates.io. Even workspaces themselves are a bit of a half-baked afterthought.
But cargo really should have started by looking at Maven for inspiration -- a) a moderated repository and b) Maven has much more expressive support for version restrictions etc.
Instead crates.io feels like it's going the way of the world npm.
Anyways, sorry if this sounds like "get off my lawn", I'm grouchy today.
└── num v0.4.1
├── num-bigint v0.4.4
│ ├── num-integer v0.1.46
│ │ └── num-traits v0.2.18
│ │ [build-dependencies]
│ │ └── autocfg v1.1.0
│ └── num-traits v0.2.18 (*)
│ [build-dependencies]
│ └── autocfg v1.1.0
├── num-complex v0.4.5
│ └── num-traits v0.2.18 (*)
├── num-integer v0.1.46 (*)
├── num-iter v0.1.44
│ ├── num-integer v0.1.46 (*)
│ └── num-traits v0.2.18 (*)
│ [build-dependencies]
│ └── autocfg v1.1.0
├── num-rational v0.4.1
│ ├── num-bigint v0.4.4 (*)
│ ├── num-integer v0.1.46 (*)
│ └── num-traits v0.2.18 (*)
│ [build-dependencies]
│ └── autocfg v1.1.0
└── num-traits v0.2.18 (*)"multiple incompatible ways to do async" is another way of saying "you can use async on a tiny microcontroller without any OS at all, or in userspace on a 128-core Epyc CPU, or in the Linux kernel"
That's a unique feature of Rust compared to any of the languages with a single executor baked into the language like C# or Javascript. It does however come with costs, and one of those costs is that people who just want to write a web app are exposed to details they would much rather not have to see.
Practically speaking, however, you are right.
There is no magic trick, just people (both hired and community contributors) who invest their experience, skill and ambition to make great technology.
I disagree.
It's more like the code has ascended (https://nethackwiki.com/wiki/Ascension). It has won because it is useful enough and mature enough to accompany the language and serve everywhere.
Depending on what it does, standard library code may be used everywhere. For example, would you say that the Rust's Vec (https://doc.rust-lang.org/std/vec/struct.Vec.html) is "dead"?
It can still get new features added. It will still get bug fixes. But its primary APIs are engraved in stone, never to change because it would break all the code that depends on it.
Granted, some code that goes into standard libraries will eventually fall out of use, or will be replaced by something better. In such cases, the code should be marked deprecated and eventually removed in a future edition of the language. In that case, it is like the code retires or could even be forked into a community-supported library in case it does still need to be used.
Adding more stuff to the interface is ok in terms of things continuing to work and over time makes the library rather large. If the new interface is better there's a good chance it's different and thus inconsistent with the previous code.
Bug fixes are tricky too. If they're semantically observable people probably depend on the current behaviour.
Languages vary in how well they handle this. Python and C++ are the usual mistakes are forever example. Perhaps rust or c# have a better strategy.
Training users to act on deprecation warnings and ensuring they actually have an alternative would be a good play. Lua manages to break things a bit over time. Python tried it once and will forever fear the flame. C++ refuses to recompile anything and fears the word ABI.
It's happening in Python 3.13: https://peps.python.org/pep-0594/
Hyrum's Law (https://www.hyrumslaw.com/) definitely does come into play, but bugs do need to be fixed. Good communication is essential so that library users can adapt. In extreme cases, jarring fixes can turn into new APIs.
>> Languages vary in how well they handle this. Python and C++ are the usual mistakes are forever example. Perhaps rust or c# have a better strategy.
Rust's Editions (https://doc.rust-lang.org/edition-guide/editions/) offer a nice way to gently break backward compatibility in order to keep everything progressing.
For Rust it can't happen, because of the backwards compatibility guarantees: you can mark an API as deprecated, and potentially make those unavailable in later editions, but code that compiled on 1.0 (2015 edition) should continue to compile in perpetuity, so the API has to stay around in the stdlib so that older code still works (we dont have an stdlib per edition, it wasn't even originally envisioned to have editions affect the stdlib, only syntax and semantics). It shouldn't be a problem, but it is an inconvenient burden.
I do want to write an article on “technical debt” at some point. “Debt” is not a great metaphor for the concept. Maybe “high viscosity code” might be better? The code cannot be changed easily to accommodate new features, so it feels like it has a high derivative. Debt is something else completely.
Of course the issue is like with real debt also - how do you estimate the costs/gains from taking it and you're under risk of going bankrupt if not managed properly.
The article's differential contribution relative to other bog-standard tech debt articles is thinking about it as derivative. In the financial world, the behavior of derivatives is more complex and less predictable. Here is an excerpt showing the tech-debt-as-derivative metaphor in action:
> ... someone called you out of that debt. If we really want to stress the financial terms this is your margin call. Your users demand action to deal with your debt.
Remember: tech debt is not "literally" debt. It is a metaphor; a useful metaphor when used mindfully.
License-permitting, right? right?
Specifically, it was a piece of partial ownership on a bunch of mortgages in 2008. The issue here is that when the junk mortgages defaulted, the CDO also went sideways.
Fun fact - before 2008, they had started making CDOs out of CDOs because of of the insatiable appetite for mortgage-backed financial products. When mortgages are underwritten for people withouth a job, and packaged in a product with a AAA credit rating, it turns out it's a profitable thing to sell