Rust required to build Gecko
groups.google.com
groups.google.com
For anybody who missed this, Project Quantum is a Mozilla project to dramatically improve Gecko. Part of this project is to bring in Servo components like CSS and WebRender, hence the Rust dependency.
More awesome info:
https://wiki.mozilla.org/Quantum
https://medium.com/mozilla-tech/a-quantum-leap-for-the-web-a...
https://billmccloskey.wordpress.com/2016/10/27/mozillas-quan...
Especially with browsers, which not everyone agrees on how they should be and desires to customise, only to find that the option to do so has been removed or a source change must be performed, is subsequently delighted to know that it's open-source so they should be able to do it easily, but then get overwhelmed and give up after they realise the effort needed just to build an unmodified version of the software themselves. They then fall back to merely complaining on the Internet, and reluctantly accepting their "fate"... somehow, I feel like some of the visions of open-source didn't quite turn out as well as hoped.
$ git clone git://anongit.freedesktop.org/libreoffice/core
$ apt-get build-dep libreoffice
$ ./autogen.sh && make
With any relatively standard Debian/Ubuntu distro, you can have your own custom build in about an hour or so. Progressive builds just a few min if you have ccache setup.
edit: I noticed you typed make. Do you perhaps mean make -j 8 or something (depending on core count)
That being said complicatedness has little to do with compile time now, has it? :)
There are some articles of Joel Spolsky explaning why office document formats are so complicated (not my word) [1]. The spec for MS doc is like 500 or so pages. The spec is longer than my video game graphics book I had in college. I recommend some research before assuming.
> That being said complicatedness has little to do with compile time now, has it? :)
Actually there is a correlation [2]. Of course it depends on what you call or think complexity is.
Believe it or not engines may not actually be as complex because they typically are based on math formulas or combinations of simple algorithms. Math is often simple and doesn't require lots of branching (if conditions). The algorithms have to be fairly simple anyway otherwise performance would be bad.
Buisiness software or actual video games on the other hand are full of branching. NetHack may not have 3D graphics or any graphics but it is an extremely complicated game.
[1]: https://www.joelonsoftware.com/2008/02/19/why-are-the-micros...
Here's what I can think of off the bat:
- Handling fonts. A whole world of hurt. Not all of this can be offloaded to OS libraries.
- Handling layout. Another world of hurt. Especially given that many formats have their own quirks.
- Handling complicated things like math and svg.
- Understanding HTML (for formatted paste and save-as-html to work)
- Understanding PDF
- Canvas for the shapes and stuff
- Codecs and stuff for images and video
- General rendering/compositing code
- Accessibility code
There's a LOT that goes into building software like this.
I worked on Microsoft Office as an intern once. The codebase is huge. Build times would be forever without artifact builds. I currently work on browsers, which are in the same situation with a lot of not-immediately-obvious costs.
Layout is in every modern game engine as UI and is even more powerful. See scaleform for example.
Math... isn't in a game engine?
HTML/PDF and other formats? Please how about DXT1-N texture formats, crunched textures, scene descriptions, meshes, animation rigging, etc. HTML and PDF are small in comparison.
Yea I'm not going to continue. What you may not have realized as a MSFT Office intern is that just because it was slow and awful, it might not have necessarily needed to be.
It's not just Office. It's also the two browsers I have worked on (a word processor in particular has a lot in common with a web browser, but it also needs to deal with authoring). There are many things that need to be done that aren't immediately obvious. Sure, a graphics engine has such costs too but it's hard to compare unless you're well acquainted with both sides.
It's basically one line to add to a config file to enable it.
I actually had a proposal for a tool that would let you hack directly on your Firefox install without needing to download the full source (which is hard to do on flaky connections and annoying on slow ones). This is already possible, Firefox ships with a zip file (omni.ja) containing its UI code. The UI code is precompiled and cached, however you can edit the code and delete its cached form and stuff just works.
It's a hack, but it makes it incredibly easy to just get started hacking on Firefox code. I could imagine there being an addon that helped smooth this process.
For more info about contributing to Firefox, see https://developer.mozilla.org/en-US/docs/Mozilla/Developer_g...
That said, what you describe isn't really an issue with Firefox. Unless you're running a very odd distro, apt-get build-dep or mach bootstrap will set you up automatically, and on Windows there's an installer with the dependencies (outside MSVC). Of coure, the people who like to customize to no end might be those running an odd distro, but something something digging yourself out of your own hole.
You could write your comment without accusing tweakers of digging themselves out of their own hole and all that unnecessary tone. In that case it would have more power to teach people something new.
In fact I'm more scared by people trying to silence him/her.
I certainly didn't want to silence the parent, but tried to point out that having a much more helpful tone will reach his message much farther.
Maybe reread this comment thread later, and compare the tone of the comments.
But that's the problem with context and text. Emoji FTW.
But they may have meant Midori or similar which actually uses WebKit, through WebKitGTK+... Which means linking against GTK+, probably GTK+3 to get the main benefits like single-page not whole application crashing.
Which depending on your custom build, may or may not be an absolutely huge overhead.
It could be a tiny change to add to an OS. Or huge.
Oh come on
I imagine an image with the full toolchain installed, geared towards using it as a kind of "firefox-compiler" CLI (so coupled with a few shell scripts), that only requires you to know how to install Docker|Vagrant. This could easily be bundled with the source code.
If people want to install all the build dependencies themselves, that's fine, but there's really no reason to make people go through that hassle. Automating and isolating it also makes it far more liely that the dependency list remains accurate.
FROM suitable-base-image
ADD ./install.sh /tmp
RUN chmod a+rx /tmp/install.sh && /tmp/install.sh
and write a script to install the dependencies that people can opt to run outside of Docker if they prefer.The point is not to take away peoples ability to pull together all of the dependencies on their own, but to make replicating the build environment trivial.
Personally I don't want install scripts etc. polluting my server or laptop setups - I always spin up containers to run builds in these days, because it means I at the end have a repeatable description of how to set it up. But once someone has done that job once, it's silly for everyone else to have to repeat it.
Docker maybe makes sense as a "compilation target" - generating a dockerfile of your dependencies to make it easy for people to use makes sense. But it's not rich or structured enough to be the canonical record of what dependencies your project needs.
I'm not arguing for it to be that, and I largly agree with you.
As my example shows, the Dockerfile doesn't need to contain any detail of the dependencies - you can "outsource" that entirely to whatever mechanism you wish, as long as you can easily add whatever you need to apply those dependencies to a suitable base image to the Dockerfile.
But most of the time I don't care about precision - I just want a means of getting a working build environment quickly. If it pulls in 100MB of unnecessary libraries, I don't care, if it gets me an uncomplicated way of building the project without having to assemble the dependencies myself.
For a lot of projects I come across, my experience is that because the list of build dependencies are often never automatically tested, it's a crapshoot whether or not their build instructions and dependency list will yield a working build environment on any given system.
"Even" a plain Dockerfile that is regularly used to rebuild the build environment is far superior to a text file of dependencies that is rarely to never verified.
If you want to go one better and storing more precise instructions and then use them to generate a build script or Dockerfile, then awesome.
The main thing is that absent automated testing of the build environment too, the list of dependencies isn't worth the bytes they're stored in. And a simple way of regularly testing the build environment is to rebuild it automatically.
My experience is that the dependency lists for projects that automatically rebuild the build environment - whether as Docker containers or in a chroot or any number of other ways - tends to be far more reliable.
And if you first can automatically build the build environment, it makes sense to make it easy for everyone to use that automatically built build environment.
It only includes the build tools though. You need the source code separately. I think I have provided enough info in the README to get you going.
Have people traditionally been running their own modified browsers? Seems a bit scary given how you'd have to be on the ball to merge security fixes into your branch regularly.
Of course, that feed of patches will stop at some point. Don't ask me what the contingency plan for those forks is...
The release engineering team has started using Nix in their build process. Once all dependencies are mapped in Nix it means you just need to install the tool itself and all the dependencies will get pulled automatically.
It doesn't solve porting to new arch and platforms (especially Windows) so it's still a good idea to keep the number of dependencies down. But on any amd64 Linux it should be fine.
Use open source projects' CI file to find which dependencies / commands are needed to compile / run the project. Run all this inside whichever container the CI system uses (circle & travis use only a few).
The issue described will never be an issue again!
In this case it looks like you can contribute to one of the servo project packages and those changes will make it downstream into Firefox, that's great as Servo has a (relatively) small and easily understood code base!
I believe it'd be mutual - Firefox would have a bright future because of Rust...and vice versa.
We definitely want to replace bits of Gecko with Servo code (in general, Rust code). That is already happening.
Servo development is still continuing. Servo and browser.html are evolving in the direction of eventually becoming a product. We don't know right now if they will. In the meantime, Gecko can reap the benefits.
That said, I think Servo has a bright future ahead as well, even before the long term. For example, matching the functionality of embedded rendering engines for hybrid apps is far more likely, and Servo has its speed as a major advantage over the competition. And who knows, it might make it into Firefox for Android :)
Back then, it didn't take long until Mozilla was stable and usable enough that I could use it as my main daily browser, and I expect the same to also happen with Servo. Of course, Firefox is far ahead in functionality, but Servo doesn't need all of Firefox's features to be successful.
The problem is that the modern web is complicated. You have a lot of features like svg which don't get used pervasively but are enough for it to impact experience.
As for Servo crashing, we don't really prioritize crash fixes since it's not a product at the moment. But please do file bugs for it if you think it's not a known crash.
(I was going to the Wikipedia article on "Animated GIF" because I wanted to see how well Servo worked with animated images, and I knew I'd find one there. Sadly, it panics every time.)
But you're right, we don't support animated gifs.
And there was also that Mozilla bug where IIRC some kind of rounding error caused 1-pixel misalignments with floats. Took a long time until that one was fixed.
There's no fundamental reason that the Rust compiler has to be dependent on LLVM. It's just a good strategic decision as it allows the Rust developers to focus on the parts of the compiler that are unique to Rust, and use a well tested backend with lots of existing optimizations and targets for the parts that aren't particularly unique to Rust. There is actually some discussion already of using a faster, pure-Rust backend for debug builds, and maybe far in the future for release builds as well (https://internals.rust-lang.org/t/possible-alternative-compi...).
Or, maybe all of the environment variables I used when I worked on Firefox just happen to be the same, but the build system is totally different.
I don't think it's a big deal that the possibly-sudo-requiring step is kept separate and not automatically invoked on build.
You get reproducible builds because the build happens in its own sandbox, ideally. But it's not absolute.
There's a continuum here: it's not awesome to have builds from source for every library if you know you're building inside a defined Docker image, for example.
So yeah, would be nice to have, but I'd be wary about making it the default.
I find it fascinating that lessons aren't being learned from build systems which are built in the enterprise like Bazel, Pants, etc (honestly, I'm disappointed one of those two wasn't just adopted by the Rust community - unclear but I think they didn't know about them?).
Similarly, notice how "Java shops" have for many years enforced ever increasingly complex build pipelines (Javadoc, Junit, JaCoCo, Findbugs, ...). But these tools are slow, end up with long standing bugs when versions bump, and generally end up limited.
I find it fascinating however, that only two of the above steps (Docs, Tests) are first class in Rust. Why only those two? Why not learn from the whole Java build pipeline? Why does the compiler's test infra not spit out metadata about lines/branches covered? Why is clippy, like Findbugs, separate from the compiler? Isn't "better code" by definition better for everyone.
It should be relatively easy for these to be built into the compiler. While some like test coverage data is really hard to get as separate tool (especially when optimizations are on). However, Rust developers have valid points like "separation of concerns". Its hard to argue with them, after all we are only on the sidelines, they are actually in the thick of it. But again, there just seems to be a divide, and its interesting to see which features/lessons are adopted, vs which features/lessons are "not worth the trouble".
FWIW (my stab at a reason for this divide), I think its largely a problem of peers. In the enterprise, you work with great developers and you work with not so great developers. Most of these are fresh college/code-bootcamp developers but there are some rare bad experienced developers. At least in the field of building languages (maybe open source in general) you get to be very selective. Essentially, you only get peers who are experts (even if some of them do start off making changes while in college or without any formal education - also keep in mind the selection bias of people curious and motivated to help a project they aren't being paid for). Compounded with few strict deadlines, this basically means everyone is on their A-game all the time. eg. "Why would you need metrics about how many lines+branches of code are tested? and you want to potentially fail the build if it doesn't meet some threshold?? Some things are hard to test! Shouldn't the implementor get to decide how much testing is the appropriate amount of testing?"
P.S. Please, anyone, correct me where I'm wrong. I too am on only one side of this divide.
So instead of choosing a large feature set at the beginning and the box it through, they choose to have a more thin compiler, standard library etc. but make it extendable, so that it can be extended and integrate with other systems. Also you should not forget that Rust is still very young for a programming language so cargo is everything but done. Cargo is more a unified interface combining the compiler and the crates package repository than a full fledged build system. Given the available resources and time for building rust, I thing choosing to (for now) not include a fist class rust coverage system, debugger etc. is a sane choice, especially because all this things are available as external tools (debugger=gdb, coverage=kcov, etc.).
Also there are discussions/issues about enabling the integration with "enterprise" build systems, they just currently don't have a high priority and happen mostly in the background. So it's probably just a matter of time until you can use rust with a "enterprise" build system like Bazel.
Wrt. clippy integration in rust, it is mostly about keeping the compiler smaller and easier to maintain. Through it should be noted, that all "important" lints are in rustc, clippy just provides more useful ones. By keeping both projects separated it is possible to iterate both of them independent of the other. Which can be quite useful in open source development. Also don't enterprise systems also have a separate linter? (And no "better code" is not better code for every one, as the definition for what good code is can change quite a bit depending on the person, e.g. wrt. variable shadowing)
Lastly I don't think your reasoning about peers holds, you don't really have the choice to be selective about people working on the project, at last not if you have open source project on the scale of rust. Through you might be able to have more strict requirements for code quality. Nevertheless given how rustc catches a lot of possible bugs, and how to community is, I would be surprised if there would not be some less experienced programmers contributing to the compiler. But then at last Mozilla also uses rust in production, so some parts have deadlines and even if not it's quite different for them then a `A-game`classification. Also it's not so, that you can't have metrics about lines/branch coverage, failed the build on certain thresholds, etc. It's just not necessary a build-in feature, through available nevertheless. Also it would be really strange if a programming language / build tool would decide which amount of coverage is ok, it's something the project manager/programmer has to decide (once) and then configure the build tool to enforce it (I most likely did misunderstood you on the last line ;-) ).
Ups, I wrote much to much.
TL;TR: 1) rust has less (programmer) resources 2) it's still young, tools like cargo, or the testing API are everything but complete. Through they do there job good enough for now.
This is expressed pretty nicely, thank you. I was struggling to express this in my other comment. rustc has all the lints which:
- Everyone mostly agrees upon
- Have rare false positives
- Are of the kind you'd expect from a good compiler
It doesn't have 150 lints because 150 lints would be annoying, and would drown out the "important" lints. I work on clippy, and I find this distinction very useful. When rustc tells you to fix something, you fix it. Clippy lints help /inform/, and you may ignore them often. But you don't always need to listen to clippy. If I wrote a bot that fixed all rustc warnings I'm sure people would be fine with just blindly running it. I cannot say the same about clippy (not because of bugs), and that's a good thing.
The plan instead is to polish clippy and distribute it via rustup.
Yes, you can add new lints with an RfC. However, the bar for inclusion is very high.
I'm totally okay with the status quo, but I very much disagree that there is scope for uplifting lints to rustc from clippy with the current policies.
On one hand, I see how it is annoying when a compiler complains about code patterns that are not worth errors and could be false positives. OTOH, I feel like a language designed for safety should try to warn about as many bugs as possible. Would any of clippy's lints pass every Rust crate tested on Crater? Maybe those could be come rustc errors.
For fun:
Lint, a C Program Checker (1977)
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.56.1841
The separation of function between lint and the C compilers has both historical and practical rationale. The compilers turn C programs into executable files rapidly and efficiently. This is possible in part because the compilers do not do sophisticated type checking, especially between separately compiled programs. Lint takes a more global, leisurely view of the program, looking much more carefully at the compatibilities.lolnope
Well, there are lint-checking-lints, and there are very few linters out there, and clippy passes clippy, so those would pass. A couple other obscure lints that check for things that in practice never happen. Too obscure to be included in rustc I guess.
IMO a lint passing all of crater is a very good reason for it to not be included in rustc.
What specifically makes you think lessons aren't being learned?
> honestly, I'm disappointed one of those two wasn't just adopted by the Rust community - unclear but I think they didn't know about them?
Bazel had its first release in March 2015, and wasn't even considered beta ready until later that year. Pants had its 0.0.17 release, the very first public one, in July of 2014.
Cargo was announced in March 2014, earlier than both.
> Why only those two?
Well, those two are necessary. Other tooling is nice, but not the bare minimum needed. We need to stay focused, but more on that later.
> Why not learn from the whole Java build pipeline?
I have not dealt with the Java-specific tools you're talking about in a very long time, but pipelines have advantages too. A few slow tools doesn't mean the whole concept is bad.
> Why is clippy, like Findbugs, separate from the compiler?
Rust cares a _lot_ about stability. Adding new lints is extremely difficult, because they can break people's builds. Clippy being a separate tool means that they can iterate on lints at a different release cadence than the compiler itself, and aren't bound by our mega strict breaking changes policies.
> In the enterprise, you work with great developers and you work with not so great developers.
I assure you, it is very much the same in open source. ;)
So, we do very much care about Rust's usage in enterprise scenarios. For example, Mozilla in many ways feels exactly like an enterprise customer of ours. It was a huge amount of technical (and social) work to get to this point, and a lot of that involved negotiating how the builds would actually work. Our perspective is that Cargo is a world-class tool for building code, but that's only possible because it's focused on _Rust code_. So the strategy should be, how can we best integrate Cargo with these more generic build systems? We've already taken a number of steps forward on this front, and there's more to come.
Most of these features lacking in Cargo are due to a lack of some kind of champion to lay out the specific needs; we don't want to just build features and hope they're useful. In the past, we've had several people come to us with "here's my issue getting Rust in my workplace", and that experience has been valuable. We're a small crew, we need to focus in order to keep shipping quality stuff. We don't have the ability to just toss some people on it, for whatever value of it, and get it done.
> Cargo was announced in March 2014, earlier than both.
Thats fair, and a part of what I meant about "they didn't know about them", because working for Amazon right out of college, I've been using things like Bazel/Docker/Travis for 5+ years now.
I actually meant it as a discredit to people who work in the industry for not exposing these tools earlier. I did a poor job of communicating that (it was after all like 3 am :P).
Furthermore, Bazel and Cargo are aimed at slightly different use-cases. Cargo is tailored towards Rust, is integrated with a Rust package repository, and can use it's knowledge of Rust to allow for builds of pure-Rust projects with a lot less configuration. Bazel is a more general purpose build system, but that means it requires a little more configuration. There's no reason you couldn't use Bazel for building a larger project that includes Rust and C, C++, Java, or other languages that need to be built, but Cargo makes it a lot easier and simpler to handle pure-Rust projects than it would be in Bazel.
Many of the other tools that you mention don't require just more support from the build tool, but also significant engineering effort on the part of the compiler, or a separate static analyzer, or other development process tools. These tools are desired, there just isn't anyone who has had the time to design and build them.
As to why Clippy is separate from the compiler, there are a couple of reasons. Adding more lints to the compiler, especially ones that are turned on by default, can cause problems as well. Any time you add a new lint, you may break people's code that has #![deny(warnings)] turned on. Even if people don't, a lot of people will still try to rewrite code to make it comply with lints; but that in itself can cause problems. I've seen many bugs introduced over the years by people trying to do too many lint cleanups at once, and making mistakes in some of their cleanups. For this reason, the Rust compiler is fairly conservative about adding new lints..
That said, it sounds like there are plans to integrate Clippy with the compiler and cargo, but probably as an opt-in separate command rather than being included in the default set of warnings: https://www.reddit.com/r/rust/comments/5ibr2a/cargo_check_ha...
Code coverage is another example. You can use some existing code-coverage tools with Rust (https://users.rust-lang.org/t/tutorial-how-to-collect-test-c...). Because there are tools out there that work, there's less of a pressing need to get code coverage integrated with the Rust toolchain; I think that it would be a good idea to do eventually, as it will make it easier to do cross-platform and out of the box, but for now it's not the highest priority.
So, I don't think that any of these omissions are due to a difference of philosophy; just a difference of focus. Rust is all about admitting that programmers aren't perfect, and that better tooling can help avoid mistakes and thus make developers more productive. Right now, that focus is on things like improving the compiler (adding MIR, which allows for fixing several bugs in the compiler and doing some Rust-specific optimizations before getting to LLVM), implementing the Rust Language Server which can be used as a backend for IDEs, improving the ability of the compiler to do incremental and parallel builds, and adding language features with a focus on making Rust easier to use.
They didn't exist at the time. cargo was built using lessons learned from other package managers, however.
> Why is clippy, like Findbugs, separate from the compiler?
The compiler didn't want to bloat its lints. Being out of tree let us iterate quickly. It lets us work on extremely useful lints which as a side effect have false positives (there are a lot of patterns it catches which might be buggy but sometimes have legit use cases that are hard for clippy to detect and ignore). It lets us add controversial lints. In general a separation between the "core" lints (which everyone usually agrees on passing, and cover a small set of bare-minimum issues) and the clippy lints (which not everyone agrees on, and cover a wide range of issues which don't always apply to your codebase) is nice to have.
There are many reasons to not want to run clippy, one of them being that's it's too strict/annoying. In Servo we can't run clippy because someone needs to go through it and update Servo over the thousands of warnings clippy produces when run on Servo (many of which are false positives). I occasionally do some of this, but clippy grows pretty quickly too so each time I do it there are new things. (I'm waiting for rustfix to mature so that I can automate a lot of this).
So, to me, clippy should be something you have to decide to use.
We do plan to make clippy part of the rust distribution (bundled with cargo and friends), it's just not happened yet since it's blocked on a bunch of things.
Clippy not being part of the default pipeline is just an artifact of its immaturity. It's still used by a lot of people despite not getting free publicity from being a default which is promising, though.
I don't see why there's much value being assigned to "first class" tooling in Rust. code coverage and clippy are a `cargo install` away.
> Most of these are fresh college/code-bootcamp developers but there are some rare bad experienced developers
There are a LOT of new/inexperienced programmers in the Rust community. You're right that there's still some selection bias, but Rust tends to be pretty welcoming. I've mentored people who've only done very basic programming in making Servo pull requests. Some of the sporadic clippy contributors are pretty new to programming.
> Why would you need metrics about how many lines+branches of code are tested?
I have never seen opposition to code coverage in Rust in that form. In general folks in Rust are happy to have more kinds of checks lying around.
It's not part of the pipeline mostly because it's a cargo install away. If your project needs it, it's pretty easy to make it part of your workflow.
There is an argument to be made that making it part of the pipeline will mean that more people will be driven to use it, which is great. You'd have to take it up with the tools team and community if you think this is something that needs to happen; I am mostly ambivalent about the idea.
I was trying to discuss the divide (maybe come up with reasons for it - potentially improve it), not the lack or need for tools. I was just providing concrete examples using Java vs Rust build pipelines. It might have been better if I had stuck to abstract examples like: "I find it fascinating when something like X is so obviously a priority to me(us), but it takes a lot of effort to convince others that it is a priority at all."
I hope that is more clear.
What's the advantage over a build system that downloads the dependencies, but which gives you the ability to prefetch all the dependencies (so you can reliably do work without connectivity after prefetching)?
On one hand frameworks and code in general becomes more and more bloated and incomprehensible these days, due to an overzealous approach in the aforementioned. On the other hand tools are getting cramped with features which were never in their original scope.
The leftpad issue was only one example of that lets-inject-third-party-code-on-the-fly mentality.
This separation allows to build the project off-line, which is not an uncommon scenario. There are environments that have the direct internet access prohibited e.g. by a company policy. There are environments that have it difficult (e.g. are behind a proxy). There are distribution package builders (RPM, DEB), which have a policy of only working on local sources.
And then there is build reproducibility. If external network is involved, the whole reproducibility idea goes out of the window. Remember the left-pad farce? Part of the cause was idiotic split to microdependencies, but part was that everybody used external network for their build process.
Fair enough, but if you have outdated dependency artifacts, it doesn't make sense to compile without fetching them, and so I think it's reasonable to make the 2nd step depend on the 1st one (and thus make it execute automatically)
I understand that keeping the 2 of them too-close might inadvertently conflate the 2 concerns (the build step becomes impossible to run without the download step) even if the design goals explictly thought of them as being able to be run independently, so this might be an argument for "builds should only build", but I'm not sold on it yet
> And then there is build reproducibility. If external network is involved, the whole reproducibility idea goes out of the window. Remember the left-pad farce? Part of the cause was idiotic split to microdependencies, but part was that everybody used external network for their build process.
Well, the build systems that I have in mind are Stack ( http://haskellstack.org/ ) and Nix ( http://nixos.org/nix/ )
Both of them have an heavy emphasis on reproducibility, and yet both of them automatically download dependencies.
(the Hydra Nix continuous integration systems OTOH run build and test steps with limited/no network connectivity, to enforce that separation of concerns)
(the microdependencies btw wasn't a problem that caused build non-reproducibility, microdependencies only made it worse when the shit hit the fan)
Also, there are 2 different reproducible build failures: it doesn't build (bad) or it builds but the artifact is not identical (worse)... downloading dependencies from the internet can only cause problems of the first type, if you have a sane way of fetching them
sane means: either you have the hashes of the dependencies to check, or you can trust the archive to always give you the same files when you ask it for a snapshot... and for security reasons you'd still check the crypto signature on the hashes to verify that nothing has been tampered in the mirror from which you downloaded
(Unfortunately, not only Npm but also Hackage afaik have that anti-feature that allows you to reupload a "bugfix" with the same version number... so we cannot always trust that the systems that we're using are sane, which OTOH is impossible also due to DMCA or other law/court shenanigans: if your dependency violates someone's copyright, it might end up disappearing from the snapshots)
I am sorry, I'd contest that. There are often cases where one might not want to use the most current version of some library (incompatibility, undesired features, etc.).
This entire forced update approach is already unpleasant enough in the context of applications (most notably on Android and IOS) but becomes unbearable for software development. While one might argue it could still make some sense for the former, as average users might stay forever with old versions and potential security issues, that argument should not be valid in a "professional" context - where the "users" know about the implications - as it should be given with software engineering.
I think we're misunderstanding each other.
What I meant is that if you have foo-1.2 in your dependencies.conf, and you have previously downloaded foo-1.1 it doesn't make sense to compile, because the result of the compilation will most likely not be what you want
You can just have the build say "foo-1.2 not found. Latest version is foo-1.1", and the developer can decide how to respond.
- "usual" case (if you agree): dependencies will be downloaded from the Internet, single command download+build is useful
- in a bank/SC environment: the developer configured a proxy/local mirror of the package repositories, and the dependencies will be downloaded automatically, single command download+build is useful
- in a bank/SC environment: manual process to obtain and add locally the dependencies, the build/dev system has limited networking: automatically download fails, single command download+build is not useful but not harmful either
Having a "prefetch" and "offline-build" commands are perfectly fine, but I don't see the reason why the default shouldn't be a "build" command that does prefetch+offline-build
Another part was using location-addressable storage for dependencies, rather than content-addressable. If dependencies used something like magnet links, IPFS, etc. then nobody would really care if some particular machine stopped serving some particular file.
Yes they would. Because such systems can only serve files that exist on at least one node, people who want to rely on such a system to serve a given file must run their own node to host every file they care about the availability of.
The topological properties of distribution are different, but the existence of the local cache to provide dependability is not avoided.
mvn dependency:copy-dependencies
And when you build a project once, all downloaded dependencies are cached locally indefinitely. That is, you can furnish a Maven build with offline dependencies.Normally you do want to simply have your build download (and cache locally) a new version of a dependency, but an internet connection is not required if you have the dependencies available locally. Maven will always look in your local repository first, and will not try to download anything, except if you are depending on a SNAPSHOT version of something (which implies that the build should always check for a newer version).
Not really. Building from local sources only or "living off the grid" as parent poster put it typically isn't just about temporary loss of Internet connectivity, but instead serves as a shorthand for a more complicated set of requirements.
Generally speaking there will be (at least) disaster recovery requirements such that a computer with a fresh OS install and access to the local very-well-backed-up archives can produce a build. And the local archives typically have exactly one entry point so that legal and process controls can be applied as they go in.
So the problem isn't how to cache dependencies for a working directory, but how to get, archive them, and distribute them to a development team easily. A lot of package mangers approach supporting this kind of use case, but still need a fair bit of customization and manual steps if not new tooling built around them. And all that infrastructure also has to meet qualification criteria (through testing).
It gets to be enough that any tool that decides to add in their own half-assed download manager will be quite annoying to anyone that doesn't use "YOLO" as their quality management system.
I think that this is an issue of properly covering a use-case, and having simple configuration options to avoid manual and cumbersome steps... But I don't think this means that a build system shouldn't default to downloading its deps.
Basically, if "build systems should only build" means
- build systems would be better off with an offline build mode available
I agree 100%... Otherwise, if it's
- build systems shouldn't have an option to automatically download the dependencies and then build
I'm still in disagreement
Of course they get intermixed in practice, package management provides build prerequisites, and completely separate tools are not necessarily conductive to getting information from the build system (or dependency list in the tarball) to the package manager. Or for the download/installation provided by the package manager to be at the perfect scope for the software build.
I, unlike grandparent poster, am not making statements of "ought" for build systems. Just providing context.
However, I can still generically say, the tightness of the build and package management limits user flexibility. I'ts not even a particularly interesting phenomenon. It's just another case of software coupling. There are benefits to high coupling and benefits to low coupling. Typically tight coupling means the normal use case works better at the cost of all abnormal use cases. It's the same-old-same-old of all software design.
----
As a side note, I will say that in my experience an inflexible or awkward to use "offline" build mode is not an offline build mode. If "offline build mode" means:
-The tarball contains (or build system generates) a list of files that can just as easily be fetched with `wget -i deps.txt` and dropped into a subdirectory manually.
That's cool. If it's:
-I have to reverse engineer a one-off package manager architecture, write a script to fetch build-deps externally and hack the build to support my particular flavor of offline build because it's not the type of offline build envisioned by the tool author
Then that's bunk.
I'm not sure they can, or at least not without creating a circular dependency. It is greatly advantageous for build steps to be first-class packages that can use the library ecosystem etc.