I’m optimistic the Astral folks will have better success than me and I support them in their efforts. They have viable, popular solutions in hand. Hopefully that helps convert others to their cause. “If you build it they will come.”
862 karma · joined May 4, 2011
I’m optimistic the Astral folks will have better success than me and I support them in their efforts. They have viable, popular solutions in hand. Hopefully that helps convert others to their cause. “If you build it they will come.”
I don't believe there are any Apple open source references for how notarization works (at least none before it was a public App Store Connect API).
There are even times when Apple's open source releases trail functionality they are shipping in macOS. For example, Apple recently added an alternative DER encoding of entitlements, which are expressed as a plist. I don't believe Apple ever published code for how the DER encoding works. Instead, we needed to use Apple's tooling as an oracle to incrementally derive the encoding.
This Rust implementation, by contrast, has all the functionality implemented in pure Rust: there is no calling out to external processes for anything. You could drop the statically linked `rcodesign` executable into a Linux container with no other files and it would work.
That's not to discredit gon or its authors: it is a fantastic tool for streamlining common functionality. But the mechanism is completely different.
Mach-O and bundles, by contrast, require a myriad of additional data structures requiring thousands of lines of code to support. To my knowledge, nobody else has implemented signing of these far-more-complicated primitives. (Existing Mach-O signing solutions just do ad-hoc signing and/or don't handle Mach-O in the context of a bundle.)
https://gregoryszorc.com/blog/2022/01/09/bulk-analyze-linux-...
Furthermore, the index is effectively a pseudo commit without a commit message. Any workflow using the index can be implemented in terms of actual commits itself.
I think because Git doesn't have strong usability in general and especially around history rewriting, many Git users feel that the index or an index equivalent is somehow a required feature of a VCS because Git's shortcomings give that illusion. However, if you use a VCS with better history rewriting (such as Mercurial with evolve), you'll likely come around to my opinion that the index can be jettisoned without meaningful loss of functionality or productivity.
As much as I would like to support efforts like this, I feel like it is ultimately doomed to suffer from usability limitations because make does not have a static DAG. Rather, the DAG evolves dynamically as make evaluates targets. There are pesky problems like $(call) and $(eval) where you can dynamically inject make expressions into a partially evaluated DAG. And targets can generate files which are later loaded via include directives.
Dumping the make database in a machine readable format (for visualization or otherwise) would be incredibly valuable for debugging and could improve understanding. But since many make files have N>>1 "snapshots" of the internal DAG during their execution, there is a really thorny problem around when and which of these "snapshots" to use and how to stitch them together. Many make files aren't "static" enough to support exporting a single, complete, and valuable usable snapshot of their internal DAG.
If debugging is the target goal, I think a potentially better approach would be to define a function that takes an output filename and optional list of targets and writes out the point-in-time build graph when that function is called. This way makefile authors could insert probes in the make file (including at the bottom of a make file so the function is called on file load) and capture the state(s) of the build graph exactly when they need to. There could also be a specially named variable holding the names of targets that when pre- or post-evaluated would trigger the dumping of the database.
Best of luck to the person doing this work. Debugging make is notoriously hard and any progress in this area will be much welcomed by its many users.
Is it good enough or are further tweaks needed? If more, my web design skills are mediocre, so actionable feedback would be appreciated.
In PyOxidizer's case, I wanted to create virtual pipelines of actions to perform. In TOML, we could create sections to express each stage in a logical pipeline. But if you wanted to share stages between pipelines, you were out of luck. With Starlark, you can define a stage as a function and have multiple functions reference it because is "just" calling a function.
I suppose I could have defined names for stages and made this work in TOML. So let's use a slightly more complicated example.
PyOxidizer config files need to allow filtering of resources. Essentially, call fn(x) to determine whether something is relevant. In the TOML world, we had to define explicit stages that applied filtering semantics: there were config primitives dedicated to applying filtering logic. PyOxidizer's config files had to expose primitives that could perform filtering logic desired by end-users. By contrast, Starlark exposes an iterable of objects and config files can examine attributes of each object and deploy their own logic for determining whether said object is relevant. This is far more powerful, as config files can define their own filtering rules in a programming language without being constrained by what the TOML-based config syntax supports.
A common solution to this problem is to invent some kind of templating or pre-evaluation of your static config file. But I find these solutions quickly externalize a lot of complexity and are frustrating because it is often difficult to debug their evaluation.
At the point you want to do programming-like things in a config file, you might as well use a "real" programming language. Yes, it is complex in its own way. But if your target audience is programmers, I think it is an easy decision to justify.
I'm extremely happy with Starlark and PyOxidizer's configuration files are vastly more powerful than the TOML ones were.
Unfortunately, this requires building a custom binary linking jemalloc and calling into Python C APIs to configure the embedded interpreter and that's probably prohibitively too much effort for many users, who view the Python interpreter as a pre-built black box and therefore see musl's allocator as an unavoidable constant. Fortunately, tools like PyOxidizer exist to make this easier. (PyOxidizer supports using jemalloc as the allocator with a 1 line change and jemalloc can deliver even more performance than glib'c allocator.)
Not what you want to hear about an operating system :p
If you e.g. `hg clone https://hg.mozilla.org/mozilla-central` to clone the main Firefox repo, your client will connect to a CDN to download the pre-generated bundle then go back to the server to pull recent changes. Bitbucket also has the feature deployed.
After I deployed this feature on hg.mozilla.org, server-side CPU load dropped by like 90%. And with IP filtering to detect clients connecting from AWS that allows us to serve URLs direct from S3 (as opposed to going through a CDN), the total bill is like $20/mo for CI usage (S3 intra-region data transfer is free). Literally thousands of hours of CPU-core time and >500 TB/mo offloaded from the Mercurial servers.
But the GIL is more than that: the GIL also spans interpreters (that's why it's called the "global interpreter lock").
It is possible to run multiple Python interpreters in a single process (when using the embedding/C API). However, the GIL must be acquired for each interpreter to run Python code. This means that I can only effectively use a single CPU core from a single process with the GIL held (ignoring C extensions that release the GIL). This effectively forces the concurrency model to be multiple process. That makes IPC (usually serialization/deserialization) the bottleneck for many workloads.
If the GIL didn't exist, it would be possible to run multiple, independent Python interpreters in the same process. Processes would be able to fan out to multiple CPU cores. I imagine some enterprising people would then devise a way to transfer objects between interpreters (probably under very well-defined scenarios). This would allow a Python application to spawn a new Python interpreter from within Python, task it with running some CPU-expensive code, and return a result. This is how Python would likely achieve highly concurrent execution within processes. But the GIL stands in its way.
The GIL is an implementation detail, not poor language design.
While the "maybe you shouldn't use Python" comment could be construed as trolling to some, there is definite truth to your line of reasoning and I agree with comment.
I absolutely love Python as a programming language for the space it is in. But as someone who needs to think long term about maintaining large projects with lifetimes measured in potentially decades, Python has a few key weaknesses that make it really difficult for me to continue justify using it for such projects. Startup time is one. The GIL is the other large one (not being able to achieve linear speedups on CPU-bound code in 2018 with Moore's Law dead is unacceptable). General performance disadvantages can be adequately addressed with PyPy, JITs, Cython, etc. Problems scaling large code bases using a dynamic language can be mitigated with typing and better tools.
Python can be very competitive against typed systems languages. But if it fails to address its shortcomings, I think more and more people will choose Rust, Go, Java, C/C++, etc for large scale, long time horizon projects. This will [further] relegate Python to be viewed as a "toy" language by more serious developers, which is obviously not good for the Python ecosystem. So I think "maybe you shouldn't use Python for this, then" is a very accurate statement/critique.
The table mapping compression levels to window sizes is at https://github.com/facebook/zstd/blob/15a7a99653c78a57d1ccbf.... The first element is the window size. I'm pretty sure you raise 2 to that power to get the window size. As you can see, level 1 is 2^18, which is 256KB - far larger than the minimum of 1KB. Lowering it will make compression faster at the expense of compression ratio.
For comparison, zlib's max window size is 32KB. That's one reason its compression ratio is so limited.
Fortunately for us, Firefox is canonically hosted in Mercurial. So, I implemented support in Mercurial for transparently cloning from server-advertised pre-generated static files. For hg.mozilla.org, we're serving >1TB/day from a CDN. Our server CPU load has fallen off a cliff, allowing us to scale hg.mozilla.org cheaply. Additionally, consumers around the globe now clone faster and more reliably since they are using a global CDN instead of hitting servers on the USA west coast!
If you have Mercurial 3.7 installed, `hg clone https://hg.mozilla.org/mozilla-central` will automatically clone from a CDN and our servers will incur maybe 5s of CPU time to service that clone. Before, they were taking minutes of CPU time to repackage server data in an optimal format for the client (very similar to the repack operation that Git servers perform).
More technical details and instructions on deploying this are documented in Mercurial itself: https://selenic.com/repo/hg/file/9974b8236cac/hgext/clonebun.... You can see a list of Mozilla's advertised bundles at https://hg.cdn.mozilla.net/ and what a manifest looks like on the server at https://hg.mozilla.org/mozilla-central?cmd=clonebundles.
A number of months ago I saw talk on the Git mailing list about implementing a similar feature (which would likely save GitHub in this scenario). But I don't believe it has manifested into patches. Hopefully GitHub (or any large Git hosting provider) realizes the benefits of this feature and implements it.
server.validate ensures that all referenced revisions from changegroups are in fact present. It prevents repos from becoming "corrupt" (in the sense that `hg verify` will complain) due to missing data.
Transitioning away after 10+ years is a sign of several positive things:
* Creating a successful project that lasts 10+ years
* Having the perseverance, patience, and willpower to guide that project for its lifetime
* Creating a healthy community whom you are able to transition day-to-day responsibility of that project to
Matt has accomplished something that most of us never will. I'm envious of what he has accomplished and that he is able to walk away from a healthy and successful project. Truly a remarkable accomplishment.
But even then, as you mention, Git doesn't scale well for them. This is why Facebook is using Mercurial. This is why the Chromium and Android projects are spread out over hundreds of Git repos. This is why Google's internal monorepo is using a Perforce-like homegrown tool backed by Google's distributed wizardry.
Git will start to see significant scaling problems beyond say 1 GB total data, 100,000 files, and X number of refs (there has been significant upstream work around scaling refs - a bunch from Twitter I believe - and I haven't measured how well newer versions handle thousands of refs).
The good news is that unless you are a software project with hundreds of developers, have large binary assets, or are doing something like Twitter and have all your developers push all their in-progress feature branches to the monorepo (madness if you ask me), a Git monorepo will scale for you. A Mercurial monorepo will scale even better. Keep in mind that most people aren't {Twitter, Facebook, Google, Mozilla}, so chances are you won't ever reach these scaling limits.
- How does the build system integrate multiple, discrete repositories into a unified system?
- What are the dependencies between the repositories?
- How are the sub-repositories laid out on disk? Do the separate repositories use separate toolchains?
- How do I decide when to update the reference to a sub-repository? Are they completely separate? Versioned as one logical entity?
- Do separate repositories reference elements in each other? Can I copy files between repositories or should certain files live in certain repositories?
These and other similar questions exist when you use multiple repositories.
(I hope you see that multiple, discrete repositories aren't a panacea and there is a counterpoint to each of your points.)
I posit that if Git or Mercurial allowed you to clone a subset of directories, the differences between a monolithic repository and a set of smaller repositories becomes indistinguishable, as a clone of a sub-directory is functionally equivalent to a standalone repository! The problem is that narrow clone is not implemented in any popular DVCS tool today (but Mercurial is working on it).
It's part of my job role at Mozilla to bring this process into modernity. We're getting there. Slowly. GitHub pull requests will be supported in some form. Hopefully by July.