Starting October 19, storage limit will be enforced on all Gitlab Free accounts
docs.gitlab.com
docs.gitlab.com
Using hosted GitLab for open source projects is looking less and less appealing.
I also posted about issue trackers on gitlab.com not allowing search without signing in a while back: https://news.ycombinator.com/item?id=32252501
Edit: An open source program that upgrades the quota is mentioned elsewhere in the thread: https://about.gitlab.com/solutions/open-source/ I don’t use hosted GitLab for my open source work, so no idea how many people get approved.
Even though the code is available, that does not mean you can use it however you want, no?
Public repos have a cost factor of 0.008 so your actualy CI quota is 125,000, instead of only 400.
It would actually make it feasible to grow on GitLab if quotas changed based on if its a public or a private repo, like how GitHub does it.
https://docs.gitlab.com/ee/ci/pipelines/cicd_minutes.html#co...
For comparison, I think GitHub just have a cap of 100MB on any single individual file, plus:
> We recommend repositories remain small, ideally less than 1 GB, and less than 5 GB is strongly recommended. Smaller repositories are faster to clone and easier to work with and maintain. If your repository excessively impacts our infrastructure, you might receive an email from GitHub Support asking you to take corrective action. We try to be flexible, especially with large projects that have many collaborators, and will work with you to find a resolution whenever possible.
Which is a bit wishy-washy, but sounds like there's room for discretion / exceptions to be made there rather than a hard cap at 5GB.
And made private repos free after the acquisition https://github.blog/2019-01-07-new-year-new-github/
Don't forget easy access to all the code to train Copilot, the AI code launderer. With my tinfoil hat on: even the private repo edition!
The source you link is 2 years old.
Yes, the information is 2 years old, but since nobody has any proof of the contrary, I'd say it still stands.
(Sorry, my upload is too slow to just do a quick check and see if I can indeed push a 100gb repo, maybe someone else can try?)
There's a 100MB limit on the size of a push, which also limits the size of a any single git object (i.e. file) to 100MB too. However GitHub supports LFS for large files, and their documentation says to use LFS for files over 100MB:
https://docs.github.com/en/repositories/working-with-files/m...
> GitHub blocks pushes that exceed 100 MB.
> To track files beyond this limit, you must use Git Large File Storage (Git LFS).
https://docs.github.com/en/billing/managing-billing-for-git-...
According to my own GitHub account ( https://github.com/account/billing/data/upgrade?packs=1 ), I'm paying $37/yr for 600GB LFS storage on top of my existing GitHub Pro subscription.
I'd love to know how you managed this! I'm paying $60/yr for 50GB, which feels exorbitant.
(I just ran that on the QEMU project and reduced the usage from 295GB to 165GB by deleting pipelines older than 1 Jan... so that's a lot of low-hanging logfile fruit gitlab could be auto-deleting.)
Build Artifacts are listed as part of what contributes to the quota (which is fair enough), but there's no way (that I could find in the docs) to manage build artifacts that are stored.
I suspect we have a large historic storage, which we don't use / need, but there's no way to browse this, no way to verify it, and no way to delete what we don't need.
I'm dreading getting a big ole bill in a few months for storage we had no way to opt out of.
Found this recipe on StackOverflow to just nuke everything and worked for me: https://stackoverflow.com/questions/71513286/clean-up-histor.... You can tweak it a bit if you want to keep the more recent pipeline runs.
For some reason I had to run this multiple times to completely remove everything; each run removes about half of all the pipelines and only when there were about 20 remaining did the script remove everything.
If you’d like to investigate now you can use the GitLab API to fetch job artifacts and their storage size: https://docs.gitlab.com/ee/api/jobs.html#list-project-jobs and calculate the summary.
I’d suggest collecting all pipeline jobs to be deleted using this API endpoint https://docs.gitlab.com/ee/api/job_artifacts.html#delete-job...
We are looking to improve the visibility for job artifacts in the UI as we get nearer to the storage limit being enforced.
I mean if I'm self hosting on digitalocean doesn't that just mean that I'm using your FOSS and doing all the work myself and completely separate from gitlab anyway other than the common codebase?
Sorry, I'm not a guru of gitlab and don't know all the common parlance.
Anything else is self-managed.
There is no clear date to when its going to be enforced.
But its definition is so wide its crazy, its basically any egress data except the web interface and shared runners.
AFAIK it will also include git clones!, so if your project suddenly gets popular the users clones will cost you many too.
Also if you use your own runner, cloning the repo to the runner will also be included in your bandwidth limit.
And since it should apply to GitLab pages, it becomes useless for anything the you want to get few visits.
Since GitLab is behind cloudflare, you might as well just use Cloudflare Pages at this point.
You make it sounds like it's horrible. I'd offer a different take in that sure, it can be a huge negative if done recklessly on GitLab's part. But if done correctly, it can actually be a good thing. I think CI pipelines doing a full clone from scratch on every build (or npm installs or other bootstraps) is extremely wasteful in terms of resources, so I'd be glad if that reduces it significantly.
I have a very old fashioned and "not recommended" setup with Jenkins that only pulls new commits, because it works out of a persistent working directory. Works wonders for a Django/ES6 project and sips bandwidth. I wish more of the modern containerized, start-from-scratch and so on setups would work in a similar way wrt the bandwidth they use.
GitLab.com Saas agents have no sticky-ness, so basically always clone.
To help prevent unneeded traffic when a new pipeline job is executed, GitLab Runner uses a shallow Git clone by default on GitLab.com SaaS that only pulls a limited set of Git commits from the current head instead of a full clone. [1]
There is a feature proposal [2] to add support for partial clone and sparse-checkout strategies that have been added in more recent Git versions. Recommend commenting/subscribing.
Please note that for GitLab.com SaaS Shared Runners, the traffic limits do not apply with mostly internal cloud traffic. [3]
[0] https://docs.gitlab.com/ee/ci/runners/configure_runners.html...
[1] https://docs.gitlab.com/ee/ci/large_repositories/#shallow-cl...
[2] https://gitlab.com/gitlab-org/gitlab-runner/-/issues/26631
[3] https://about.gitlab.com/pricing/#what-counts-towards-my-tra...
We understand people don’t want their quotas being filled by something outside of their control. For Open Source projects, we have GitLab for Open Source which contains higher limits from the Ultimate tier (250GB Storage, 500GB transfer/month). In addition, we intend to look into other ways to address your concern such as counting only your own traffic or allowing you to limit external traffic.
Does it? We use GitLab at work, and we have to use local runners (for regulatory reasons + GitLab runners don't support what we need anyway).
Unlike other CI/CD providers (e.g. teamcity), GitLab Runners don't have a local git cache. So if you need to do a clean clone for an important build (gitlab runners don't clean up well after themselves), you need to re-clone the repo from GitLab.com.
With a 1:2 ratio for storage vs bandwidth (10GB storage, 20GB bandwidth per month), assuming using 1GB in latest commits, and 2GB total repo size (e.g. a vendored dependency that doesn't change frequently):
- 2 runners - Clean once a week - 5% of bandwidth per shallow clone - 40% of bandwidth per month on CI/CD alone.
Leaving space for 6 clones by developers (hope you don't upgrade your dev machines often).
If you're using a submodule in multiple projects you're going to tear through your bandwidth.
> Transfer is the amount of data egress leaving GitLab.com, except for:
> Paid plans only: self-managed runner transfer and deployments. This is determined by transfer authenticated by either a CI_JOB_TOKEN or DEPLOY_TOKEN
(https://about.gitlab.com/pricing/#what-counts-towards-my-tra...)
I guess someone has been backing up their movie collections to Gitlab or something.
I replaced it with a fixture that was 2GB of space characters, which compressed down to about 3KB. I know there's a canonical file bomb zip file that's under 1K but there's clever and then there's clever.
IMO GitLab does the wrong thing. They should have enforced those limits from the beginning. And if they didn't, they should've eaten those expenses or at least grandfathered old repos.
Why should they eat the expenses of abusive users? As well they clearly have and are finally taking action about preventing it. Your entire post seems extremely entitled to their money.
However, sometimes tracking these things are necessary, and since there isn't an obvious companion technology to git for caching large media assets ("blob hub?") or tracking historical build output ("release hub?"), devs abuse git itself.
I wish there were a widely accepted stack that would make it easy to keep the source in the source repo, and track the multi-gb blobs by reference.
It's basically one click/command away.
The plugin is installed out of the box in many git distributions now. Many hosts support it today, including Gitlab, which is relevant to this article's discussion: https://docs.gitlab.com/ee/topics/git/lfs/
GitLab is $60/m (edit, per year? unclear in pricing page) for 10GB storage and 20GB/m transfer. So, I can checkout my repo twice a month!?
GitHub is similar, with a 1:1 ratio for storage and transfer. So you can checkout your repo once a month.
> tracking historical build output ("release hub?")
There are OCI artifact registries.
For me, there is rather issue that when all you have is a hammer, everything looks like a nail.
Especially since one of the bugs I'm likely to fix is, "why is this library so damned big?"
Artifact size depends on what garbage from the project infrastructure gets included into the artifact. So simple things like not having a fully populated ignore file will cause things to get included into the project or the artifact.
People who don't care about file size don't care about file size. If the artifact is ridiculous, sometimes the repo is ridiculous too. Therefore if you take a bunch of projects with outsized artifacts, you are going to have above average repository sizes as well.
Because it sounds like you want to use 100% of the service and not pay for it.
In a way I'm almost glad people will have to start cleaning up after themselves.
The communication sent to impacted users via email and future in-app notifications includes only the applicable enforcement dates and limits.
Issue to follow: https://gitlab.com/gitlab-org/gitlab/-/issues/368150
- If I tag a docker image with multiple tags, and then push it to Gitlab, each tag counts towards the storage limits even though SHAs are identical. eg 100MB container tagged with "latest" and "v0.5" uses 200MB of storage.
- The storage limit is not per repository, but per namespace. So 5GB free combined for all repositories under your user. If you create a group, then you get 5GB free combined for that group. Does this include forks? Does this include compression server side?
- The 10GB egress limit per month includes egress to self-hosted Gitlab Runners in free tier. Consider this with the 400 minutes per month limit on shared runners.
These limits feel less like curbing abuse and more like squeezing to see who will jump to premium while reducing operating costs. Is this a consequence to Gitlab hosting on GCP with associated egress and storage costs? Is this a move to improve financials / justify a market cap with fiscal storm clouds on the horizon? Is this being incentivized by $67m in awarded stock between the CFO and 2 directors?Stock history over last year for GTLB (since IPO in 2021?): https://yhoo.it/3QaExCs
From the golden era of 2015: https://about.gitlab.com/blog/2015/04/08/gitlab-dot-com-stor...
> To celebrate today's good news we've permanently raised our storage limit per repository on GitLab.com from 5GB to 10GB. As before, public and private repositories on GitLab.com are unlimited, don't have a transfer limit and they include unlimited collaborators.
> If I tag a docker image with multiple tags, and then push it to Gitlab, each tag counts towards the storage limits even though SHAs are identical. eg 100MB container tagged with "latest" and "v0.5" uses 200MB of storage.
Any "duplicated" data under a given "node" (be that the root namespace, a group or a project) counts towards the storage usage only once. So images latest and v0.5 would only represent 100MB in their namespace registry usage, not 200MB.
> Does this include forks?
The registry data is not copied/duplicated when one forks a project. So this is not applicable. But even if it was, as long as the fork and the source are under the same root namespace, any "duplicated" registry data across the two would only count towards the storage usage once.
> Does this include compression server side?
Yes, the measured size is the size of the compressed artifacts on the storage backend.
We should be updating the docs shortly to make these answers more transparent!
> something went wrong while loading usage details
On my free group’s storage page.
Quoting the requirements:
---
Who qualifies for the GitLab for Open Source Program? In order to be accepted into the GitLab for Open Source Program, applicants must:
- Use OSI-approved licenses for their projects: Every project in the applying namespace must be published under an OSI-approved open source license. [2]
- Not seek profit: An organization can accept donations to sustain its work, but it can’t seek to make a profit by selling services, by charging for enhancements or add-ons, or by other means.
- Be publicly visible: Both the applicant's GitLab.com group or self-managed instance and source code must be publicly visible and publicly available.
---
[1] https://about.gitlab.com/handbook/marketing/community-relati...
Maybe it's a generational question, but I would be pretty happy with 1-2GB (i.e. git history, and one artefact per platform).
I agree with this. This move makes perfect sense. I think GitLab does a lot already for the people that don't pay for the service.
---
In order to be accepted into the GitLab for Open Source Program, applicants must:
- Use OSI-approved licenses for their projects: Every project in the applying namespace must be published under an OSI-approved open source license. [1]
- Not seek profit: An organization can accept donations to sustain its work, but it can’t seek to make a profit by selling services, by charging for enhancements or add-ons, or by other means.
- Be publicly visible: Both the applicant's GitLab.com group or self-managed instance and source code must be publicly visible and publicly available.
---
("can’t seek to make a profit by selling services, by charging for enhancements or add-ons")
I think it's quite fair to say you're not going to give free services out to open source projects that are seeking to fundraise beyond covering their costs.
Well I guess I would have to split out my one little private repo I use for my dotfiles into a separate account in order to qualify.
Also my team's biggest repo is a 2.5 GB checkout but gitlab (self-managed) reports it as 185MB "files" and 353 MB "storage" (no CI/CD artifacts).
It's in Python so runs pretty much everywhere *nix out of the box.
Thanks for sharing this tool to help cleanup the Git history. Please be aware that it will rewrite the history, which could be very impactful to existing branches, merge requests and local clones. A similar approach is described in the documentation: https://docs.gitlab.com/ee/user/project/repository/reducing_...
If you need additional help with analyzing the storage usage, please contact the GitLab support team. You can also post cleanup questions on our community forum. [3]
[0] https://docs.gitlab.com/ee/ci/pipelines/job_artifacts.html#w...
[1] https://docs.gitlab.com/ee/user/packages/container_registry/...
[2] https://about.gitlab.com/pricing/faq-paid-storage-transfer/#...
At least gitlab is not deleting any data, just rejecting pushes if you're over the limit.
As far as I know, the quota itself isn't new, it just wasn't enforced in the past.
Basically you need to either be in the top 0.1% of highly active accounts or do something specific that requires a lot of disk space, but for most people 5G doesn't seem so unreasonable.
I'm curious what kind of project would even need such a repository size. From a distant view this sounds like heavily mismanaged build artifacts in the project's git history; or abused storage for free CDN of video data or similar.
I was re-reading the announcement but as far as I can tell it only states repository size everywhere.
The linked cleanup instruction articles also mostly show repo size (git gc et al) and nothing related to the releases section of a repository. [1]
I'm also still unclear about what exactly is part of the package registry storage and whether or not releases in the repository are part of that.
[1] https://docs.gitlab.com/ee/user/usage_quotas.html#manage-you...
That's over twice as much as the disks to store that much data cost. Every month.
Or we pay $20/month to Gitlab. And I can't figure out how the quotas will intersect with "professional", if at all.
For us Open Source devs, neither is a good option. Although I have heard good things about sr.ht / sourcehut. And for the service, it appears to be fair https://sourcehut.org/pricing/
As long as it's on the internet MS can scrape the data so not being on Github is only a temporary defense.
If it has impact on CoPilot or not, is not clear.
I think SFconservancy's articles about CoPilot are very helpful.
Provoking question: so, will it be Ok to train an AI on leaked Microsoft code the publish the models (not the code)?
Of course, MSFT won't accept this, but w know that big corps are hypocrite by default.
You don't even need to pay gitlab, free tier can do for most stuff and there's a generous sponsorship for Open Source.
I don't get all the hate Gitlab is receiving these days.
FreeCAD files get big. And to accommodate easy printing, STLs are also needed.
KiCAD for board layout and schematics can also get larger. And remember, these also have 3d board components too.
Presentations are naturally larger.
Full high-rez pictures eat storage like you wouldn't believe.
Code is small, thankfully.
One such device I have created is hovering around 4.5GB for a full reproduction for the current snapshot. And if you do the command to pull the whole history (and not current), its around 10GB. And, I'm not sure if GL is counting the whole history, or the current? And are they pruning old after a certain date?
GitLab stores the Git history with all data revisions as commits. To reduce the repository storage [0], older Git commits can be pruned but this would rewrite the Git history as an invasive change [1]. To store larger files, it is recommended to use Git LFS [2]. Alternatively, you can use the generic package registry to store data. [3]
[0] https://docs.gitlab.com/ee/user/usage_quotas.html#manage-you...
[1] https://docs.gitlab.com/ee/user/project/repository/reducing_...
[2] https://docs.gitlab.com/ee/topics/git/lfs/
[3] https://docs.gitlab.com/ee/user/packages/generic_packages/
Linux:
$ git clone https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
$ du -h linux/.git
...
2.8G linux/.git
Git: $ git clone git://git.kernel.org/pub/scm/git/git.git
$ du -h git/.git
...
120M git/.git
DefinitelyTyped: $ git clone https://github.com/DefinitelyTyped/DefinitelyTyped
$ du -h DefinitelyTyped/.git
850M DefinitelyTyped/.git
home-assistant/core: $ git clone https://github.com/home-assistant/core
$ du -h DefinitelyTyped/.git
380M core/.git
The above is also only taking into account repo size, while the GitLab limit applies to everything including release artifacts, CI build artifacts, hosted containers, etc. Many of which GitLab currently provides poor or sometimes even zero support for implementing purging.I also really got hooked on their code/project management philosophy, rethinking a lot on how I want to run collaborative projects in general
Driving away individuals is apparently their strategy now.
Sad. I used to teach new developers starting with Gitlab pages.
But after CoPilot I still dont want to use it, on the contrary, I'm switching to GitLab because all the free stuff on GitHub are just unrealistic, and are just to play the long game.
Also I dont want to invest my time and effort into a proprietary platform, like GitHub Actions, and Gitlab CI and most other core features of GitLab are Open Source.
https://sfconservancy.org/blog/2022/jun/30/give-up-github-la...
If GitLab didn't add the 5 person limit per group most people will be fine, but now with the upcoming bandwidth limit, the 5 person limit, and many others.
Its almost impossible to grow on the platform Excexpt with the OSS program, which requires Lawful agreement with each org member.
> Namespaces on a GitLab SaaS paid tier (Premium and Ultimate) have a storage limit on their project repositories. A project’s repository has a storage quota of 10 GB.
Even it's not mentioned as a change nor in the timeline, but that limit does not exist currently.
Repo 130.82 MB Artifacts 12.74 MB Wiki 51.20 KiB Everything else 0 bytes.
Been like this for at least a couple weeks.
This one is public, but I have others that are private with negative values.
Example of a public project with negative project storage: https://gitlab.com/strivinglife/book-raspberry-pi
The headline number is 1.1G but the container registry is 16G.
Edit: You should definitely add a simpler way to delete old pipelines. Having to mnnage it yourself through the API is a pain. (https://stackoverflow.com/questions/71513286/clean-up-histor...)
No blingbling no social-media drama, just clean straightforward code hosting.
It's a shame it's come to this, but I'm confident GitLab didn't make this choice lightly. It must be done in order for them to stay afloat.
Thank you GitLab team for your efforts. I hope you guys are successful in your future endeavors.
[1] https://about.gitlab.com/blog/2022/03/24/efficient-free-tier...
[2] https://www.theregister.com/2022/08/04/gitlab_data_retention...
[3] https://www.theregister.com/2022/08/05/gitlab_reverses_delet...