Exploding Git Repositories
kate.io
kate.io
Oct 12 15:47:52 x99 kernel: [552390.074468] Out of memory: Kill process 7898 (git) score 956 or sacrifice child
Oct 12 15:47:52 x99 kernel: [552390.074471] Killed process 7898 (git) total-vm:65304212kB, anon-rss:63789568kB, file-rss:1384kB, shmem-rss:0kB
Edit:Interesting. Linux didn't kill Chrome, it died on its own.
Oct 12 15:42:21 x99 kernel: [552060.423448] TaskSchedulerFo[8425]: segfault at 0 ip 000055618c430740 sp 00007f344cc093f0 error 6 in chrome[556188a1d000+55d1000]
Oct 12 15:42:21 x99 kernel: [552060.439116] Core dump to |/usr/share/apport/apport 16093 11 0 16093 pipe failed
Oct 12 15:42:21 x99 kernel: [552060.450561] traps: chrome[16409] trap invalid opcode ip:55af00f34b4c sp:7ffee985fb20 error:0
Oct 12 15:42:21 x99 kernel: [552060.450564] in chrome[55aeffb76000+55d1000]
Oct 12 15:47:52 x99 kernel: [552390.074289] syncthing invoked oom-killer: gfp_mask=0x14201ca(GFP_HIGHUSER_MOVABLE|__GFP_COLD), nodemask=0, order=0, oom_score_adj=0
Seems Chrome faulted first, but it was probably capturing all signals and didn't handle OOM. Then next, syncthing faulted and it started the oom-killer which correctly selected 'git' to kill.How would Chrome 'handle' an OOM anyway? As far as I'm aware, malloc doesn't return ENOMEM when the system runs out of memory, only when you hit RLIMIT_AS and alike.
Took me a good day's worth of debugging before some bright spark piped up and said "wait, you said you were on x86-32...?"
...yeah, I use really old computers.
I fixed my old video card, a GTX 560, and wanted to see what it could run. I loaded steam and PUBG said "invalid platform error". It took me a moment. I hit alt-pausebreak, presto, Windows 32-bit. Whoops.
Hadn't had that problem in a long time except at clients running ancient windows server versions complaining about why Exchange 2003 won't work with their iPhones anymore "it used to work and we didn't change anything!" (Yeah... but the iPhone DID change--including banning your insecure 2003 Exchange protocols.)
https://pcpartpicker.com/products/memory/#Z=32768002&sort=pr...
https://www.newegg.com/Product/Product.aspx?Item=N82E1682023...
https://github.com/Katee/git-bomb/commit/45546f17e5801791d4b... shows:
"Sorry, this diff is taking too long to generate. It may be too large to display on GitHub."
...so they must have some kind of backend limits that may have prevented this for becoming an issue.
I wonder what would happen if it was hosted on a GitLab instance? Might have to try that sometime...
> Project Goals
> Make the git data storage tier of large GitLab instances, and GitLab.com in particular, fast.
[0]: https://gitlab.com/gitlab-org/gitaly
Edit: It looks like Gitaly still spawns git for low level operations. It is probably affected.
Someone will probably have to actually try an experiment with Gitlab.
My naive question is whether CLI "git" would need or could benefit from a patch. Part of me thinks it doesn't, since there are legitimate reasons for each individual aspect of creating the problematic repo. But I probably don't understand god deeply enough to know for sure.
Visual Studio Team Services has a fundamentally different architecture, but we do some similar mechanisms despite that. (I should do some talks about it - but it's always hard to know how much to say about your defenses lest it give attackers clever new ideas!)
attackers will try clever new ideas anyway if their less clever old ideas don't work :P
Counting objects: 18, done.
Delta compression using up to 4 threads.
Compressing objects: 100% (17/17), done.
Writing objects: 100% (18/18), 2.13 KiB | 0 bytes/s, done.
Total 18 (delta 3), reused 0 (delta 0)
remote: GitLab: Failed to authorize your Git request: internal API unreachable
To gitlab.example.com: lloeki/git-bomb.git
! [remote rejected] master -> master (pre-receive hook declined)
error: failed to push some refs to 'git@gitlab.example.com:lloeki/git-bomb.git'
I had "Prevent committing secrets to Git" enable though. Disabling this makes the push work. The repo first then can be browsed at the first level only from the web UI, but clicking in any folder breaks the whole thing down with multiple git processes hanging onto git rev-list.EDIT: reported at https://gitlab.com/gitlab-org/gitlab-ce/issues/39093 (confidential).
https://github.com/cocoapods/cocoapods/issues/4989#issuecomm...
Clone (--no-checkout):
$ git clone --no-checkout https://github.com/Katee/git-bomb.git
Cloning into 'git-bomb'...
remote: Counting objects: 18, done.
remote: Compressing objects: 100% (6/6), done.
remote: Total 18 (delta 2), reused 0 (delta 0), pack-reused 12
Unpacking objects: 100% (18/18), done.
From there, you can do some operations like `git log` and `git cat-file -p HEAD` (I use the "dump" alias[1]; `git config --global alias.dump catfile -p`), but not others `git checkout` or `git status`.[1] Thanks to Jim Weirich and Git-Immersion, http://gitimmersion.com/lab_23.html. I never knew the guy, but, ~~8yrs~~ (corrected below) 3.5yrs after his passing, I still go back to his presentations on Git and Ruby often.
Edit: And, to see the whole tree:
NEXT_REF=HEAD
while [ -n "$NEXT_REF" ]; do
echo "$NEXT_REF"
git dump "${NEXT_REF}"
echo
NEXT_REF=$(git dump "${NEXT_REF}"^{tree} 2>/dev/null | awk '{ if($4 == "d0" || $4 == "f0"){ print $3 } }')
doneHad the pleasure of meeting him in Singapore in 2013.
Still so much great code of his we use all the time.
Edit: see my comment below before you downvote me.
Here is information about some of the other problems with AMP:
https://www.theregister.co.uk/2017/05/19/open_source_insider...
https://danielmiessler.com/blog/google-amp-not-good-thing/
https://ethanmarcotte.com/wrote/ampersand/
(I do use uMatrix to block 3rd party JS.)
For example:
% ulimit -a
-t: cpu time (seconds) unlimited
-f: file size (blocks) unlimited
-d: data seg size (kbytes) unlimited
-s: stack size (kbytes) 8192
-c: core file size (blocks) 0
-m: resident set size (kbytes) unlimited
-u: processes 30127
-n: file descriptors 1024
-l: locked-in-memory size (kbytes) unlimited
-v: address space (kbytes) unlimited
-x: file locks unlimited
-i: pending signals 30127
-q: bytes in POSIX msg queues 819200
-e: max nice 30
-r: max rt priority 99
-N 15: unlimited
% ulimit -d $((100 * 1024)) # 100 MB
% ulimit -m $((100 * 1024)) # 100 MB
% ulimit -l $((100 * 1024)) # 100 MB
% ulimit -v $((100 * 1024)) # 100 MB
% git clone https://github.com/Katee/git-bomb.git
Cloning into 'git-bomb'...
remote: Counting objects: 18, done.
remote: Compressing objects: 100% (6/6), done.
remote: Total 18 (delta 2), reused 0 (delta 0), pack-reused 12
Unpacking objects: 100% (18/18), done.
fatal: Out of memory, malloc failed (tried to allocate 118 bytes)
warning: Clone succeeded, but checkout failed.
You can inspect what was checked out with 'git status'
and retry the checkout with 'git checkout -f HEAD' yes | head -n536870912 | bzip2 -c > /tmp/foo.bz2
I would imagine you could do something really creative with ImageMagick to create a giant PNG file as well that'll make browsers, viewers, editors crash as well.Admittedly I don't know that much about the inner-workings of git, but off the top of my head, perhaps something with traversing the tree depth-first and releasing resources as you hit the bottom?
This is essentially something that can be expressed in relatively few bytes that expands to something much larger.
Imagine I had a compressed file format for blank files "0x00" the whole way. It is implemented by writing in ascii the size of the uncompressed file.
So the contents of a file called terrabyte.blank is just ascii "1000000000000" ... or the contents of a file called petabyte.blank is "10000000000000"
I cannot decompress these files... what is the solution?
Naive solution, just write to the end of the file and make sure you have enough disk. More sophisticated solution, shard the file across multiple disks.
That seems to be the problem. I mean, if an object expands to something much larger to the point that it crashes services just by the sheer volume of the resources it takes... That is pretty much the definition of an attack vector of a denial-of-service attack.
Being able to express trees efficiently in a data format is an useful feature, but it requires the code processing it not to be lazy and assume people will never create pathological tree structures.
If the tree object format was required to store its own path, then you wouldn't be able to repeat the tree a bunch of times. The in-memory representation would be the same size, but you would now need that same number of objects in the repository. No more exponential fanout.
But that would kind of defeat the purpose of Git for real use cases (renaming a directory shouldn't make the size of your repo blow up).
git clone https://github.com/Katee/git-bomb.git --barehttps://github.com/Katee/git-bomb/tree/master/d0/d0
(INB4 The article suggests Github is aware of this repo, so I have no qualms posting this link here.)
(Yes you would need to add a loop detector for paths and resolve ".." differently but it's not like doing this is conceptually hard.)
Just click here: https://codeload.github.com/Katee/git-bomb/zip/master
...I clicked Download a few seconds ago.
GitHub is still thinking. :/
Edit: After about a minute I got a pink unicorn.