How Git servers work, and how to keep yours secure
gemini.nytpu.com
gemini.nytpu.com
- users/organizations
- issues
- PRs
- milestones
- releases
- wikis
- activity/contrib graphs
I have it running on a $5/mo Linode instance for some of my personal projects.(I've done a cheaper version of this -- except for the immutable part, and the separation of accounts between devices -- in the past using SSH+SVN to a home server, and it was great.)
I was thinking immutable from the perspective of a device. A given device can pull branches of certain repos, and make commits to the branches. But a device's user account on the git server doesn't have permission to affect past commits. So, for example, if my dodgy Linux smartphone is compromised, a hypothetical person who isn't being nice can't do anything to my backups, other than make bogus additional commits.
Maybe each device has its own branch (e.g., `big-laptop`, `little-laptop`, `smartphone`, `media-server`), where they can commit their changes, and maybethey can pull from main/trunk. And then the physical console for the git server lets me inspect and merge changes from the different devices, so that other devices can pick up those changes.
I thought about starting with Gitlab CE, but that's pretty big, so, even if the features could be made to do what I want, I don't know whether I'd always be running too many vulnerabilities that defeat some of my purposes.
And anyway, seems like good practice, and shouldn't be hard to do, and should fit with workflows already familiar from work.
What would be bad practice is to give less-trusted devices (e.g., Linux development phones, or some disposable PC on which I had to install some sketchy software) access to all my files and backups.
Using git this way might be a simple (given we have to know git anyway) way to give the goodness of backups and selectively syncing various kinds of files both ways with the less-trusted devices.
If your concern is data loss/malware, anything on the git level is going to be insufficient (but can still be useful of course, as you said)
I’d echo the suggestion of zfs snapshots replicated on a separate mirror of disks. I can recommend zrepl to set up the snapshotting/replication/pruning part. Syncoid is another popular one.
But considering the case of a malicious got contributer, access to any of my devices and ssh keys is already a wayyy bigger issue to begin with, and likely entails restoring the rest of the system to a known-secure state due the sheer number of files an intruder could have tampered with outside of version-controlled directories.
Any sort of backup of the git repo would also achieve the same thing.
If you are super paranoid, you could also do git over email ala the linux kernel on sensitive repos and only apply trusted patches yourself.
A few things ...
First, 'git' is built into the rsync.net platform and you can do anything you like with it, remotely, over ssh:
ssh user@rsync.net "git clone git://github.com/freebsd/freebsd.git freebsd"
I personally track a number of repos I consider important and keep my own source trees up to date without running git locally.Second, the ZFS snapshots that are taken, nightly, of your entire rsync.net account are immutable (read-only) so if you clone/update your git repos into your account, they are protected from ransomeware/mallory.
Third, we finally have LFS / git-lfs support which pleases me greatly.
However, in late 2005 / early 2006, when we spun it out[1] as a standalone corporation and registered the domain name, etc., I did request, and receive, explicit permission from the authors/maintainers of rsync to adopt, and use, the rsync.net name.
[1] rsync.net began operation in 2001 as an add-on feature to JohnCompanies which was the first provider of the VPS as we now know it.
receive.denyNonFastForwards
"If set to true, git-receive-pack will deny a ref update which is not a fast-forward. Use this to prevent such an update via a push, even if that push is forced. This configuration variable is set when initializing a shared repository."Shameless plug: https://github.com/CGamesPlay/git-remote-restic
Yes, Gitolite can do this.
R is read access only
RW is read and write access
RW+ is read, write and the ability to overwrite history (rebasing)
Gitolite provides a command-line UI only; you can use it with gitweb or cgit to allow people to view repositories in a web browser.
No issues or pull requests or fancy stuff like that!
Well, to be fair, you don’t need gitolite for that. You can just do that with any account and the right .ssh/authorized_keys file. prgmr.com uses/used(?) this for out of band access to VMs for example.
The latter half of that is particularly striking to me, given that he immediately dives into a shortcoming of nginx... rather than reaching for Apache, he works around nginx's shortcoming.
I can think of a few things I dislike, but all of the failures so far were self-inflicted. Like forgetting to update packages that I self compiled outside of supported repository.
It's not completely autonomous when it comes to upgrades, you have to think for a while before hitting "yes" after seeing the list of packages to update, but so is not Debian in the long run... (some of the servers I have to manage are 13 years old or so, and going through major dist-upgrades is never that pleasant either. It's a bit more hassle, because I don't trust the major version upgrade so I have to run them on a backup VM first just to see whether some issues will crop up).
But having the latest versions of the programs is great, I don't have to second guess myself when writing new programs (will it be compatible?), can use the latest kernel APIs, etc.
I've had bizarre networking bugs pop up on arch that I've not had elsewhere. I also just don't like rolling releases as much as I used to. Since the majority of what I do can be containerized, I prefer a much slower release cadence for my hosts, and anything that requires more up to date packages just gets thrown in a container.
Basically, I think my server workflow just doesn't line up with how arch works. On desktops it's great, since I'll always have up to date video drivers, desktop environments, etc, but on a server I don't usually use things that require the latest and greatest software. I figure as long as it's still getting bug fixes, I'm probably fine.
I know git is distributed by design. So if I want to push code to a pair of servers for better availability, I can do it explicitly:
git push <remote1> <branch>
git push <remote2> <branch>
But what if I wanted to make this transparent but still highly available, such that the remote URL in git push <remote> <branch>
is actually backed by a HA cluster?Some of the software and ops to make this happen is Github's secret sauce. I'm not looking to compete with them, but would love an open source solution that had a better uptime than a single digital ocean droplet running debian. Ideally, I could get there without green-fielding raft consensus shims into a modified git binary.
git remote set-url --add origin $second_url
Not quite HA cluster levels of redundancy, but it's also way simpler to set up. * distributed file system like gluster or ceph for repos
* clustered db (eg replicated Postgres)
* redundant instances of gitea
* load balancing
Am I missing something?The trouble comes when the system receives two simultaneous pushes to the same branch. When ceph goes to merge them, which one wins? There has to be a distributed write mutex. Perhaps this mutex could be acquired in a pre-commit hook on the gitea nodes, but it's absolutely necessary (in addition to the other clustered services) to prevent silent data loss or corruption.
post-receive hook[1] can be used to automate that,
[1] https://git-scm.com/book/en/v2/Customizing-Git-Git-Hooks#_po...
You'll need some forwarding solution, or use an SSH reverse tunnel to punch through CGNat.
Use something like ngrok or localhost.run. For example if you're using gitea, host it on localhost:8080 then run this:
ssh -R 80:localhost:8080 localhost.run
This will forward localhost:8080 to <subdomain>.localhost.run
Now you just pass <subdomain>.localhost.run to colleagues and they'd be able to connect to your self-hosted git instance.
Now do the same for ssh port.
Caveat that your traffic would route through localhost.run, so it's best to not use this for anything serious, or alternatively host your own reverse SSH tunnel on a VPS somewhere.
> alternatively host your own reverse SSH tunnel on a VPS somewhere.
To make a quick version, on a VPS or somewhere, install OpenSSH server. Modify your sshd.conf file adding,
GatewayPorts yes
Then you can use something like this, ssh -R 8080:localhost:22 user@server.example.com
After that, you can use, ssh user@server.example.com -p 8080
from any other computer and it will connect you to the machine you ran the ssh -R from.Carrier grade NAT ISPs sometimes have global scope ipv6s assigned, and if the other endpoint has ipv6 support, too, you can breakout easily using the assigned ipv6.
Rather than that I would recommend reading up on DNS exfiltration techniques [1] and things like pwnat [2] that use faked SNMP reply packets that make routers think they forgot to let a data packet through for hop traces.
And if you have the time, I'd recommend to use websockets as a tunneling protocol because it's very flexible in its payload size and allows compressions via websocket extensions and the srv flags. I wrote a detailed article that explains the WS13 protocol and all its quirks [3]
Additionally to that it's good to know the limitations of a SOCKS proxy, hence that's what most "easy to use" implementations provide. Spoiler: forget ipv6 via socks5 proxies. I also wrote a detailed article about its quirks [4]
I'm currently experimenting with the idea of a DNS protocol implementation that uses multicast DNS service discovery to find local peers and that uses DNS exfiltration techniques to breakout of a CGNAT, but I'm not there yet to write a detailed article about it. It's current research for my stealth browser project.
[1] https://blogs.akamai.com/2017/09/introduction-to-dns-data-ex...
[2] https://github.com/samyk/pwnat
[3] https://cookie.engineer/weblog/articles/implementers-guide-t...
[4] https://cookie.engineer/weblog/articles/implementers-guide-t...
" Using a t440p base as my laptop, best laptop for the buck. bought it as a 4300m model with a dual core. now it has an IPS display, better coreboot+bios update, 32gb ram, i7-4712, 2x 512gb ssds plus a 4tb hdd. all together cost me less than 600eur. hackintosh compatible if necessary, though it's running Arch these days. "
If it's via modded coreboot revision, please do mail me the file when possible @: delio_man@abv.bg
10x in advance and sorry 'bout the Spam!
https://git-scm.com/book/en/v2/Customizing-Git-Git-Hooks#_po...