Shit – An implementation of Git using POSIX shell
git.sr.ht
git.sr.ht
Five hours later, it had turned into less of a simple explanation of "look how simple these primitives are, we can create them with only a dozen lines of shell scripting!" and more into "oh fuck, I didn't realize that the git index is a binary file format". Then it became a personal challenge to try and make it work anyway, despite POSIX shell scripts clearly being totally unsuitable for manipulating that kind of data.
Anyway, this is awful, don't use it for anything, don't read the code, don't look at it, just don't.
Why do you think it's the wrong way? I sit somewhere in between and think that some people want to know the details and learning from inside is a good idea. But some other people want to simply be users and for the tool to get out of their way - and that's also good. So if the docs or the UX make either way hard or less effective, that's on the docs or the UX to improve.
Maybe if I play my cards right, I'll use git for the rest of my career. If not, maybe there will be something new eventually, but I imagine that the concepts learned in mastering Git would still be useful.
Every one of the projects that I interact with regularly are in a Git repo on some kind of Git hosting service and the projects are run by people who understand/use Git regularly. For those projects, switching to Mercurial is a net loss, even just considering the time it takes to migrate the codebase + related processes (think CI, issue queue integration, even the repo hosting itself).
Sure, I could use hg-git, but that doesn't gain me much either: now I'm the guy with the weird setup. If something goes wrong with my setup, it's too weird for other people to help with. If something goes wrong with somebody else's setup, I'm not that helpful because I have a weird setup.
Mercurial works great, has a sane CLI and can both manipulate and interoperate with git repos.
Many of the tools mentioned, at their core, are manipulating the same DAG data structures.
What do Google and FB use as their VCS? Also, any source? I don't mean this to be an attack, I am just genuinely amazed by the statement.
https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
But I'm a user of it - it works even if I don't understand exactly how and nobody tells me that I learned using the mouse "the wrong way" because of it.
Git is 100x more complex, if not more.
Give me a break. I'm sick of people glamorizing the idea that you should have your hand held through each step of every tool you use and never expend any effort on becoming an expert in the tools of your trade. Git is an engineering tool, designed by and for professionals. Imagine this kind of obscene complatency in other fields.
There's often sentiment that people should know more, but I don't think I've ever seen anyone saying starting programming from high level is "the wrong way".
Example from out of it: doctors learn both how to use USG and how it works. But in every case, I've seen it in that order: practice, then internals.
You're talking to the wrong crowd with me, you know. I disagree with this approach, too. Maybe we start with "hello world" to get a taste, but the first thing we should do is start breaking it down.
$ git remote -v
But to show all branches it's
$ git branch -a
It's like it's been purposefully designed to be obtuse.
But the argument is being put forward that to understand the command lines you have to understand the internals.
`git remote` and `git branch` list the remotes and branches respectively.
Adding -v makes both of these verbose. It will additionally show what each branch/remote is "pointing at".
Adding -a to `git branch` shows remote tracking branches in additional to local branches. This not normally interesting so the default is to list only local branches.
There's plenty of weirdness in Git, but honestly my main complaint is that the interface is awful and the documentation makes Dostoyevski look modern and sleek.
A version control system is tackling a non-trivial problem. Go learn it properly, otherwise you'll be one of the 'users' that, at best, will be stumped on trivial issues, losing productivity and running to others for help. At worst, you'll be making bad decisions and dragging down your team.
Would you also say that you don't need to learn anything and can just "guess" while working with a programming language?
I have no problem with people on my team asking each other for help, and I definitely don't consider it bad for productivity when they do. If someone on my team suggested people asking them for help was bad I would bring it up in their next one to one because that's a really bad sign something is wrong.
If everyone on my team decided to learn the internals of git so they didn't need to ask one another when a problem arose I would be genuinely concerned about how the team is working.
The only difference is that grasping the cursor will probably take you minutes because it's a simple concept, and grasping the data structure takes a bit more effort because it's just a more complicated topic.
You don't need to know the implementation details of git, but you need to know the data structure it operates on, cause otherwise you're just walking in the dark.
But you do need to understand that the mousepad doesn't correspond to points on the screen, and you have to learn to treat it more like a treadmill than anything else. Going back in time and thinking about it from a rollerball perspective can help with that -- new users have a tendency to use something like 90% more space because they don't grok that for long movements they have to pick up the mouse.
People are bringing up the mouse as simple because they're used to using mice. But hand anyone a mouse for the first time and you'll find out that they aren't simple. They're just doing comparatively less than Git, so the problem space is slightly easier to tackle. And that's even ignoring the hand-eye coordination problem we take for granted, and that can take weeks for someone new to computers to get over.
Talking about internal mechanics is broadly useful when teaching computer literacy -- everything from mice, to copy/cut-paste, to shift-selection of files, to the file browser itself benefits from trying to build a systemic, mental model of some kind of behind-the-scenes abstraction.
Who knows, maybe Linus will have another epiphany, while Microsoft somehow mismanages GitHub and squanders all the goodwill away. Then, a group of upstarts...
That said, wanting to learn git's internals for the sake of knowledge is fine as motivation.
Because the porcelain is a nonsensical pile of crap on it’s own so you really can not make sense of it from the top down, it actively resist that approach.
That’s not an assertion that it’s a good thing mind, just that it’s the only one: learning git from the bottom up is much easier than top-down, and people who dislike that approach are simply hosed.
You could, but its easier if understand that basic abstractions which are its fundamental building blocks.
That’s complete nonsense. A table is not a low level implementation detail of sql it’s a core feature.
And I don’t have to known how tables are represented on disk or what they store exactly to acquire a good intuition of how things work.
SQL is in and of itself an abstraction decoupled from the underlying concerns of implementation and execution. Something git’s porcelain definitely is not.
> I said that the internals were so simple that you could implement a workable version... inside of an afternoon. So I wrapped up what I was working on and set out to prove it.
Been there. Done that. With other things, not git.
I suspect many others here have too.
> Five hours later, it had turned into less of a simple explanation of "look how simple these primitives are, we can create them with only a dozen lines of shell scripting!" and more into "oh fuck, I didn't realize that ...". Then it became a personal challenge to try and make it work anyway...
Yep. Been there too. Done that too. Again, with other things, not git.
I suspect many others here have too.
Also, you need to generate a tree out of something. Could just hash the entire worktree every time, but that would be pretty lame.
(Note that index doubles for other functions like merging/conflict resolution, but I never thought that was a good thing, and could be separated out.)
Why? I prefer a each branch to have its own staging area and working tree, which maps better to my mental model of "branch as an under-development feature".
Currently my workflow to achieve this involves a lot of stashing.
Someone proposed splitting this up the other day, but even that would come at the cost of performance and an increase of complexity.
Maybe it isn't because it would necessitate creating a lot of blob objects as you staged and unstaged changes, which might not get garbage collected for some time. I can't see any other reason.
Agreed. I always said the best git tutorial is https://www.sbf5.com/~cduan/technical/git/
> The conclusion I draw from this is that you can only really use Git if you understand how Git works. Merely memorizing which commands you should run at what times will work in the short run, but it’s only a matter of time before you get stuck or, worse, break something.
I say this as someone who uses git regularly, and who prefers it to all other version control systems I have tried:
A tool that breaks the principle of encapsulation by forcing you to grok its internals if you are to have any hope of understanding its arcane and inconsistent usage syntax is frankly not a very good tool.
By contrast, I don't understand how vim works beyond the base conceptual level (keypress goes in, character shows up on screen or command is executed) and yet I don't have any trouble using it. I don't need to know vim's internals to use it effectively. Vim is a good tool.
It would have been nice if there was a simpler shell around the complex machinery for those (us?) who don't want to do crazy stuff, who don't need to be able to do crazy stuff and who could settle for only the simple 90% of the tooling like we do with the alternatives, but are forced to use git for external reasons.
But if you ignore the non-ms movement, shortcuts and advanced transforms, it is still a text editor that everyone may use. You can’t put your text (text, not a current mode!) into a state that looks okay but requires a vim guru to continue or start over because something is broken in the model. That’s different from git issues where working copy looks okay, but the branch and merge are broken in subtle ways.
>Plain text is just a lot simpler than a VCS repository.
Than a Git repository, not a VCS one. Not saying that VCS = plain text, but much simpler models exist for merging teh codes.
How much time have you spent trying to figure out how to change the font size in Vim, rotate text 90° in Vim, recalculate a formula in Vim, or insert an image into a document you're editing in it? If the answer is “none”, you probably have a pretty deep understanding of the data model Vim manipulates, even if you aren't aware of it.
I'm not sure if I'm missing some features in Vim or you're actually pulling my leg by forcing me to notice that I know more about text than I care to admin :-)
(I'm not OP, BTW, just a random passer-by)
I run `git listdead` after I merge and delete a branch and do the next fetch (`git pull --rebase`). That lists the branches that can now be deleted.
Then I run `git prunedead` and it actually removes them.
Previously if I ran `git branch` it would list every development branch I had ever created in nearly a decade of work. Now it lists maybe ten branches.
listdead = "!sh -c \"git branch -vv | grep ': gone]' | awk '{print \\$1}'\""
prunedead = "!sh -c \"git branch -vv | grep ': gone]' | awk '{print \\$1}' | xargs git branch -D\""Looks like peepcode was acquired, but the book was open sourced: https://github.com/pluralsight/git-internals-pdf
I'm reading through it now. It starts with the fundamental git structure and works up from there.
Things began to click for me as soon as I read this in its intro section:
> Beginners to this workflow should always remember that a Git branch is not a container of commits, but rather a lightweight moving pointer that points to a commit in the commit history.
A---B---C
↑
(master)
> When a new commit is made in a branch, its branch pointer simply moves to point to the last commit in the branch. A---B---C---D
↑
(master)
> A branch is merely a pointer to the tip of a series of commits. With this little thing in mind, seemingly complex operations like rebase and fast-forward merges become easy to understand and use.This "moving pointer" model of Git branches led me to instant enlightenment. Now I can apply this model to other complicated operations too like conflict resolution during rebase, interactive rebase, force pushes, etc.
If I had to select a single most important concept in Git, I would say it is this: "A branch is merely a pointer to the tip of a series of commits."
But this is not actually correct because a branch can often point to a commit that is not the tip.
The point is that a branch is simply a pointer to a commit that automatically encapsulates all of the parent commits.
As opposed to any other VCS? Feels like that model is the only one that works with SVN too. I struggle to see how "branch is a container of commits" is a viable model to begin with.
One of the biggest hurdles my co-workers have had learning git after having used svn for years is the "bucket of commits" mental model they've built up for branches. A common question is how to merge a single commit.
In git you can lose track of what came from what branch when you start merging multiple back and forth, this does happen with svn.
I have a bash function for it (with a ton of other customizations, but it boils down to this):
function pwlog() {
git log "$@" --graph --oneline --decorate --color | less -SEXIER
}
pwlog --all -20
(...in that "less" command, "S" truncates instead of wraps lines, one "E" exits at EOF, "X" prevents screen-clearing, and "R" is to keep the color output. The second "E" does nothing special, it and "I" (case-insensitive search) are just to complete the word) #export PAGER='less -SEXIER'
#export GIT_PAGER='less -SEXIER'
git config --global core.pager 'less -SEXIER'
git config --global alias.l 'log --graph --oneline --decorate --color'
# git diff ~/.gitconfig
git l
core.pager: https://git-scm.com/docs/git-config#Documentation/git-config...> The order of preference is the $GIT_PAGER environment variable, then core.pager configuration, then $PAGER, and then the default chosen at compile time (usually less).
Now the name makes even more sense. I first read it as sh/git, but reading it as something that starts inside and slowly works its way out is now my preferred explanation of the name.
¿
I wonder if the project (or likely, another one) might be better served by implementing the index using plain text (or whatever else might be more natural for shell wrangling) to elucidate the conceptual structure rather than matching git literally.
PS: The name is very apropos. One doesn’t see too many such fitting opportunities — feels warm and fuzzy to see this one well used :-)
write_hex() {
hex="$1"
echo "$hex" | sed -e 's/../&\n/' | while read -r hexbyte; do
printf "\\x$hexbyte"
done
}
you can imagine other shenanigans with xargs or something, but I think this strikes the best balance between performance and readability (as far as shell script goes).read_int16 and read_int32 don't work on big-endian systems, or if int is 16 bits instead of 32. the latter issue can be easily fixed by explicitly specifying -td2/-td4, but the former issue is not so easy. I think it requires either figuring out the endianness beforehand, or better, something like this:
od -An -tx1 -j"$offs" -N4 "$path" | while read a b c d; do
echo $(((0x$a << 24) | (0x$b << 16) | (0x$c << 8) | 0x$d))
done
oddly, this is used in ls-files already. and yes, I checked: 0x$a is POSIX, and the arithmetic evaluation size must be at least a signed long, which is at least 32 bits.'for x in $y; do printf "$a%s$b" "$x"; done' is equivalent to 'printf "$a%s$b" $y' (assuming neither a nor b contain format specifiers). similarly, 'for i in {1..100}; do printf "$a"; done' is equivalent to 'printf "$a%s.0" {1..100}'. unfortunately, brace expansion is not POSIX, but these are both significantly more efficient (both in code size and execution time) than the loop methods.
sha1sum is not POSIX. I think shell arithmetic provides you enough tools to implement https://en.wikipedia.org/wiki/SHA-1#SHA-1_pseudocode directly, although it may be slightly slower than a C implementation. awk is probably faster than shell.
"'for x in $y; do printf "$a%s$b" "$x"; done' is equivalent to 'printf "$a%s$b" $y' (assuming neither a nor b contain format specifiers). similarly, 'for i in {1..100}; do printf "$a"; done' is equivalent to 'printf "$a%s.0" {1..100}'. unfortunately, brace expansion is not POSIX, but these are both significantly more efficient (both in code size and execution time) than the loop methods."
I didn't understand why you write that part at all, considering the goals of the program we discuss (which is to demonstrate some git primitives in POSIX compliant shell code).
Yes, indeed, thanks for pointing that, it's documented:
"The format operand shall be reused as often as necessary to satisfy the argument operands."
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/p...
So what’s your secret?
Having to learn about the internals is a giveaway that the tool suffers from poor encapsulation.
To me git is definitely one of those tools where one should satisfice and not learn it deeply, because it's not worth the effort. One can successfully stick to a simple workflow and ignore anything git astronauts come up with, like git flow if they want to keep their sanity and focus on what matters - creating quality software. And almost any team has some git fetishist which will be thrilled to help when things go south. And if they don't, it's probably for the better.
Not necessarily. This may mean (and I think in this case, it does) that people are too afraid to learn about those "internals" - or should I say, the mental model behind the tool (and then some of those people write tutorials for others, perpetuating the problem). And with a "monkey see, monkey do" approach, people can fail at anything, up to and including tying their own shoelaces.
There is no such thing as a perfect encapsulation. Not in programming, and especially not in the physical world. "Internals" are ever-present and leak into view all the time. A good abstraction is just one that you can use day-to-day without constantly minding what's going on behind the scenes.
More importantly though, when you're just learning a bunch of git commands in isolation ("monkey see, monkey do"), you're not learning a tool/an abstraction - you're just learning its interface. That's sometimes OK, but in general, for effective use of an abstraction it's better to learn what moving pieces it is abstracting away. Which in case of Git is that it's a DAG. DAGs are kind of fundamental in programming, too; it's good to understand them.
It's a logical fallacy to use this as an excuse for a tool having bad encapsulation/abstraction.
Yes it's impossible to have a perfect abstraction. No that doesn't mean we shouldn't try harder.
> Which in case of Git is that it's a DAG
And one step beyond that, that it's a Merkel tree. It's key to understanding stuff like "if I change a commit, it changes all commit after that" or "if I move (cherry pick, rebase) this commit, I'm creating a new one, not really moving".
Not really, not every block chain is a Merkle tree. Since Git history is not linear, you can’t order the commits in any canonical way. You definitely can order them in some way (like "git log" does) and then construct a tree for that list of hashes, but this is not really useful computation. Git repo integrity is verified simply by HEAD commit hash because you normally clone the entire repository anyway.
$ git cat-file -p d8defd0bb0062ed541de173a2aec834b64d6adbe
tree cbdb56fe9bb1766d8fc2b2e53c9c934efbacbf1c
parent a8a1049c06d100a3f926a82414e6addb9b9af5e8
author ilammy <alexei@cossacklabs.com> 1581533224 +0200
committer ilammy <alexei@cossacklabs.com> 1581533224 +0200
fixup! Avoid unsigned overflow in length computations
And compute the commit hash manually to verify that: $ sha=d8defd0bb0062ed541de173a2aec834b64d6adbe
$ cat <(echo -ne "$(git cat-file -t $sha) $(git cat-file -s $sha)\0") \
<(git cat-file -p $sha) | sha1sum
d8defd0bb0062ed541de173a2aec834b64d6adbe -
Git prefixes object content with object type ("commit" in case of commits), its size in bytes (textual, decimal), terminated by a null byte. And hashes all of that get the commit hash.Moving to Git was the more difficult step for me. I prefer it now. For a while I had to move back to Subversion and I hated it.
Just the idea pains me. Missing git add -p, cherry-pick & rebase -i so much I immediately put git-svn on if I have to go back...
Also, it makes telecommuting easier, asynchronous team work so much simpler...
I think the key 'abstraction' that people don't understand is cherry-pick. I can't explain clearly in fine details /how/ it works, but it is the base of so much of git's power...
It should be possible to learn the commands for creating a branch, uploading our changes or making a commit like it was/is possible for all version control tools and then move on with our professional lives, which likely revolve around writing software and not fumbling with git.
By the way, I love your conversation about Merkle trees below; it was one of the most surreal things I've read lately. :-)
The problem revolves around the fact that, despite same name being used, git!branch != svn!branch, git!commit != svn!commit, etc. They serve related purposes - but not the same, because the concepts behind them are different. Learning a tool means learning those concepts. So in the process of learning "git commit" and "git branch", you're supposed to pick up on the "pointer to a node in a DAG" thing - otherwise you haven't learned "git commit", you've learned something else that's vaguely similar. And then you'll have difficulties when its behavior goes against your expectations.
Now that git's very popular, the teachers have become arrogant and are claiming that our mental models for how VCS work are wrong and we should instead adapt our thinking to the git internals. In almost all other professions a confusing tool is scorned, but only developers are expected to learn how all sorts of weird contraptions work and then anyone who can't keep up is scorned instead.
git is almost 15 years old and here we have yet another attempt at clarifying how it works to the masses. Why are there so many git GUIs and tutorials and attempts to clarify how this tool works? It's a freaking VCS, not rocket science. git took something that used to be straightforward and doable by any developer and turned it into an over-complicated mess.
Now here's a question for you: why do you defend this anti-developer tool instead of siding with fellow developers?
Or that the problem space is inherently difficult.
Coworker: "I really hate Git, it's so hard to understand what's going on internally."
Git guy: "Did you read the documentation?"
Coworker: "Nope."
Git guy: "Did you read Git - book?"
Coworker: "Nope."
Git guy: "Did you read Scott Chacon's 'Pro Git'?"
Coworker: "Nope."
Git guy: "..."
As a matter of fact, git is hard to use.
The question is: keeping the same level of functionality, could it have been made easy to use?
In this particular case, git is a version control tool and it supports various typical operations for such tools. One should be able to learn the commands and then successfully use the tool. If that's not possible, I continue to assert that there's a problem with the tool.
Not to vouch for having Git store things in human format. But I often think about how inefficient JSON API's and YAML storage formats are (in parse time) just to be a benefit of a user debugging it or discovering the API through a browser. But since most people use a JSON prettifier plugin or a tool like Postman anyways, what is the benefit of the line format being character strings? Wouldn't a binary package not be just as easy translatable into human readable JSON formatted output as a compacted JSON string is?
One benefit is that I can look at an arbitrary file/response and be able to tell with a fairly high certainty whether it's JSON, YAML, or TOML, but there's no way that I tell whether it's messagepack, bson, or protobufs.
But I think you should be able to detect the type of format for binary encodings just as well as there spec is pretty specific. Maybe not at glance as a human, but that is the point I'm making. Should all line formats be made absolute human readable and parsable at glance just to the benefit of debugging at the cost of performance. Where with just a simple lens tool you can look at the data in a completely different (human friendly) view. Tools like this already exist in the form of WireShark, only they mostly operate at a deeper level.
Everybody's brain is different but I actually understand all of git's internals (the "plumbing") but it doesn't help me with the git commands (the "porcelain").
Yes, I know that git is a DAG (Directed Acyclic Graph), and that HEAD is a pointer, and the file format of BLOBs and SHAs, etc. If I were to implement a DVCS, I would inevitably end up reinventing many of the same technical architecture decisions that Linus came up with. But none of that insider knowledge really helps me remember git syntax if I haven't been using it in more than a month. Even though I grok git's mental model, I still can't answer the top-voted "git" questions on Stackoverflow without a cheat sheet: https://stackoverflow.com/questions/tagged/git?tab=Votes
The git UI and unintuitive syntax is just too hard for me to remember unless I use it every day.
In contrast... In vi or MS Word, I can effectively modify text without digging into underlying "rope data structure"[1]. In databases & SQL, I can "INSERT INTO x" without learning the "internals" of b-trees[2]. In Photoshop, I can stack layers without learning the math "plumbing" of alpha blending[3]. And yet for some reason, Git in particular needs people to learn it "inside out" more so than other tools. Not sure why Git needs this cognitive prerequisite.
[1] https://en.wikipedia.org/wiki/Rope_(data_structure)
[2] https://en.wikipedia.org/wiki/B-tree#B-tree_usage_in_databas...
[3] https://en.wikipedia.org/wiki/Alpha_compositing#Alpha_blendi...
People regularly use the limited subset of git that Github permits without learning how it works internally. If only the 'edit this file' button were powerful enough to do actual work... It isn't, and that's the other problem: internals knowledge actually helps you do day-to-day versioning tasks. The reality is that one day two developers will submit PRs that conflict, and you'll have to find a way to merge them both, and knowing how rebasing works inside absolutely helps you. The analogy is more MSWord style-stacking than manipulating ropes directly, because Git does manage to completely hide some of its guts. (Object storage, compression, transfer come to mind.)
In vi and Word, you're not worried about state changes outside of saving the current state and, possibly, undoing some number of steps. In a VCS, you might need to check out, merge, or compare arbitrary states from the history, and doing this inherently requires a deeper understanding of how the history is stored. A good VCS should expose these internals in a clear way. In my experience, working with even 1 teammate immediately requires you to have some mental model of how your VCS deals with merging different histories.
That said, it's up to the VCS's interface to make these things clear. Git's mental model is simple enough, and the porcelain can do some of this stuff very well, but CLI is arcane; I end up storing extremely common functions as shortcuts because I'd never remember them or want to type them even though I use them dozens of times a day.
True, but it is very helpful when designing your indexes.
The code is Ruby, but there's enough explanation for each snippet to be able to follow along in whatever language one prefers. I had no problems with translating to Go, for example.
So many people say they "know" git, but then I watch them work and it's "git commit -a" all the time and "git clone" when something doesn't look right. It's really amazing how people refuse to learn this essential tool.
Drew even livestreamed the whole thing, but I don't think he's uploaded it to a PeerTube instance yet.
I can see that it's open source[0], and I'm very tempted to copy it. I'm already in the midst of migrating all of my video hosting to peertube, but I don't have a solution I'm confident in for livestreaming other than Twitch -- especially because when in the rare instances where I do stream coding sessions they can go up to 5 or 6 hours, at which point archiving and storing that video starts to look a lot more costly.
>This is the website for my self-hosted livestreaming platform (aka bag of hacks dumped into a server).
PeerTube is nice in theory but in practice it's been really really unreliable for me.
WORST—-and I mean WORST—case, could crank out a clone during a hack week.
(And yes, I've seen, and boosted, your Mastodon tirade, and am ... apprehensive in commenting here.)
(they were not, this is a shitpost, they're actually called Merkle trees)
Damned edit window...
Archival/storage of video shouldn't be that costly.
With 5400rpm drives (better for archival than more or less any other type of storage media, including faster hard drives), it looks like the going rate is about a United States cent per gigabyte. Two for 7200rpm drives from manufacturers that seem to produce the most reliable drives on the market, consumer-side.
A setup that could survive through a reasonable amount of drive failure, then, seems to be relatively inexpensive, so long as you're not trying to archive your video In The Cloud®.*
*Someone Else's Computer
That's irrelevant, though, given the person's question was about storage and archival.
Clouds are indeed selling that, but I think that's false advertisement. At least from here (western Balkans) it looks this way.
Latency directly impacts bandwidth, which impacts quality, since all current-gen user-facing live streaming protocols that matter (HLS, DASH) are layered on top of HTTP (on top of TCP), and that's already the best trade-off for end-user delivery today.
For VOD it's less of an issue since you can just maintain a larger buffer, but with live that's a trade-off with being closer to the live edge or choosing poorer quality. It works OK for some cases, it's bad for others (like sports, or when letters on the screen become illegible due to compression artifacts).
Building your own CDN off of el cheapo VPSs is theoretically viable, the beauty of HLS and DASH is they're 100% plain old HTTP, so just drop Varnish, add GeoDNS on route53 and off you go. Actually I'd love to have the time to try that :)
Here, the roundtrip latency is ~14ms within the country (e.g. from here to capital city), and 40ms to the closest AWS or GCP datacenters (both are in Frankfurt).
Delivery is cheap. Big clouds just mark it up by criminally high 500%+.
Although all the code was in git (the real git), deployment involved a magical shell script someone in IT had written ages ago. Only after a bunch of rocky, outage-causing deployments did we have the sense to start digging into this magic script.
It turned out it was just a bad, thousands of lines long re-implementation of git using perl and mysql that captured and stored diffs and rsync'd them to production. ...And this was well into the era of CI/CD and infrastructure automation tools.
Eventually, the company brought in new IT leadership that put an end to that kind of nonsense, which freed us in marketing to buy purpose-built PaaS for our needs.
Obviously do this kind of crazy stuff as you please on the side. But at work, build only what uniquely adds value to your company. For most, dev tools are probably pretty low on that list.
I think your anecdote is really just an example of bad management, not bad tooling. Could easily be that some management doofus had prohibited git from being installed on the 'production' system.
A table saw by itself is mostly useless. With a fence and a miter gauge, it becomes useful. With a push block, stop block, subfence, outfeed table, infeed table, featherboard, crosscut sled, tenon jig, and dado set, it is the single most useful tool in a woodshop. Keep in mind it is still one tool and all those accessories are not "tooling", they are accessories to a single tool that increase what you can do with it. The tool always works the same way, and anyone can use it with any of those accessories in any woodshop in the world. In effect it becomes a new, larger solution, made up of many features that extend the utility of the tool.
That's not really what we have. Mostly what we have are jigs. A jig isn't an outfeed table or a crosscut sled. It's a hack for a particular job. If you need to make one specific cut 300 times, you nail together some scrap wood, dial in the miter gauge, angle the saw blade, and make your cuts. And the jig is scrap once again.
But in the heady new world of "DevOps engineering", the jig is now "tooling", and we pat ourselves on the back that we were able to nail some scrap wood together and claim it created business value. Of course, it's not a shitty jig like in the bad old days of shell scripts ("ha ha! remember when we were productive with this simple code that was portable and not gigantic or complicated? how foolish!"), because instead of making it out of scrap wood, we now make it out of scrap steel with a MIG welder. We're advanced now.
And I'll go further. The fact that most woodworkers make their own tablesaw extensions is illustrative of the problem: craftspeople like having fun with their toys. Is there value and experience and dollar savings you get out of making your own crosscut sled? Sure! But it'll also take you 1-2 days of buying parts, measuring, cutting, gluing, clamping, drying, aligning, and finishing. Any business with any sense should have paid $100 to just buy a complete crosscut sled made of aluminum with a good design that will last forever. But they are too dumb to notice they're spending an inordinate amount of time and money on craftspeople making jigs.
I wish the ghost of W. Edwards Deming would rise from the grave and call us what we are: bullshit artists.
That is arguably what P/S/SaaS is, instead of building bespoke or Chef'd instances. But I'm not one to reimplement something in F77 in Julia just to have something to blog about.
Regardless, sometimes a DSL is just what you need, and you'd better have someone who likes creating compilers do it. Otherwise it's like when builders do wood things without talking to a carpenter first.
In which case I agree, for almost all cases. A carpenter trying to build their own saw or chisel would be laughed out of the workshop, and rightly so.
However, the mark of a master craftsman is one who can identify when custom tooling are necessary and knows something about how to build them.
Early contender for Best Analogy 2020.
I return to my original point: if your company is working on something where "production" is novel/unique/a differentiator, then you probably need to invest a little time in how you manage and deploy to production (e.g. you need something that's more than just a jig, and you can't go down to the store and buy it because it literally doesn't exist).
There is probably a certain point in scaling an engineering org (maybe 50+ devs) where you inevitably have to devote some engineering time to this anyway (e.g. you adopted/bought a tool that requires non-trivial maintenance and customization).
If, on the other hand, you're working on something where production and deployment are a well-understood--maybe even commoditized domain--then you should direct your precious engineering time elsewhere.
Marketing website infrastructure has its nuances, but it's well-understood. CRUD apps that talk to databases are a similarly well-understood area.
It's fantastic to what lengths some of us go for a good pun :-)
1. Cache invalidation. 2. Naming things 3. Off by one errors.
This person has #2 nailed down, hard.
Still, fun stuff.
In case there are others like me, this will save some clicking:
I completely require it in shell scripts in general. Good checking of errors is really a fundamental requirement of professional programming.
I don't mean any offense by that - sourcehut is fine - just what popped into my head.
I have it because when I make a command line mistake, I more often say "shit" than "fuck". It feels more natural to type what I'm actually thinking. :)
But for a toy project like this, it seemed appropriate.
I use libgit2, but the code is a bit difficult to read and hack. I know there is a ruby book implementing from scratch.
I want to understand the git internals.
https://news.ycombinator.com/item?id=19540845
(to GP: lots of links related to reimplementing git in there)
Source: https://git.kernel.org/pub/scm/git/git.git/commit/git?h=v0.9...
Better name and contraction of "shell git" ? Or maybe this is not something in common usage. Dunno.
They're much less concerned about their language now.
I don't mean to overstate the case - it's not a swearword or the sort of thing you'd really censor, just a playground term for a mean person. All the same it's an ugly word and (however irrationally) this is one of the reasons I prefer Mercurial to this day.
* Here in our project we use shit, and it's very good. * do "shit pull" / "shit pop" * Sorry, I don't have shit in my machine.
But, again, english is a funny language.