AWS CodeCommit
aws.amazon.com
aws.amazon.com
Github is awesome for science. But my code and workflows often are just as reliant on the data I have as the code I've written. With Github, I always have to treat my data separately. A different workflow, a different storage location, different (and often manual) versioning.
Now, I can start integrating my data into the same workflow. No size limits. Just drop it in and version it like everything else. 1GB? No problem. 500GB? Just pay the money.
This is especially awesome because, as a scientist/dev, I do not want to stand up my own infrastructure and servers/vms. I don't want to do linux updates. I don't want to worry about things going down. It just needs to be there when I need it. When I don't need it anymore, I'll take it down. Done and done.
I've used git to push around a lot of binary application packages and it's very nice. Previously I was copying around 250-300MB of binaries for every deployment--after switching to a git workflow (via Elita) the binary changesets were typically around 12MB or so.
We use it for auto deploys from autoscaled instances cuz github has poor uptime.
You can get a self contained cli (jgit.sh) from here: https://eclipse.org/jgit/download/ . It does not have an eclipse dependency.
you can then just add an s3 remote and push to this remote as part of your CI flow.
Here is a decent article that describes the end-to-end: http://www.fancybeans.com/blog/2012/08/24/how-to-use-s3-as-a...
> For best results, clean should not alter its output further if it is run twice ("clean→clean" should be equivalent to "clean"), and multiple smudge commands should not alter clean's output ("smudge→smudge→clean" should be equivalent to "clean").
So this isn't the fault of clean/smudge filters, just the way they were used with git-media.
The most frustrating problem was that filters are executed pretty frequently throughout git workflows, e.g. on `git diff`, even though assets rarely ever change. The added time (though individually small) created a jarring experience.
I'd also be curious how git-bigstore addresses conflicts. It seems like a lot of the filter-based tools out there don't handle them well for some reason.
I'm no EC2 expert, but I believe you still have to do linux updates if you are running on AWS.
Packs happen automatically sometimes per https://www.kernel.org/pub/software/scm/git/docs/git-gc.html
Only because svn downloads the plain text of the server's version of every single file from the revision, so that "svn diff" doesn't have to hit the server.
edit: Come to think of it, the feature they could potentially implement their side is somehow making it so binary files have their history purged somehow. I'm not a git expert but if they are able to trick git into think the files are new, and have no history, each time they change then I guess that could work.
As others have pointed out Git is not a great way to do big data versioning. Git is almost exactly the wrong tool for the job, almost any other revision control system handles large file versioning better because you're not expected to clone entire archive histories to everyone.
[1]: https://docs.aws.amazon.com/AmazonS3/latest/UG/enable-bucket...
It would have several advantages for the user:
1) You could pick your issue management separately from your code browser
2) You don't have to worry that the best tooling (Github) has some of the least reliable storage.
3) If it stored the data in git, it would work offline!
The big downside is that having to "assemble your own" is more complicated for the user, and that it isn't Github/BitBucket's business model...
to me this would mean an issue tracker, and some kind of merge request/review tool. these could be paired with the likes of Gollum, providing a wiki.
One thing though, I would really love to see them support more than just git. Mercurial shouldn't be hard to support, as it has largely similar concepts, and even though SVN is not as popular as it was at its peak, its still widely used and a very good choice for some workflows/use cases.
In the same vane, I would love to see a decent stand-alone code browser that supports different VCS.
On multiple VCS: On the one hand, git is so dominant (for better or worse) that I don't know if the complexity is worth it. On the other, perhaps git is so dominant because github is so dominant. The answer in open source land is probably to set up a reasonable abstraction layer and let interested people provide their own implementations of the VCS-interface.
My key focus for this would be on a more unix-philosophy approach (separate tools that do one thing well).
I like the idea of using a VCS repo (again, ideally any of those supported not just git) to store related information such as issues/tickets, etc however I don't know what the practical constraints of that would be, my first thought is that issues/tickets stored primarily in a (version controlled) filesystem (and then presumably indexed on a server for web views, searching etc) would not be as intuitive as say a wiki like gollum.
If the issues/tickets tool used a flexible SQL model (e.g. choice of mysql, postgresql and sqlite for instance) I would be happy with that, as the data is still quite open (to the system, owner, not as much to the individual developers admittedly)
Customers fear vendor lock-in, so an open platform/ecosystem for git could flourish, like java. Hopefully they'll have a better business model for it than Sun did. Maybe amazon's?
The fact that so much of the world's private source code is on a single public website (GitHub), which has repeatedly been hacked via relatively easy exploits, is pretty frightening to me.
Not sure how that translates as a business model, as no one owns git (like Oracle owned oracle). However, although Sun didn't make money from java, everyone else did - so there is a business model, just as a user of git, not an owner. e.g. reporting tools based on git; CI tools, code quality tools.
So maybe the answer is to just to treat git as infrastructure, for the app you sell, as opposed to trying to make money from the infrastructure itself. Linux is a similar product.
[1] http://fossil-scm.org/ by the author of SQLite.
In theory Github should be leading the way on this (per its open source mission).
Built on-top of it (like how Gollum is a wiki using git for storing documents) is not necessarily bad, but a built-in system is the worst possibly solution IMO
git clone https://github.com/snowplow/snowplow.wiki.git
hg clone ssh://hg@bitbucket.org/BruceEckel/python-3-patterns-idioms/wiki
Also, a project atop git does issues: https://github.com/jeffWelling/ticgitThe idea is interesting to me.
S3 has 9 9s of durability and is "closer" to where our servers run so there are less possible points of failure. We have had 0 issues since migrating to s3 backed git repo 2 years ago.
We still use github for day-to-day scm mgmt as you guys have tons of extra value add.
It's interesting how they seem to be adding all the stuff in the workflow that we are okay with just being good enough and don't benefit from bells and whistles. Github is very cool for managing an open source project, but for my own stuff just about anything which makes managing my repos relatively simple and saves me a bit of time is fine. I don't need much for an interface or features. Just being able to push a button and having a hosted repo ready for other users and with instructions for less technical contributors is awesome.
If I'm already on Github, I'm probably staying there. Competition is good though. Competition will push Github to continue doing more to differentiate themselves as being something much greater than simply hosting Git repo's.
I wrote a blog post about this: http://blog.justinsb.com/blog/2013/12/14/cloudata-day-8/
The git object data itself goes into a blob-store (like S3); it can be stored without strong consistency. It turns out you only need to keep track of a very small amount of metadata consistently (the refs). Riak, etcd, DynamoDB or the Google Cloud DataStore would all be good choices, I think.
I was working on an open-source implementation of Raft as part of this (called barge), but it isn't as reliable as the alternatives above - yet!
Given Amazon's completely different philosophy, it's natural for them to implement git on S3.
On another note, is it just me or has Amazon really been ramping up with their releases lately? It feels like there is something new each week at the moment.
That said, doesn't seem to offer a compelling alternative to Github, based on what they're currently saying. There could be value in tight integration with other Amazon tools, but that seems like it'll come more from intentional lock-in than added value.
It's somewhat annoying I know, but I'm not going to go into more detail as I'm not sure about the legality of my position should I do so.
Edit: On review of their website, I had missed the announcements for Amazon CodePipeline and Amazon CodeDeploy, which between them provide all the missing functionality I hinted at above. And so yes, it looks like this is the internal tool I was referring to.
We've got an amazing Builder Tools team here at Amazon, I must say.
I have limited (but at lease some) experience in working with GitHub Enterprise, and for the longest time, their answer to backup was to turn the entire system offline while you performed the backup. I believe they have since improved in this area, but it was clear that while GitHub itself offers an extremely available service, GHE is severely lacking in this respect.
CodeCommit on the other hand promotes high availability from the start.
That is correct. On earlier versions of GitHub Enterprise backing up data was quite a hassle. For a consistent (repository) backup taken on the VM level you didn't have to shut it down, but you've had to switch the appliance into maintenance mode, effectively preventing people from getting things done. This isn't the case anymore though as we've shipped new backup utilities [1] and support for HA setups with the 2.0.0 major release [2] some time ago.
Something like http://aws.amazon.com/s3/pricing/ combo of storage + requests
The only thing that Bitbucket devs wont add is gist :(
Leaving the conversation about small plans aside, we can set up a plan for you that has as many private repositories as you need. Just email sales@github.com and we can get it set up for you.
I'm not sure how the overwhelmingly dominant player in retail and hosting moving to become the overwhelmingly dominant player in source hosting can count as "competition", except in a Gatesian sense.
But I'm not sure if this gives you much compared to just pointing Mercurial/largefiles (or git-annex) at S3.
Does git have scaling issues? I know git manages the 60mil+ SLOC of the Linux Kernel, that's what it was designed for.
I don't see how even medium/large size enterprise will have difficulty. Or am I missing something?
Or are they?
[1] https://www.digitalocean.com/community/tutorials/how-to-use-...
I also look forward to Perforce's response. For versioned binary files Perforce is still the best. Competition, hooray!
That aside.. This looks great. If I'm not fully satisfied with my setup (right now I'm not) I might make the move to this. Lower friction is always welcome
I think you're right
If they use custom permissions ( directly tied to IAM ), then it would be another closed solution like github.
There will also be a number of companies where a "small" DevOps teams leverages a large number of cheaper systems (owned, leased or whatever) for lower cost.
Do you have any candidates for service termination?
Correct me if I'm wrong, but most failed Google products had no direct revenue stream other than advertisements.
If you were an up and coming fast food burger joint that was going from 1 shop to 2 shops, and possibly more shops, you sure as hell would want to ask McDonalds for it's strategic advice and further advice on how to run your burger shop as well as possible. They've been there, they've done it, not that you need to follow their guidelines, but you definitely want their knowledge in the back of your pocket. They've got about 70 years of experience.
Same with Amazon, who has to keep Amazon.com online or they lose massive amounts of money.
They converted 100% in 2011 according to Jon Jenkins https://www.youtube.com/watch?v=dxk8b9rSKOo#t=449