Processing 40TB of code from 10M projects with a dedicated server and Go
boyter.org
boyter.org
Amen! This is why I am learning Go at the moment and considering using it instead of Python for admin and data processing tasks on a fleet of servers. The single binary deployment makes it a lot easier for users to adopt. Python misses out on a lot of use because of the inability to do this. And No! I do not want to pip install a lot of stuff on the servers just to be able to run this script once. Heck, some of these servers don't even have access to public internet to be able to pip install whatever.
Yes, I have looked into pyinstaller and Nuitka. They threw up some errors that were indicative of deeper issues that I didnt feel like a good use of my time to debug. I'd rather choose a language that has this as a priority/design goal instead.
My initial thought was Python, but it needed a couple of 3rd party dependencies, and there wasn't an overly clean _and_ simple way to copy the script in from the mixin and run locally.
So I shrugged, rewrote my script in Go, and then used a multi-stage build to copy in the binary and nothing else.
Ending up with a single statically linked binary was cool, even if Go does some stuff that made my eyebrows quirk a tad (I still can't believe that an idiomatic set in Go is map[T]struct{}...)
Did you consider, and stop me if this suggestion is completely wild, Java?
People are now irrationally scared of using dynamic libraries, OS packages and even directories, like virtualenvs.
Instead, a simple solution is replaced with containers, or by rewriting tons of code in the new hyped language.
Are we trying to create job security through unnecessary complexity?
As a python/go dev, I can assure you that go is not much more complex. And as a server admin, it's much easier to deploy.
It does have drawbacks when compared to python, but it can often be the optimal solution.
The server might have the wrong Python version available (2.x vs 3.x, multiple evolutions of 3.x with big features added in each point release). Et cetera, et cetera.
Or you could just do it with bash or a static Go binary and be done with it. Portable, works pretty much everywhere.
That's the point of using virtual environments, so that you can run the Python version and libraries that you need. Also, as of 3.3, Python ships with venv which means you don't need to separately install virtualenv anymore. It's all very portable.
It really isn't as portable as one might wish.
Not all operating systems carry Python >3.3 by default, that would mean installing backports or unofficial repositories on production machines.
Or you can use a static binary that just works.
Imagine a situation where you'd need to perform an operation on a hundred servers, would you either transfer one static exacutable or build an virtualenv in every machine and download correct versions all of the related libraries?
I had heard that the Python packaging and deployment story had got a lot better in recent years, but this sounds like it still falls far short of table stakes.
We're finally getting around to migrating away, and we've settled on Twitter's github.com/pantsbuild/pants which is like Google's Bazel except that it's jankier in every way _except_ that Pants supports Python while Bazel doesn't (Bazel advertises Python support, but it doesn't work as the many Python 3 issues in their issue tracker can attest). The nicest thing about Pants is that it builds pex files (github.com/pantsbuild/pex) which are executable zip files that ship all of your dependencies except the Python interpreter and shared objects.
I'm still not very satisfied with this solution, and it's still far, far worse than Go's dependency management and packaging model, but it's a dramatic improvement over Pipenv, virtualenv, etc (gives us both reproducibility and performance).
The point is that the static binary created by go is a single file that you copy into place and run. It is a lot simpler.
See this repo: https://github.com/unixtreme/D3Edit. It has Linux/Mac/Windows virtualenvs and it works without any additional setup
Basically build a virtualenv and then pack that into a .deb, install the .deb on the target machine and you have a self-contained package that requires no external resources to install.
That said, using golang is WAAAAY simpler. The virtualenv/deb solution, while it works, is very sketch.
I've had some success making a Python file an executable. Though, I do understand your gripe with needing to install dependencies.
It can produce a single .pyz file which has all dependencies except the interpreter itself inside. Actually this format is compatible all the way back to Python 2.6, it's only the zipapp convenience scripts that are new.
https://docs.python.org/3/library/zipapp.html#creating-stand...
[1]: http://widgetsandshit.com/teddziuba/2010/10/taco-bell-progra...
And with one dumb word choice, I suddenly can't share this otherwise good piece with most audiences.
Some of the cloud providers have free hosting for public data sets (people who use the data incur cost to download/process the data). I'm not sure if this would qualify.
* https://aws.amazon.com/opendata/public-datasets/ * https://azure.microsoft.com/en-us/services/open-datasets/
https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPay...
That’s an astronomical number. And that’s only what we can all see. Can’t imagine how much more is private.
Also, it look’s like 20% of code is comments. Which feels just about right.
I don’t think the raw number is what’s so impressive, as much as the fact that there’s more code public on GitHub than I could comprehend in a lifetime.
Amusingly, that tiny bit of backward compatibility can lead to vulnerabilities as outlined here: https://www.acunetix.com/blog/articles/windows-short-8-3-fil...
With variably length extensions, there isn't a good reason to not just use the full name.
> the filename “15” & “s15” made the top 50 most popular filenames
Anyone know why?
https://boyter.org/posts/an-informal-survey-of-10-million-gi...
15 is also a common number in coding problem sets that people post to GitHub. It’s a stretch but it might be something to do with 2015 being a year that has a lot of coursework commits, and 15 is a pretty low number (i.e. higher probability a student posts solutions to #1-15 than #1-19).
The same goes for schoolwork. Many people do online MIT and Stanford courses that are from previous years ~2012-2016, but there has been less time elapsed for students to post answers from 2016,17,18,19)
This is mostly conjecture, so I hope someone has a better answer!
Plain 15 is trickier. GitHub's search is too fuzzy to see a pattern at a quick glance, but again it seems that JavaScript is the culprit with a disproportionately high number of filename:15 results compared to filename:14 or filename:16.
It seems generating a `react/angular` project produces many files, that almost never get changed, so it would be interesting to know the most duplicated files...
I think this could also give valuable insight into how to make the language/frameworks better or simpler...
[1] https://github.com/boyter/scc/blob/master/languages.json
- The largest .c file is actually the CATH database, not sure why it has that extension
- The large .cpp file is actually C++ but has a kind of 'data as code' approach defining bonds between atoms
https://medium.com/google-cloud/github-on-bigquery-analyze-a...
Spark would have a been a simple option to do this kind of processing, with less lines of code, and could also run on "spare compute". Same goes for the "How does one process 10 million JSON files taking up just over 1 TB of disk space in an S3 bucket?": there are appropriate file formats for storing and querying big datasets, text/json is simply the least efficient option and likely the cause of the "$2.50 USD per query" number...
https://boyter.org/posts/an-informal-survey-of-10-million-gi...
Nicely written. :)
Stored as protobuf, I estimate it would be 8x smaller. A custom binary format would be smaller again. Not only is that smaller, which saves storage and transfer cost, it's also proportionally faster to process.
NOT Go, but Go Template!! A subset of Go is causing more anguish than all of Python and only a sliver less than all of Rust.
Says it all really ;)
"Parsing 10TB of Metadata, 26M Domain Names and 1.4M SSL Certs for $10 on AWS"
http://blog.waleson.com/2016/01/parsing-10tb-of-metadata-26m...
We used a lot of the same tech: python, golang, S3, 32 core machines. Anyway, nice read and good use of Hetzner.