How I sent 300k emails through Github's API in a matter of minutes
badlogicgames.com
badlogicgames.com
To all watchers of the libgdx repository: i’m terribly sorry and hope i didn’t interfer with your work in any way
This is meant as a cautionary tale about using Github’s API on a repository with quite a few watchers (460 in this case).
Earlier this year we migrated our code from Google Code to Github. We didn’t have a good migration plan for the 1200 or so issues back then, so we kept them on Google Code. We now have about 1700 issues on the tracker
Today i finally wanted to tackle the issue tracker migration, using a Python script [1] i found on Github. The script requires one to specify a Github user account that owns the repository the issues will get migrated to. I did a dry run on a fork of the main repo using my Github account, fixed up some issues in the script, and validated things to the best of my abilities. Things looked good.
Then i ran it on the main repository. Luckily i was watching our IRC channel. After about 4 minutes, people started to scream. They each received 789 e-mails from Github. Every single issue i migrated, and every single comment of each issue triggered an e-mail notification to all watchers of the main repository.
This wasn’t apparent to me during the dry runs, as i used my own Github account. The script posts all issues/comments with the user account i supplied, so naturally, i did not get any notification mails.
I stopped the script after 130 issues (4 minutes), and immediately started sending out apologies and a mail to Github support, to which i haven’t received an answer yet. I send roughly 300k mails through their servers in a matter of minutes. If i hadn’t watched IRC, i’d have send out about 4 million mails to 460 people within an hour.
Let me assure you that i’m extremely sorry about this incident. I know that things like this can interrupt daily workflows quite a bit, even if getting rid of those mails is not a Herculean task. I’d be rather upset if a repo maintainer pulled something like this on me. Please accept my deepest apologies.
The lesson for Github API users: think hard about the implications of automating tasks through the Github API if you have more than a few watchers.
The lesson for Github/API designers: consider safe-guarding against such issues in your API, in case other idiots like me pull off something similar in the future.
The implementation you'd probably want to see is for the API's email notifications to be batched if more than small-n trigger within a short period of time. Then the end user gets a notification that "1700 updates have occured" instead of getting 1700 emails.
The API should also be set up with exceptional event detection. Spikes which are several orders of magnitude above normal should get paused in the task queue and flagged for immediate manual review. Users do things you don't expect.
etc ...
In resume: human fails happen. You where smart enough to realize and stop it. That puts you over the average.
I'm glad this can serve a learning experience and no actual damage was done. Sweeping email, while annoying and little bit disruptive, isn't the end of the world and those who choose to understand, will do so.
Some years ago the company I was at had a policy of sending an email on every error - to every developer, so nobody actually checked them. One day the 3 main servers went down, all at the same time. Transpired that a) someone had introduced a bug into the code that completely 500'd a site b) it got through QA on to live and c) there was a relay chain for the email servers (I have no idea why). By the time we'd worked out what was going on the hard drives on all the machines had reached capacity and all the live systems were down. Of course the timeouts that started happening as the load came on just cascaded the issue. Gotta watch those automated emails :)
It's actually a PITA to overcome issues like this on a technical level because you have to run something akin to buffer queue, that works similar to how "debouncing" works.
The best approach as I have found is to...
- You rate limit events as they happen... So you might let 5 events through (within 10 minutes), and then start to rate limit them by adding each item to a queue which you will merge down every 10 minutes (but that exponentially/incrementally back off each time you exceed the 5 items, so the next queue takes 30 minutes before it's popped, and then 90 minutes... etc)
- So for example, you might have an instant pop from queue where less than 10 events have been triggered within 10 minutes.
- Then if more than 10 events have been triggered, you add each item to a queue, and after X minutes, you pop each item off and send a bulk email.
----
It's a real pain to manage such a system, because your "typical" job server, such as Gearman, doesn't let you add a "delay" on jobs...
Ideally, you'd want to make sure that you ignore any new events for at least X minutes... So you are left with the only option of running another pseudo-queue system just to catalogue all of your throttled events.
Let's talk strategy. How else do you guys handle instant email notifications?... without this spamming issue. PS. I'm referring to GitHub implementing this strategy, not the OP in case there was any confusion.
The poster didn't use "send email" API, he was just automatically importing things, and every import triggered emails nonsensically.
I guess there's a commercial reason for this: http://giovanni.bajo.it/post/60836467126/github-is-missing-i...
I was going to make the same mistake myself (thanks OP!). Is there a workaround?
Not that I know of. I added a note to the README about it sending lots of emails to help others avoid accidentally doing this, though.
I doubt that's the case. I bet you that if GitHub were to break out their revenue by plan that the vast majority comes from business plans and GitHub Enterprise. I would be astonished if any of these customers would ever export their dormant repositories for storage.
I'm sure that it's just a matter of them prioritizing this vs. the million other things that are on their backlog. As with most other software companies, your best bet would probably be to start emailing their support with requests for it so that they know it's important to the community at large.
Sure enough, several minutes later all of the pages were updated. All 50 or so pages in each of the 15 spaces. And everyone who had ever touched one of those pages got an emailed for that page.
The nice thing about the Confluence API is that you can specify "minor" updates to prevent exactly this scenario from happening.
I guess since GitHub is built on the git foundation, adding some sort of "silent" flag might not be as easily possible, but certainly it's desirable.
Plus it's up to the developer. We can't have one-click-do-all buttons for everything.
Always figured I'd do Cocos2d-x or Unity for any serious game I do next. I used Cocos2d before and written Unity plugins before. I even have a contractor working on a Unity project right now. Will have to give libgdx a few extra points when deciding in the future, though, for having a caring maintainer.
I actually wrote an OpenGL game engine for Android back before any of the later things came out like Replica Island, AndEngine, the Cocos2D port, etc.. Almost makes me wish I'd open sourced it. It did have some awesome stuff like batching all the sprites with similar draw states together into one draw call.
Glad to see other Android "old-timers", would love to see what you came up with back then.
You may want to turn off auto-subscriptions to repositories you have push access to: https://github.com/watching
Thank you.
I loved it. A history of rum, including all of the politics around it (like the role it played in the slave trade and American independence), great read :)
note: the site is down so I don't know if this is a reference to the original article