Building bots to mend badges, or how to get your GitHub account suspended
movermeyer.com
movermeyer.com
I created a bot that would scan for private SSH keys to connect to AWS and other services, it also warned about leaked software licenses for SublimeText and other popular programs at the time. While many people appreciated the initiative, it was not taken the best way by others. Ultimately, GitHub suspended my account and I had to explain what was all about.
One year later, through my employer, I created another bot to scan for security vulnerabilities in projects written in Ruby, Python, PHP and Node.js; this time I already knew that I would need to contact GitHub beforehand to make sure what were the limits of the "automation". They simply stated that — at the time — no automation was allowed, which was quite surprising because CI is automation. Travis and other services are allowed to do things there so I didn't understand why my bot was different.
I reported to my employer that we would need to shutdown that project and move on to something different. One year later, I find that GitHub implemented a (semi) vulnerability scanner for a selected group of programming languages, warning the repository owners about problems with their software dependencies. I cannot be mad about this, it's their service, but it still made me a bit angry.
Ps. Another reason why we mostly host on BitBucket ;)
Checking that pull-request and comparing the URL of the user with the blue avatar, you notice that the username marked as "bot" is actually an App [1] but the user [2] — linked at the bottom of the right sidebar — is a regular GitHub account (with no "bot" marking). I am not sure how are they different at a technical level, maybe the App is a web-hook and the User is the one interacting with the API. In any case, they are two different things.
I just don't think there's any meaningful claim that CI explicitly configured by maintainers using officially supported channels should be lumped in the same category as automated scraping of repositories for creating PRs.
Sure, there are many definitions of automation that would include both of these things, but I think it's obvious what GitHub intends in practice.
Perhaps they didn't mean automation per se, but CI is certainly automation, automatic merging, as in not manually merged by a human.
Perhaps a better rule would have been, no automation on repos not controlled by yourself.
Build your pipeline so that a human has to approve each automated action being taken. The difference between a bot making 5000 network requests in a day and a human making 100-200 semi-automated requests isn't a whole lot in terms of throughput, but makes an enormous difference in terms of quality control and not stepping on toes.
I really wish more companies would do this. Fully automated business procedures that make demands on human attention are just plain awful. Humans should be interacting with humans, machines should interact with machines. Machines can help the human, and the human can help the machines, but interaction points should only be between two of the same types of entity.
Every time I've ever suggested human rate limiting though, I get looked at like I'm a moron. Even when I build the entire workflow myself and tune it so it only takes a relatively tiny amount of time to clean massive amounts of data, people just don't want to do it. It's beneath them. Even when it creates a massive difference in the quality of the product / service you're offering.
So, PRs then?
Just out of curiosity, what's the biggest job of this type you've personally handled?
We're talking thousands of restaurants that we wanted pictures of food from, each of the restaurants had dozens of images we could pull. So tens of thousands of images needed to be sifted through, I figured with the right tooling, myself and my cofounder could put together something really nice that would only need an hour or so of maintenance a day to keep up.
So I built a pipeline that used very basic and easy to build and maintain 'dumb' Rails asset pipeline pages to present data for sifting. Go to the endpoint, it shows you the name of the restaurant and a bunch of images, you select one, type in a name for the dish, and it saves it to the database and puts up another page of images.
It took me bitching up a storm to get him to even look at it. He complained about how long he thought it would take, while I just got to work. Took maybe three weeks to prototype our app. One thing I learned in the process is that if you're looking at a bunch of Southern food, for some reason the picture of shrimp and grits always looks the most appetizing.
I was well on my way to classifying and figuring out novel ways to present the data when I had to make the determination that there wasn't good cofounder fit. So now I work with CNN.
But now all my side projects revolve around ways to get human attention to improve automated tasks. I suppose one of these days I'll get the right idea and/or the right cofounder and I'll give it another go.
There's a wealth of usable information out there on the web that one can build businesses on top of if one only wants to apply a little elbow grease to clean it and turn it into data. It's far easier to scrape data with a regular web browser with a custom browser extension than to try to build out headless infrastructure. But no one wants to do it.
Our main initiative was creating a heuristic based classifier (think lots of regex). At my own initiative, I trained ML classifiers while we worked on it. As development went on, the ML classifiers were rapidly catching up with the heuristic based one. Unfortunately it was kind of a one off data processing task, and when time ran out the regex machine was still in the lead.
I was modestly proud of the legalese DSL generator I wrote up. The lawyers didn't even know they were writing coffeescript as they typed out what documents were, what key dates were, etc. :D
That coffeescript formed the basis of our accuracy testing suite. It was as fundamental as it was huge. That team ended up creating a couple thousand tests in less than a month.
The task of verifying correctness of a pull was much more time consuming than this one. Even if it only took 30 seconds to verify (optimistic), that's 290 hours. Which isn't necessarily all that much for an organization of verifiers (or Amazon Turk), but it is a lot for an individual.
Maybe that should be the cost. But perhaps some things you might be fine with letting a bot do (after manual verification of a statistical sampling, and thorough testing).
The project is currently on hold.
The entire purpose of automation (meaning, automation in general, not just business automation) is to take humans out of the equation for mundane and repetitive tasks[1], and have them deal only with exception scenarios and edge cases cannot be properly handled by the machine (either by design or due to system limitation).
Guess what happens when you have humans approve each and every automated action like you suggest? You defeat the purpose of automation, and users end up hating the system because mundaneness and repetition are reintroduced, except in a different context.
[1]The reason you want to do this is because the more mundane and repetitive a task is, the more likely people are to make mistakes, and mistakes can be costly. In fact they are often more costly than the labor itself.
But - picking the right photo for some restaurant, as GP stated in another comment? Make good UI and perform manually. Alternatively, train NN in background - but it might not make sense in a startup world where the effort for this would be prohibitively high. In the end it's the photos that matter to your business, not that fancy photo picker algorithm that took ages to develop and that you can't sell to anyone else.
That’s not a good example though because it is unlikely to be repetitive and mundane. It’s a decision most restaurenteurs make at most a few times for each restaurant.
I could image though a system where there was some sort of community managed github bot. Developers could submit pull request to the community service to fix common issues. Github would then run the service nightly themselves. Developers could opt-out of the service if they wanted. Something like this could be very handy for many things - security issues, typos, broken links, etc.
The main issue is one of discovery.
Though I imagine you could build an application to notify users of new fixer applications. Maintainers would opt into that for their accounts/repositories, it would then match repositories & applications submitted to it and ping submitters when an application looks… applicable.
> Github would then run the service nightly themselves. Developers could opt-out of the service if they wanted.
It would be just as bad as TFA's.
I have looked into the pull request and discovered that this is a variant of "Bug #4" from the blog post. It happens when the third-party renames their forked repo. At this point, the names don't line up and my bot doesn't realize that the two repos fork to the same location.
I have manually fixed my merge request for your repo and will be writing a script to look for others that might have had a similar experience.
Sorry once again.
This sounds like a terrible idea! I wouldn't want an automated bot trying to auto-correct my work in this manner
It just opens a PR though, it's not like it actually breaks anything.
The one big issue is that it's spammy.
At the same time, I wonder how you could contact maintainers to see if they're interested, I feel opening an issue would be just as spammy, to say nothing of DM-ing maintainers, or tagging them on issues in your own repository.
You'd want this sort of behaviours to be opt-in, but at the same time you'd be limited by awareness. It's not like this is a big/complex change so chances are the maintainers just don't know about the issue or how to fix it.
And since the bot is only creating pull requests, I don’t see any harm: worst case for my repo, it would brake the readme but I double check it just like any other pull request, realize that it messed up, and fix it myself (but I would be thankful the bot noticed the broken link and I have a motivation to fix it).
What is a bit problematic about this bot would be that, due to a bug, it starts spamming (creating 1000 of pull requests) flagging false positives, etc).
Furthermore, it is also important where to draw the line. A bot that notices something is broken and offers me a fix is ok. A bot that notices I use a working service X and offers a pull requests to use service Y could be problematic because it might be useful but might also be annoying (because it is advertising, and service X might be good enough for me.
For my part, I manually corrected all of them and apologized to the maintainers for any inconvenience. The corrected pull requests were accepted, and the bot went on to submit correct fixes to several hundred other repos.
There's always the opportunity for bugs, but once they were ironed out it was able to happily submit correct fixes for hundreds more. I think that makes the idea worth something.
Maybe the other service would do it if you emailed them?
I've reached out to them to try and see whether they want help.
[UPDATE: they're benevolent and I'll be working with them on this. Cheers]
I must say this made my day... It's one of those occurences when one uses many many many hours for a project that in the end could be solvable by a small amount of money ($10) and a few e-mails. Must admit I didn't think of it either when I was reading the blog post. :)
Maybe a solution would be for someone to create an app like Greenkeeper, but which promised to start by doing five things, but to add more over time, informing you of each new thing and letting you opt out at any time: a list of checkboxes.
GitHub doesn't seem to have an ambition as for being the world’s monorepo either, the features they’ve been building is not usually in line with that. I think GitHub should consider creating a team of people that think about the next 10-20 years of open source development and how THEY can carry the flag in terms of innovation in being a world-scale code repository.
Whilst fixing problems is fine, spamming developers is not. It would be interesting to find out more about this 'bot'.
Check out the few times he has over 200 contributions in a single day, most of those are issues opened with a bot and so on.
All one needs to do is take a step back and take stock of how many real-life situations we find unsolicited anything acceptable, and the real potential for pitfall would've been clear.
Even if Shields doesn't support the specific third-party service integration you're looking for you can generate a badge using the incredibly simple image API:
- https://img.shields.io/badge/hacker-news-orange.svg
- https://img.shields.io/badge/rate-limited-red.svg
The API is open source: https://github.com/badges/shields
I started the project although it's maintained by Thaddée Tyl and Paul Melnikow these days. Here's a bit of backstory: http://olivierlacan.com/posts/an-open-source-rage-diamond/
But yes, this was a particularly trivial problem to tackle. It was meant to be a stepping stone to a truly useful bot that I was working on. However, that has been put on hold.
FWIW, the maintainers themselves gave lots of positive feedback on the project.
If your bot has output, always make your bot act like a person. That means messaging, and that means timing. Even in the best case, if your bot uses few resources and always perfectly does the right thing, people don't like bots.
> There are four very important things that any automated message needs to do in order to help avoid aggravating people: Be Accurate and Useful, Be Honest/Open about being a bot, Have a mechanism for feedback, Be Friendly
No. God no. There are two important things that you need in a PR: Don't act entitled (op did a great job there) and don't waste my time (op failed hard at this). Everything else is bad. Telling someone that your pull request is coming from a bot only hurts your goal. In the absolute best case, they treat your PR like any other. In many cases, though, knowing that a message was automated will get you instantly reported for spamming regardless of how helpful you were.
> Automated messages should describe themselves as such.
This is off topic and therefore violates the "don't waste my time" principle. It also has a tendency to engage the gag reflex.
> It should be the opening line.
Having multiple lines for something so small violates the "don't waste my time" principle. And definitely don't start your message with something that is off topic.
> announcing it as automated helps explain why they are receiving the pull request
They are receiving the pull request because something is broken and you are fixing it for them.
> Have a mechanism for feedback
They can put feedback on the PR. This violates the "don't waste my time" principle.
> I ended up settling on the following message for the pull requests:...
Holy crapballs that's verbose. This definitely violates the "don't waste my time" principle. "Fix broken badge by pointing to working URL foo [see: bug_report_link]". Boom. Done. It's easy to read, easy to understand, and easy to approve.
> Note that the last paragraph is only included in the message if the README includes the “download count” badges. I debated working out a system to delete these badges automatically
You should have either skipped them or maybe filed an issue instead. "The download count badge in README is broken because the foo API no longer exists". Not a whole paragraph.
> Do not make automatic unsolicited pull requests.
Most pull requests are unsolicited, and GitHub has an automation API for pull requests, and their ToS doesn't prohibit unsolicited automation, just "excessive" such, so this is probably the wrong takeaway.