New tools for open source maintainers
blog.github.com
blog.github.com
The "reasons" Github offers for minimizing comments are: Spam, Abuse, Off Topic, Outdated, and Resolved. "+1" isn't exactly any of those. (I guess they're kinda "Spam" but they're not unsolicited commercial messages.)
(Personally, I would consider them spam.)
I also think they're still valuable when you have Github and another VCS host, where one mirrors the other automatically.
I was astonished that they didn't do this when they introduced the emoji reactions feature.
I think it could be done on a per repository basis, like, basic behaviour rules just like at forums and other such communities. It's then up to the maintainer / moderators / whatever to allow or disallow +1-style comments, and what the 'punishment' would be. They could then also tweak a setting that auto hides comments if they're shorter than X characters.
Well then the user would learn nothing. A better solution would be to notify them that this is wrong, delete their comment, and then prevent them from commenting on github.com for a period of time (1 hour? or more?)
If you don't have anything more to add to the discussion, but it seems like the issue is not being considered, a "+1" comment seems like the most effective way to communicate the need to the devs.
Ideally a thumbs up reaction would do the same thing, but does it? Do devs use the sort by "reaction" feature in real life?
I have not yet had the chance to be on the other side. I'm pretty sure I'd check that list occasionally, but I bet I'd forget a lot of the time. Especially if I had a busy project. Sorting by "most recent" or "number of" comments would be more likely. And that would make a "+1" comment more effective.
Namespace retirement: Agreed with Nullability, just disable username reuse, it causes a lot more oddities than just cloning projects from unknown sources, old comments link to incorrect users and the like, it's awful.
Accidental PR protection: So this won't affect anyone, if I read it correctly, who types something in the PR description? As long as that's the case, fine, but I definitely don't want to overly impose upon the drive-by contributor... as an occasional drive-by contributor. (Sometimes I fix spelling errors, someone's gotta do it.)
Seems like making a PR out of someone else's "private" changes now requires two PRs: first into my fork, then to upstream.
Not that this is a common use case, but I did consider this two or three times the last five years...
* For an alternative of having shared ownership of open source, see Zed Shaw's post about Launchpad https://web.archive.org/web/20120318224723/http://sheddingbi...
What does: "the changes are not explained in the commit body" actually mean?
Is this opt in?
If you don't like a pull request close it or ignore it. But disabling accidental and “drive-through” pull requests sounds very bad to me.
On popular projects this happens often enough to become annoyance for maintainers.
I don't see any case when opening a pull request from someone else's branch and without a description could be done intentionally.
It does add an extra step if you wish to PR someone else's code (you now have to fetch it from their remote, push it to yours, and then make the PR vs directly cross-repo PR between two repos you don't own).
It's not a heavy extra step though and I don't see this as removing the ability for any number of small PRs.
We get at least one of those a day and they are user error, spam or something. Definitely a useful feature for popular open source projects.
Thus a shed load of commits to try and resolve the difference.
Hopefully this new behaviour does fix the problem. On the other hand, I have a feeling it might alienate potential contributors who just aren't that familiar with git yet and could instead do with some guidance (ie how to use the correct branch to base changes on).
I guess we get to find out. :)
How does GitHub identify "bot accounts"? I thought there is just one type of user accounts. You can give the API tokens generated on that account to a bot. But there is no flag in the settings to make it a "bot account".
Minimized comments: I run the mailing list and can enable moderation at any time.
Retire namespace: I control the web server and so every aspect of the URL after the domain name.
Unwanted pull requests: no different from unwanted anything else.
Furthermore GitHub is great for quickly looking over a project when trying to determine if it suits my needs and for getting an impression of how complete it is, how maintained it is and what the code is organized like.
I love GitHub, it is the single best thing that happened in software development for a long while.
GH is not perfect, but it is awesome nonetheless.
Cloned the repository? You can now declare yourself the canonical source for that repo instead of Github, host it from your own machine, and let others pull directly from you. Not that you'd usually want to do that, but it's available as an option if you do.
If Github closed down tomorrow, we'd lose a lot of code. If Github announced that in one year they were going to close down, then we'd only lose private repositories that no one bothered to clone.
There is the issue of, well, issues. And pull requests and wikis and all that.
But wikis are, as far as I know, just git repos themselves, so that's not too bad.
Issues and pull requests and all of that sort of data probably has to be pulled out through the API into some format that's hopefully useful in other places. Could easily lose a lot of work here if Github went poof.
Not that I think they will or that I want them to, just something that's important to consider when a service becomes a monopoly becomes critical infrastructure.
Github imposes a Terms of Service; I don't have to agree to any Terms of Service to use my own site. I am the Terms of Service.
There is all sorts of cruft in the ToS. Here is something I just spotted: "GitHub does not target our Service to children under 13, and we do not permit any Users under 13 on our Service." If a good programmer has produced a good patch, I want it, whether they are twelve or 92; and they are welcome to post it to a mailing list or send it to me directly. GH have to have this kind of rule because they can't 100% control what goes on their site, 24/7. (Really, if they were smart, they would make that 18.)
Another: If you believe that content on our website violates your copyright, please contact us in accordance with our Digital Millennium Copyright Act Policy. No spectre of DMCA hanging over my own site.
Abuse or excessively frequent requests to GitHub via the API may result in the temporary or permanent suspension of your account's access to the API. I can hammer my own server with as many requests as I want.
GitHub has the right to suspend or terminate your access to all or any part of the Website at any time, with or without cause, with or without notice, effective immediately. GitHub reserves the right to refuse service to anyone for any reason at any time. No comment required. Today, you have a bustling project; tomorrow it's gone.
The age restriction is 13 years because of COPPA. It is there because of (sensible) Federal law.
https://en.wikipedia.org/wiki/Children%27s_Online_Privacy_Pr...
> No spectre of DMCA hanging over my own site.
A C&D could be sent to your hosting provider, your ISP and/or the registrar of your domain.
All of whom have a contract stating that they'll pass off all DMCA and Abuse messages to me, and I let my lawyer handle them directly.
Major downside: You miss a huge percentage of possible contributors.
Mailman just replaces @ with "at" which spam harvesters could easily reverse; I made it better for my users; see below.
Kazinator's got your back!
0:webserver:/var/lib/mailman# quilt applied
privacy
pipermail-to-lurker
0:webserver:/var/lib/mailman# cat patches/privacy
Index: mailman/Mailman/Archiver/HyperArch.py
===================================================================
--- mailman.orig/Mailman/Archiver/HyperArch.py 2012-11-29 21:26:57.000000000 -0800
+++ mailman/Mailman/Archiver/HyperArch.py 2012-11-29 22:27:51.000000000 -0800
@@ -93,6 +93,8 @@
True = 1
False = 0
+def hide_domain(s):
+ return re.sub(r'([\w])([-+,.\w]+)([\w])@([\w])([-+.\w]+)([\w])[.]([\w]+)', '\g<1>...\g<3>@\g<4>...\g<6>.\g<7>', s)
def html_quote(s, lang=None):
@@ -281,10 +283,9 @@
try:
i18n.set_language(lang)
if self.author == self.email:
- self.author = self.email = re.sub('@', _(' at '),
- self.email)
+ self.author = self.email = hide_domain(self.email)
else:
- self.email = re.sub('@', _(' at '), self.email)
+ self.email = hide_domain(self.email)
finally:
i18n.set_translation(otrans)
@@ -412,9 +413,7 @@
otrans = i18n.get_translation()
try:
i18n.set_language(self._lang)
- atmark = unicode(_(' at '), Utils.GetCharSet(self._lang))
- subject = re.sub(r'([-+,.\w]+)@([-+.\w]+)',
- '\g<1>' + atmark + '\g<2>', subject)
+ subject = hide_domain(subject)
finally:
i18n.set_translation(otrans)
self.decoded['subject'] = subject
@@ -466,7 +465,7 @@
d["in_reply_to_url"] = url_quote(self._message_id)
if mm_cfg.ARCHIVER_OBSCURES_EMAILADDRS:
# Point the mailto url back to the list
- author = re.sub('@', _(' at '), self.author)
+ author = hide_domain(self.author)
emailurl = self._mlist.GetListEmail()
else:
author = self.author
@@ -574,10 +573,8 @@
if mm_cfg.ARCHIVER_OBSCURES_EMAILADDRS:
otrans = i18n.get_translation()
try:
- atmark = unicode(_(' at '), cset)
i18n.set_language(self._lang)
- body = re.sub(r'([-+,.\w]+)@([-+.\w]+)',
- '\g<1>' + atmark + '\g<2>', body)
+ body = hide_domain(body)
finally:
i18n.set_translation(otrans)
# Return body to character set of article.
@@ -1049,7 +1046,7 @@
author = self.get_header("author", article)
if mm_cfg.ARCHIVER_OBSCURES_EMAILADDRS:
try:
- author = re.sub('@', _(' at '), author)
+ author = hide_domain(author)
except UnicodeError:
# Non-ASCII author contains '@' ... no valid email anyway
pass
@@ -1221,7 +1218,7 @@
text = jr.group(1)
length = len(text)
if mm_cfg.ARCHIVER_OBSCURES_EMAILADDRS:
- text = re.sub('@', atmark, text)
+ text = hide_domain(text)
URL = self.maillist.GetScriptURL(
'listinfo', absolute=1)
else:
Index: mailman/Mailman/Utils.py
===================================================================
--- mailman.orig/Mailman/Utils.py 2012-11-29 21:26:57.000000000 -0800
+++ mailman/Mailman/Utils.py 2012-11-29 21:26:58.000000000 -0800
@@ -438,9 +438,9 @@
When for_text option is set (not default), make a sentence fragment
instead of a token."""
if for_text:
- return addr.replace('@', ' at ')
+ return 'address-hidden'
else:
- return addr.replace('@', '--at--')
+ return 'address-hidden'
def UnobscureEmail(addr):
"""Invert ObscureEmail() conversion."""Niceties are off-topic there by definition so you don't have to worry.
Git writes e-mails out of commits: see the "git format-patch" command. You can practically send that straight to the mail command.
I remember I wanted to check out what was happening in Postgres, since it's my favourite database. But the sheer complexity of getting to the code diffs and comments and all that... I quickly gave up, because the environment was rather unfriendly to newcomers (actually, I didn't completely give up, I just browsed the Github mirror, which was the only usable part of their version control).
And I'm not blaming them, it works well for the core devs, but they can't expect many new contributors (and contributions are even things like documentation, tests etc.!). Which is a pity.
And I'm even worried about those... what will happen in a few years/decades from now, when the generation of devs that are using mailing lists is not active anymore?
Unfortunately, email lists are not easy to browse if you're not subscribed and browsing past messages from before you described isn't that easy. The easiest way I've found around that problem is to use an email to NNTP gateway like gmane. Though that makes me wonder why project discussion isn't managed via a newsgroup as opposed to an email list.
Not the same as a mail client, but once you opened an individual email there's threaded structure.
> and downloading the mbox file requires one to authenticate with the server.
It does, but also says “Please authenticate with user archives and password antispam” - it's just to make automated extraction of email addresses a bit more work.
Yes, but it's much more cumbersome to navigate the thread compared to using an actual mail client. If web pages of mail archives could present an overview pane for navigating messages, it would go a long way to making them usable in a browser as opposed to just posting links to the immediate parent and the set of children for a given message.
Though I wish that having a NNTP gateway (or using NNTP as a primary method of communication for a project) was the standard as opposed to using a mailing list since it would make viewing past messages before subscribing to a newsgroup trivial.
It does, although discovery of the feature as well as the feature itself could use some improvement. In the header there's "Thread:" and the associated button displays the structure. You can also open a whole thread, but I don't find that particularly useful. The download mbox link inside a message will give you the whole thread as well.
For my own site, I rejected the awful Pipermail (default archiver with GNU Mailman) with its brutally ugly presentation format and text-only.
I found an alternative called Lurker, with a better web interface and some of the features you're mentioning, like threading.
I further customized Lurker. I gave it a button bar with custom icons, and hacked it to support in-line HTML (when people post HTML to the mailing list, it is integrated into the archive as HTML). Because that is dangerous, of course, I wrote a comprehensive HTML cleaner for that.
I welcome HTML in my mailing lists; it's time for mailing lists to retire "text only" rules. It improves the formatting of text, letting you use typewriter fonts, italics and bold appropriately, and real bullets.
Here is an archive message with in-line images, another benefit:
http://www.kylheku.com/lurker/message/20131110.004036.5ed128...
Here is discussion where you can see a branched tree view:
http://www.kylheku.com/lurker/message/20121127.210005.0ca78d...
Lurker mods: http://www.kylheku.com/cgit/lurker/
HTML cleaner: http://www.kylheku.com/cgit/hc/tree/
This lexically analyzes the HTML and filters out all unwanted tags, or unwanted attributes of otherwise wanted tags.
mbox archives reveal people's e-mails; you have to control access to those.
https://sourceforge.net/p/lurker/mailman/message/31487163/
TL; DR: Irrationally against all HTML in mail archives, even if properly cleaned. Thanks for Lurker, though!
That does look better compared to other web interfaces I've seen in the past (though you still have to hover over the icon to view details of a message like the author and subject).
One decent interface I've seen is public-inbox (https://public-inbox.org/git/). It presents an overview with the subjects and threading and allows for downloading mbox files or using an atom feed link.
What makes our git own git repo unusable?
https://git.postgresql.org/gitweb/?p=postgresql.git;
I personally find that quicker to get around in than github.
You have to swipe the cursor over "https://git.postgresql.org/git/postgresql.git" and copy; how lame.
CGIT is so much nicer. It also serves up tarballs out of tags. If you do a "git tag foo-123" and push it out, then on your CGIT interface, there will be a foo-123.tar.gz URL for download (as well as .bz2 and .zip), pulled right out of the repo.
Still, either is light years ahead of github's presentation of repos.
> What makes our git own git repo unusable?
I know; WTF? From grandparent's comment I got the impression that Postgresql must be using CVS, or just tarballs in a FTP directory or something (and that someone kindly made a git repo and put it into GH).
Same advantages of self hosting, simple UI for inexperienced users, all the ability to control your own stuff.
What do you find terrible about them? When browsing email lists like the one used for developing git itself, I find it very easy to go through various patch sets and follow the discussion they have on each commit. My client makes it easy to search for subjects, expand or collapse threads, search for authors, etc.
The emails themselves are easy to read and don't require an excessive amount of scrolling to go through even if the patch set has multiple commits. It's quite easy to see comments on a patch set inline with the patch itself.