GitHub dropped Pygments
greghendershott.com
greghendershott.com
The approach seemed to be, if things break, people will report it and we’ll fix it.
While this may not be the best approach, the number of languages supported is too high for a person to check each one manually. Generally, I imagine they wouldn't expect a change like this to break anything significant.
[people] use it as a portfolio. [..] To suddenly doink the appearance of people’s portfolios is unfortunate.
It is very unlikely that syntax highlighting errors in GitHub will affect someone's chances of getting a job.
Sure, this switch could cause some issues but they don't seem to be severe enough to kick up a fuss over.
Just switching the library and breaking things at a whim is problematic.
Also, the number of languages may be 316 (including some oddballs like "Unified Parallel C") but that's still a possible number to check for at least for major, obvious breakages. Still, for people that do use Unified Parallel C, adequate highlighting might just be the reason to choose that platform and use it to write your blog in, instead of writing a custom highlighter for prism.js.
Sorry, if you business is code and you decide to support 316 languages, expect people to hold you on that promise.
That said, also: errors happen. But that isn't a reason to give them a pass, just not to put too much weight on such things. It doesn't break the platform at large, but terribly inconveniences some users, and they are very right in being upset, too.
What promise did github make?
Half the difficulty of running a business comes from customers with a sense of entitlement not understanding this.
Taken to the other extreme, a constantly breaking Github doesn't have any value proposition.
Promises are not always explicit and not at all related to laws.
The next Transformers being something to look forward to is a promise, but that doesn't mean not liking it is in any way wrong or that I could sue someone for it. I can choose to leave the movie and never go see one again. The producers of Transformers certainly owe me nothing, but they also cannot tell me how to feel about their handling of the material.
Github promises the most awesome code hosting around. That's a very different thing for many people. And to some people you can live up to the promise, to some people you cannot and it's perfectly fine for those to feel let down. Calling all those people "entitled" is, quite frankly, insulting.
As I said: putting this like it is the end of the world is overreaching, but it is a valid complaint and a valid sentiment. Saying that this is the most important thing on Github, because it happens to be specific your problem is entitlement.
Finally, it isn't true that you are only bound by spelled out things. The law many people cite so often has the concept of Good Faith: http://en.wikipedia.org/wiki/Good_faith_%28law%29 and similar fun things that extend beyond that. So Github _does_ owe me beyond their ToS. (I appreciate that this is probably not a case covered by this)
It's a common sentiment in these circles that only the rules written on the contract are the ones that count, while nothing can be further from the truth, widely varying from legislation to legislation.
the other half the difficulty of running a business comes from making tacit promises and then acting annoyed that the other side holds you to them.
I'm not sure you understand the network effect and how that ties into Github's business.
Some of my repositories that use Logos[1] are now incorrectly classified as a combination of Ruby and Scala[2].
For many languages this is a significant and distracting degradation in the presentation.
I could understand GitHub removing highlighting completely because they feel speed is the overriding priority. That would be even faster than what they're doing now. Languages would look "plain" instead of "wrong". Not my first choice, but a reasonable choice.
The situation now is that they've replaced a library that had been handling highlighting thoroughly, with a variety of text-editor lexers that mostly are not. People like me who already contributed to Pygments, aren't feeling motivated to do this all over again for no good reason. So it seems likely the lexers will remain poor for quite a long time. Which is unfortunate.
Finally, at the time I wrote my blog post, I was speculating about the motivation because GitHub hadn't explained why, yet. Someone later did explain ("because speed") in the issue thread.
If a few minor languages hardly anyone uses as compared to the whole site might need some fixing, this still seems like a win from Github's side of things since per the graph the change did in fact significantly improve render times.
For Racket 99% uses either DrRacket or Emacs. This implies that the lexer deployed is very rudimentary.
Any pointers besides the TextMate documentation for writing lexers are welcome.
For instance, TM syntaxes can legally have recursion loops in them, which TextMate will cut so that the app doesn't spin into infinite recursion. But the precise way that it does this is a mystery.
The pygments design is better for static syntax highlighting.
For static there's tons of choices. Pygments, prism.js, GeSHi for PHP, etc. Any idiot can write a static highlighting system. But none of these can be used in-editor.
For dynamic highlighting, there is only one game in town and that's tmbundles. Only TextMate has support for the 100s of languages in existence, including the new ones that pop up each day.
I would love to replace tmbundles. I know just how to implement it. But the problem is, who is going to write all the long-tail language support? VHDL, Pascal, GAP, AtScript, Julia, ...
- - -
Interesting you mention BBEdit. I have a test file I call "the behemoth" which consists of a python file with 32000 copies of this:
""" """
The challenge is to insert """ at the top of the file and see how the text editor cries in pain. It's torture to a syntax highlighter.To pass, the editor must
1. Load the file quickly
2. Have smooth scrolling inside the file, even after making the change.
3. Color the quotes properly through the end of the file, before and after.
To my knowledge BBEdit is the only editor to pass the test. Emacs is a good 2nd place.
Then depending on exactly how popular a language must be to be supported you could end up breaking language support for quite a lot of them.
And if it does, the problem lies not with the syntax highlighting, not even close...
That said, I don't think anyone will bin a candidate because Github didn't highlight his or her code properly.
I know the argument: Someone, somewhere has a copy of each repo checked out, so we (the nebulous "we") could reconstruct everything from the diaspora of ".git" directories.
It just bothers me to think how dependent OSS has become upon GitHub.
There is a silly joke about that even. It usually goes like --
"Gee, I wish someone would invent a decentralized version control system".
Fortunately the bug was closed.
Is there any GitHub-esque outfit waiting in the wings that provides free OSS hosting?
https://www.fogcreek.com/kiln/
https://about.gitlab.com/gitlab-com/
I must admit, I rather like bitbucket, you can get free private repositories with them too.
"To create a new project you simply register at SourceForge and then submit a new project request. Most projects are approved immediately, and you'll typically get an email notifying you of the approval in ~ 24 hours "
Used by projects like Qt, GnuTLS, Haiku and CMake.
sircmpwn@homura ~/s/K/kernel master> ssh irc.sircmpwn.com git init --bare kernel
Initialized empty Git repository in /home/sircmpwn/kernel/
sircmpwn@homura ~/s/K/kernel master> git remote add backup irc.sircmpwn.com:kernel
sircmpwn@homura ~/s/K/kernel master> git push backup master
Counting objects: 4735, done.
Delta compression using up to 8 threads.
Compressing objects: 100% (1765/1765), done.
Writing objects: 100% (4735/4735), 7.62 MiB | 650.00 KiB/s, done.
Total 4735 (delta 2954), reused 4699 (delta 2930)
To irc.sircmpwn.com:kernel
* [new branch] master -> masterWhat is our definition of a "point"?
If you have an external hard-drive backup of your laptop, that's 2 points of failure, right? If someone else has 10 external hard-drives that they keep in different places, that's 10 points, yes? But what stops you from calling all those hard-drives "a giant single point of failure"? If all of them are destroyed, the data is lost.
I just don't get these arguments... The chances of all GitHub data being lost is probably less likely than BitBucket and SourceForge combined.
I'm thinking: shut down due to the business model not working, or some other business-model variant, such as gradually getting intolerably crappy like SourceForge. Cash flow isn't great; let's introduce "sponsored" code. Etc.
- 3 images of your data (1 original + 2 copies).
- 2 types of media (e.g. extHD and DVD)
- 1 offshore location (e.g. bank vault, your parents' house,etc.)
Of course it's not a golden rule, but it can prevent catastrophic failures.A single issue tracker with hundreds of issues for all these projects. A single database of commiter permissions. Etc.
I'd love to see all that metadata in a separate git repo (just like wikies).
To start with, pull requests could be implemented in the main git tool. They're no longer experimental and many, if not most, git users rely in them in some form or another (Github, Gitlab and BitBucket all support them). Folding them into the core would just standardize all the implementations.
It would also be good to define protocols for collaboration. Off the top of my head, that could mean a fork:// browser protocol that would allow a BitBucket user to fork a Github repository and seamlessly submit pull requests back upstream. Some of that is possible today, but there would be some new requirements around federated authentication to enable this (i.e. how to allow a user who is registered with a different service to create a pull request).
If the mechanics of interoperability are standardized, people will develop competitors to Github and things will get more decentralized. But, as mentioned above, Github has a lot of momentum and no incentive to cooperate with other providers. The only way they're really forced to work with others is if these changes and standards are coming from the core Git project.
Even a duopoly (!) would be preferable to a single vendor.
The initial drafts of CSS 2.1 on the other hand, were published in 2002, yet it was 2011 by the time it became a full reccomendation despite having been in use for a long period by that stage. CSS3 colors (i.e. rgba() etc.) also became a full recommendation on the same day.
> Two months earlier, in March of 1998, CSS 2 had become a W3C Proposed Recommendation (PR) which meant that it was considered "done" and was simply awaiting a procedural W3C member review and vote. For all practical purposes, nothing else was going to get fixed in CSS 2. There wasn't a Candidate Recommendation (CR) phase back then, as evidence by the fact that no-one (including my Tasman team as part of Microsoft Internet Explorer 5 for the Macintosh) was able to implement CSS 2 as specified. The problems in CSS 2 were far more severe than mere errata - we had to develop a full revision to fix it.
http://tantek.com/2011/160/b1/css-2-1-css3-w3c-open-web-stan...
Browsing the issues list, this isn't just "fringe" languages, either. Perl, PHP, Go, and Clojure all appear to have regressed to some degree.
It is really easy to highlight simple things (keywords, numbers, ...). However when it comes to more complex scenarios (e.g. where the type of a word depends on the previous one) then the singleline regex based mechanism shows it weakness. Due to that many language support plugins will yield wrong results when you start to split things like function declarations over several lines, even though it's perfectly legal in the languages. Some things can be worked around with the start/end regexes, but nesting those multiple levels deep can get quite akward and I don't think that they were thought of for things beyond braces and multiline comments.
Therefore I don't know if Githubs move here is a really good choice. However I think their main motivation might be that this file format already has such a big ecosystem due to Textmate, Sublime and Atom and the parser has a high performance so that they went for it.
https://github.com/atom/highlights#using-in-code
(Generic identifier looks the same as class identifier looks the same as a property in an object literal looks the same as a string. I guess maybe orange is just the color used left of an equal or colon, but in that case the string color should be different.)
https://github.com/github/linguist/issues/1717#issuecomment-...
>By using TextMate grammars we also get some nice features like highlighting SQL inside Ruby heredocs. But the main motivation was improving performance.
That's the nice thing about actual software with actual versions that you actually install: it doesn't change out from under you at someone else's whim. No sane person would use a "Cloud" C compiler. Of course, GitHub is just a mashup of online backup and Facebook, so it doesn't matter if it breaks.
Many companies provide existing users with advance notice of technology migration, so they can plan and adapt.
Why is advance customer notice incompatible with Github ops planning?
Github did the most efficient thing: break it and let the people who actually care fix it.
Does the vendor or the customer decide what's important to the customer?
> Why not give existing users the choice of migration and timing?
Edit: -20 downvotes. Previous record was -3.
They are nearly opposite situations. A minority trying to silence a majority vs a majority trying to silence a minority.
In both situations the majority wins.
So no, I think downvotes for a comment that is polite, intelligent, on topic and part of a discussion was the wrong use of downvotes.
I still don't agree with it of course :-) User choice is a bugger on cloud services especially where they are shelling out to run the pygments lexer.
On top of which is probably the interesting issue that there is no longer a binary on/off for most features. AB testing, feature toggles, stayed rollouts all mean we never quite know which version of a service we are running.
Keep in mind that HN has cultural norms for downvoting, it is explicitly not for "silencing" opinions.
Those cultural norms are not written into any of the faqs or guidelines and they only exist within a subset of the users of the site.
I don't know why you were downvoted so heavily! I would have thought that a few downvotes would have been enough.
Cultural norms usually aren't, that's why they're culture.
I normally wouldn't understand this type of thing (others say they don't see the problem and it's quite clear where they are coming from), but in a way I _do_ see the author's point of view. When you build something people really care about, any change, no matter how minor, has the opportunity to impact someone. That's why we all build things, isn't it?
[1] - https://github.com/shurcooL/go/commit/6aad35a0a60fd67927f446...
If Racket syntax highlighting was causing performance issues that were noticeable to Github, performance must have really sucked. Why should Github let Racket drag down its capacity?
How would a "fringe" developer feel about that vendor's brand after their language has become successful? Would they like/trust/recommend the vendor to upcoming developers?
Reading their thread, it sure sounds like they want to work with the community to make Racket code viewing a good experience.
That doesn't fill me with a lot of confidence in this thing.
If a vendor cannot be trusted to host your content without regression, why would you trust that vendor to supply your mission-critical editor?
If this were the case, it would be slightly evil, but very clever business-wise. But we don't know, so all this is is a small conspiracy theory.