[1] https://news.ycombinator.com/item?id=16525735
[2] https://github.com/atom-archive/xray/tree/master/memo_core
There has been some work done by people from the ML@B over at Berkeley (?) to greatly improve github-linguist, the tool GitHub uses (I think) to identify programming languages in GitHub repositories. The project is named lexicon (used to be hosted at https://lexicon.github.io/ but now it 404s...) and the results seemed very interesting.
Tried pinging github support about open-sourcing this work, but it seems it's not going to happen. I had very friendly and helpful people answering me, and I totally understand not wanting to open-source something. I just feel a bit sad for the lost opportunity.
"Every day, millions of files are uploaded onto Github, a coding repository where users share code and collaborate on projects. But how do you tell what language those files are written in?
Identifying programming languages are surprisingly hard. Symbols used in one language often have different meanings in another. For example, ‘#’ in python indicates a comment, while in C it indicates a preprocessor command. Even worse, code in one language can actually contain code in different languages, such as HTML, which can contain CSS and/or Javascript.
At the moment, Github uses a giant checklist to identify unique quirks in the language. For example, if the code contains “:- module”, then it’s probably"); background-size: 1px 1px; background-position: 0px calc(1em + 1px);"> Mercury. However, most languages simply don’t have enough unique quirks for this method to be accurate enough.
The current solution the Github team at ML@B came up with is to use a machine learning algorithm called a Naïve Bayes Classifier. To optimize the program, they used Github’s checklist as a guideline for choosing the correct language rather than a hard-and-fast rule. Currently, the team is working on scraping Rosetta Code, a large repository of code in hundred of different programming languages, for data that they can use for training the program and testing its accuracy. Eventually, they will want to see if implementing other models, such as neural networks, can improve accuracy."
I'm certain I saw more advanced claims of success on this project and that github knows about it, since last time I asked they were 'still talking about it internally', but every link I had is dead...
WELL now that I googled seriously, I think I was too pessimistic, there seems to be some movement on this front at github : https://github.blog/2019-07-02-c-or-java-typescript-or-javas...
Excited to see what's coming out of this new endeavour... Just sent an email to github support, fingers crossed.
[1] http://yjs.dev/
And, not only adding features is doing something; improving reliability or just maintaining the site, monitoring is also work...
So dismissing pre-MS GitHub might not be really appropriate.
CircleCI has raised 115M, hire the best people, and are struggling to release a half-baked UI remake that has been in progress for seemingly ever(that is now being forced on everyone this quarter supposedly).
Not sure what the take away is.. Hard to make sense of it externally; interesting things going on under the surface perhaps.
Given a re-write though, I think the way it's moved forward is a bit of a train wreck:
* It's incomplete information wise.
* They threw the baby out with the bathwater as far as UX. Nice things like when you visit a job with failed tests it high-lights that fact and failures; gone.
* They implemented the new log viewer with a virtual scroller. Browser search can search into it, and it has no integrated search. You can select past the buffer size of course. Zero extra value to the user here.
* It seems like the people working on the project have little product empathy. I would almost think it's been outsourced. When very obvious missing feature or issues(compared to the current UI) are pointed out, we get platitudes like "Thanks for letting us know as we iterate on this. We are all in this together!".
* Supposedly it's now being forced on the user base Q1 2020 in this state.
* As a nit, it's uninspired and ugly as sin IMHO from a design standpoint. It looks like an off-the-shelf Atlassian SPA theme. And janky. Somehow the sidebar got wider AND dropped the test explaining the icons?!
The whole thing seems like a boondoggle of epic proportions. Did I mention they seem to also be simultaneously rolling out GraphQL; because GraphQL?! I'm somewhat surprised and saddened there aren't more people pushing back on it and leaving feedback in their forums. Meanwhile we can't filter entire workflows, and we can't filter for PRs. Sigh.
I do like many aspects of their service, so it's a shame to see something play out like this :/
We made the switch, but quickly looked at other options. I can't have a vendor willingly introduce major, breaking changes.
Prior to Microsoft, it seems that the devs had free reign to develop whatever features they wanted, completely undirected by any sense of product management.
We stopped using their enterprise product (which they shipped as an obfuscated ruby blob) because they wouldn't listen to us and help us with the problems we were having.
Primary push was the Dear GitHub letter and the isaacs/GitHub repo which documented hundreds of such small issues. Also, the GitHub refined plugin, features of which GitHub keeps reimplementing now.
Also the big one for me is their Desktop Client. It has constantly improved and it makes my life easier. It has only gotten better after MS acquisition. They added rebase and merge support recently. I still do stuff from the command line (interactive rebase the most), but the Desktop Client really improves my workflow.
To me this largely looks like they're actively listening to their own developers, or have a structure in place to allow devs to take the initiative to address the needs they're facing.
But I have zero insight in how MS works internally, so I can only guess :)
Source: MS employee now works at GitHub