GitHub-Next
githubnext.com
githubnext.com
Is it a conference? A team? An initiative? How does it work?
Whatever this is supposed to be selling me on, they're doing a terrible job because I can't figure out what it is!
Their team consists of engineers and researchers.
I'm seeing a pretty good description:
" GitHub Next investigates the future of software development. We explore things beyond the adjacent possible. Tools and technologies that will change our craft. New approaches to building healthy, productive software engineering teams. "
Or maybe they just updated the site in the last hour. It's got a list of team members and events and stuff.
I've read the whole page and I'm not sure what it is. I suspect it's a think tank, but I'm not certain.
It's an old trick first called out in The Dilbert Principle that you get promoted by being associated with sexy sounding projects, the best case is a sexy sounding project that has vague objectives.
Given the blurb above it sounds like there is very little substance to this and lots of style. Expect it to be sunset or fade into obscurity in 3-5 years.
EDIT: For context their stated mission is > GitHub Next investigates the future of software development. We explore things beyond the adjacent possible. Tools and technologies that will change our craft. New approaches to building healthy, productive software engineering teams.
Notice the lack of anything that will actually be produced, key words that you are dealing with BS is "explore", "approaches", and "change our craft" without any details on what any of that means. If they were producing or doing something they'd say that. As is it's meaningless marketing drivel that could translate better as "We've made up a fake job to play with cool tech, write ill-informed meaningless blog posts about what the next cool thing everyone should do is. We're also light on technical skills so we are going to focus on teams and projects not actual tech"
I don't begrudge them the fact they weaseled their way into this job good for them, but I don't expect much, and If I am wrong and these are actually super smart talented individuals that finally got the freedom to make something happen then I apologize for my dismissiveness. That being said when companies get acquired by MS they don't tend to be known for their innovation and ground breaking research going forward.
This is slightly worrying if this is the direction Github are heading.
I want to know all the callers and callees of every function. This shouldn't be too hard, we already have find references via LSP.
Turning this into a graph would make it significantly easier to manage the entry and exit points of a code base and inform architecture decisions, refactors, type checking, hot paths etc.
The AST module was super handy as you'd expect. The script would optionally take some filters to reduce the size of the generated graph, and then it sent all the info to Graphiz (it emitted DOT, too, so it could be version controlled!!)
It was extremely fun, highly recommended.
Yesterday I tried to use Datadog's Github integration for stacktraces and it asked me for "access Github on my behalf".
It's been the same since the beginning of Github - they leave integrators with no better options, and users with an ambiguous UI dialog / docs that downplay the scope being granted.
Sooooo maybe fix your own stuff before making such grandiose claims?
They also provide a direct git integration, which as far as I can tell just is a reduced version of the GitHub one, with a featureset that seems reasonable if they only have the pure git data.
Microsoft has been super litigious in the past when it came to copyright violation starting all the way back with Bill Gates' letter in Byte magazine about those pesky pirates. To see them do this makes pirating MS software fair game from here on. They could have asked nicely, instead they just took.
It is. Laws are adapted based on widespread technological capabilities and progress.
As an example, if it is easy to create real voice or signature using AI models - they should no longer be considered effective evidence for contractual reason instead of enforcing that it is illegal to forge it. That is not going to work.
Past shouldn't dictate what we allow tomorrow.
All laws are made in interest of someone.
Does the copyright apply to AI models since they are out of scope and weren't widespread when it came into force?
Does the proposed benefit in the original law apply in practice?
Are they more beneficial than the progress allowed by AI models who use them as training data?
Is the copyright law practically enforceable on output generated by AI models?
https://en.wikipedia.org/wiki/Berne_Convention
Copyright is what FOSS depends on. For Microsoft to shit all over GitHub contributors rights is despicable.
Does the same apply to a human? Do we now define copyright violation differently for computers? I don‘t know the perfect answer here. But I‘m not so sure we should have standards that change depending on if a program is doing it or a human is doing it. Perhaps a bad standard to begin with.
I do tend to learn towards thinking „company uses publicly available, open source code in product“ is somewhat of a nothing-burger though.
EDIT: After more reading on the subject, I'm willing to accept that copyright infringement is unlikely here. This link [1] was the one I found most convincing.
However, I would still shift the goalposts and look at this ethically, and I still think it's wrong that Microsoft is profiting from code with licences like GPLv3. This is a whole other topic, though.
[1] https://www.technollama.co.uk/is-githubs-copilot-potentially...
Copilot's API is surfacing snippets of work without licensing information attached alongside. It can be shown in discovery that Copilot does access the origin work.
The sooner this is slapped down, the sooner we can avoid addressing the even more troubling question that exists today: is someone who used Copilot to throw together a bunch of code infringing copyright of works where those portions originate?
This is a complex problem with no satisfying conclusions... how could one be violating copyright if they never accessed the 'copied' work to copy? Copyrights aren't patents. Infringement requires copying.
Using Copilot launders the user's awareness of the origin works, yet making the Copilot users liable for widespread "accidental" copying would be troubling.
That you got the code from and entity that stole it somewhere else doesn't really matter. Generative models should respect copyright for their sources, and using a generative model to create new works that you intend to claim copyright on is stupid: someone may well show up one day with ironclad proof that you used their code without permission.
Also a reminder that outside the copilot debate, the online rights movement has largely been pushing for scraping, deep linking and transforming scrapped data to not be considered copyright infringement, regardless of any TOS on the site being scraped.
To me, co pilot is a exactly that, a scraper that has scraped public websites and is now presenting me the scraped data in an alternative and often transformed form. It’s my responsibility as a developer to ensure that my released product complies with applicable copyright law, but copilot and the use thereof is not in and of itself copyright infringement.
That a tool can be used to create infringing work or infringe on copyright in general is no more a valid argument against co pilot than it is against CD burners, de-drm tools, vcrs, kodi or plex, scanners or any number of day to day items that have the ability to infringe copyright if the user uses it for that purpose.
https://decoded.legal/blog/2021/06/github-copilot-initial-th...
https://fossa.com/blog/analyzing-legal-implications-github-c...
https://felixreda.eu/2021/07/github-copilot-is-not-infringin...
Until then all you have is opinions, mine is pretty straightforward: if the generative model can be made to work without first training it on other people's code then it isn't copyright infringement, if not then it is transforming one set of works into another.
The only thing that might let GitHub off the hook is their terms of service, but that might mean mass exodus from GitHub because if they interpret you using GitHub to host your code as a blanket permission to do with that code whatever they want then that's clearly not the original intent of the service.
If Microsoft buying GitHub claims that gave them a blanket license to do as they please with the contributions of millions of FOSS contributors then they are still just as bad as they were in the past.
Almost every GitHub repository comes with a license file, even GitHub should have to abide by that license, otherwise the whole thing is pointless.
If Microsoft/GitHub want to field the argument that they own the rights to all of the code uploaded to GitHub then I'm perfectly fine with that, the only problem I see with that defense is that it will likely kill GitHub overnight.
As for the jury argument: that's fine, but juries aren't lawyers either. I'm not sure if that should weigh as a positive or a negative for Microsoft.
Finally, regardless of the legality: there is such a thing as ethics and in my book you don't appropriate a large body of work from a whole community without so much as a by-your-leave. There have been other threads on HN regarding this and it is interesting to see the various opinions, even so if Copilot is challenged legally than I'll be cheering on the party bringing the suit.
Time shifting: recording TV on VHS to play later.
> GitHub Copilot is trained on billions of lines of public code.
> In one instance, GitHub Copilot suggested starting an empty file with something it had even seen more than a whopping 700,000 different times during training–that was the GNU General Public License.
https://github.blog/2021-06-30-github-copilot-research-recit...
This indicates that they are training it on github's public repositories and at the very least including 700,000 GPL licensed projects or code files. Since the GPL is one of the most "restrictive" open source licenses one can assume they are not caring about the licenses much.
I've heard people saying this since the mid-70's.
I lump it into the same trash bin with flying cars and orbiting space hotels, and "90 minutes from New York to Paris — undersea by rail." Things envisioned by artists that will never happen in my lifetime, or yours.
There is no precedent on if a computer reading your code or looking at the image is fair-use or not.
The code being covered by copyright and the code being publicly accessible are two different things.
It was easier, at least - but probably gives some nice guarantees about the statistics of "public" code, with the norms and conventions you're "used to", because you're used to the internet's coding norms.
Company code can be pretty bad, even at Microsoft, often riddled with hyper-verbose variable names and strange design patterns.
1. Code editor is full of telemetry and costs money, only accessible over the internet.
2. Code regressing into lots of boilerplate and automatically copied stack overflow answers by copilot, programmers using less critical thinking skills.
I don't trust Microsoft either.
By enabling increased complexity (via Language Server Protocol, Copilot, and even GitHub itself) devs get locked-in to the MS ecosystem. It reminds me of Braess's paradox ("adding one or more roads to a road network can slow down overall traffic flow through it" https://en.wikipedia.org/wiki/Braess%27s_paradox ).
Increasing our ability to generate (but not comprehend) complex systems is also intrinsically dangerous (beyond the "rent seeking" of MS) because complexity itself is a kind of cost or overhead. This is not to say that the more complex system cannot result in efficiency gains that outweigh the cost to maintain that complexity. (If that were true there would be no multicellular life, eh?) It means complexity should be carefully justified in terms of economic/engineering considerations.
The dependency between the modules seems like a nice addition to me. I don't think CodeScene has that one. Can't wait to try this on our bigger projects.
I never found a really good way to visualize large codebases and the dependencies between the modules, does somebody have something for this?
> We moved to https://github.com/githubnext!
https://github.com/githubocto/flat still using old name
Why not deploy as next.github.com subdomain?
Also, various boring realities around SSL termination made deployment difficult in a github.com domain. This was the expedient solution. Not phishing!
If you’re interested, check out https://codeapprove.com
Heads up, that page has no `<title>` tag so the browser tab is `githubnext.com/`. That is a _VERY_ minor nit, but still an SEO ding (you're github, that doesn't matter much), and a rough edge that could be buffed out.
Bonus points for adding a favicon too. :)
Well, the title is precisely the only mandatory element in a valid HTML5 document, even if forgetting it seems harmless.
Edit with details from https://www.w3.org/TR/2014/REC-html5-20141028/document-metad...
> Note: The title element is a required child in most situations, but when a higher-level protocol provides title information, e.g. in the Subject line of an e-mail when HTML is used as an e-mail authoring format, the title element can be omitted.
From https://www.w3.org/TR/2014/REC-html5-20141028/document-metad...
> If it's reasonable for the Document to have no title, then the title element is probably not required. See the head element's content model for a description of when the element is required.
So strictly speaking, if it's meant to be used as a traditional web page, you should really have it (obviously), but it's not strictly required.
> If the document is an iframe srcdoc document or if title information is available from a higher-level protocol: Zero or more elements of metadata content, of which no more than one is a title element […].
> Otherwise: One or more elements of metadata content, of which exactly one is a title element […].
So it is required, not just suggested, for a web page, but not for all kinds of html documents; TIL. The parser still tries to parse head contents before body contents even if you omit the head tags, so a doctype followed by title is the shortest valid full page.
I didn't mention the doctype because I believe it isn't strictly speaking an element, just a preamble, but you're right, it's required as well.
Funny thing, when I read “the only mandatory element in a valid HTML5 document”, I interpreted “element” in its generic English sense (piece, thing) rather than its HTML sense (node of type element, as distinct from text/comment/doctype/other-types-only-found-in-XML-syntax nodes).
As far as sources are concerned, the HTML spec is maintained by WHATWG, not W3C. The relevant citations start at https://html.spec.whatwg.org/multipage/semantics.html#docume....
The normative reference on the necessity of <title> is in the content model for the head element:
> If the document is an `iframe srcdoc` document or if title information is available from a higher-level protocol: Zero or more elements of metadata content, of which no more than one is a `title` element and no more than one is a `base` element.
> Otherwise: One or more elements of metadata content, of which exactly one is a `title` element and no more than one is a `base` element.
For the rest, you’re correct: the DOCTYPE is the only always-mandatory thing in a valid HTML document.
That’s like calling a cheese sandwich without any cheese a cheese sandwich.
(This is different from XML syntax, where omitting the start and end tags means omitting the element as a whole so that there will be no head tag in the parsed result.)
Directly from the specification:
> 1.2 Is this HTML5?
> In short: Yes.
— https://html.spec.whatwg.org/multipage/introduction.html#is-...?
$ dig githubnext.com
; <<>> DiG 9.11.3-1ubuntu1.17-Ubuntu <<>> githubnext.com
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 29398
;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 65494
;; QUESTION SECTION:
;githubnext.com. IN A
;; Query time: 1458 msec
;; SERVER: 127.0.0.53#53(127.0.0.53)
;; WHEN: Fri Aug 26 17:20:13 CEST 2022
;; MSG SIZE rcvd: 43Glad to help