What's Going on with Language Rankings?
redmonk.com
redmonk.com
Not only that, they are impacting the actual languages people use. Why use some new / esoteric language that an LLM doesn't know much about, when you can get the same job done much faster using a language that the LLM knows well and can debug?
It's going to get increasingly difficult to bootstrap new languages unless they are available in LLMs.
Rust is wildly easy to generate (idiomatically so).
I've been building in our Rust monorepo without LLM tools for years and just tried them out yesterday -- I was utterly blown away.
Rust might become more popular with LLM accessibility.
do you need help? are you getting paid?
Would you have replied the same way had they mentioned the ability to use LLMs to generate Python? Would you equally imply an accusation of a Python developer of being a shill for mentioning it?
Also the LLM won’t have been trained on the vast corpus of debugging sessions, blog posts, etc of existing languages. I’d expect the quality to be very poor compared to the current top players.
(It’s effectively the same phenomenon as folks sticking to languages with a big hiring pool.)
It really illustrates (to me) that LLMs are "Plausibility Machines". They produce plausible answers. Plausibility is informed by repeated observability over large sets. We asses something as plausible when it matches the majority of our observations. That's how the parlor trick of seeming to "tell the truth" is pulled off. The venn overlap of "plausible" and "true" is pretty high. But when "truth" is in flux, you start to see the feet of the wizard behind the curtain.
It was definitely pretty bad with liveview for a while, while that was under heavy development and API changes, but it's had a data update that mostly fixes that, and if you specify the version it's helpful there too.
I’d say it’s still uniquely good with Python, but I find it roughly as useful with Elixir as it is with some other very popular languages.
It makes sense in theory, but this is the type of thing I want to see data on before I consider it a possible problem, even.
No sense in worrying about things that you can't control, can't prepare for and have not occurred yet, after all. :)
Right now, today, when choosing a library, top of my mind is "has this been around long enough to be in GPT4's training set?"
Because to give up the 5-10x productivity boost GPT allows for, a library better be really good.
I've seen a couple companies offer a RAG enabled ChatGPT for their SDK, but without a large number of stack overflow and medium posts detailing how to use the SDK in different scenarios, well, there isn't much help an LLM can offer that I cannot just get from whatever is on a company's site.
Back in the day Microsoft used to pay boat loads of money to have absurdly complex sample applications shipped alongside new libraries and SDKs. The sample apps were often fully featured end to end solutions that had been gone over by testers as well.
The idea being, since there wasn't any really good online help back then (aside from news groups and such, which were useful if you knew about them), you had to rely on sample code or on books, books which Microsoft also commissioned to be written.
Maybe someday we'll get back to that point, companies paying money for large volumes of sample code to be written, but this time with the goal of having it ingested by LLMs.
I don't see how this is anything new. Similar tension has always existed slowing adoption of new technology. I've resisted picking up new languages because I could implement something faster/better in languages I already know. I've been skeptical about some because the ecosystem just isn't there yet. You said "a language that the LLM knows well and can debug" when sometimes the resistance is that the new language does have as good of tooling/debugging support.
There is nothing new here and if anything LLMs could help. I've noticed the community seems pretty understanding/accepting of low quality outputs of some of these LLMs. The small corpus of a new/esoteric would mean even lower quality, but I don't think that'd be a deal breaker for some folks.
I haven't really used copilot with, say, Rust or Haskell, but if the average quality of code is better in those languages, maybe it will take less coaxing to actually suggest good code?
Either way, LLMs are great at generalizing between human languages, so I expect they transfer a lot between programming languages too. I think your language has to be pretty weird for copilot to not get a grip on it with enough context.
If you provide high quality context, it's more likely to produce high quality code too - but it can also spiral into creating bad code if it accidentally starts off on the wrong foot. The more bad code you have, the more bad code it suggests. That's why I think it's very important to clean up the code it suggests, even if it's just cosmetically.
There are lots of answers that are flat out wrong 5+ years after they were answered. There is no way to go back and "unmark" that answer.
> Modern C++ is really a different language and yet there is a lot of not modern code out there that AIs don't know is bad.
This is a human problem and why I won't use C++ anymore. Instead of "modern C++" we needed a new language to replace C++ and leave the old nastiness behind. We'll see if any of the new replacements like Carbon get any traction.
Hmm, I can see it going the opposite way - maybe the availability of LLMs means that having an active user community in your timezone is less important, and makes semi-obscure languages more competitive.
For me, GPT4 means Rust is approachable, something I would never have considered a couple years ago. There’s tons of documentation online, but it’s great to be able to get the “vocabulary” of my question from GPT4 before looking it up online. Rust compiler errors are pretty good, but a LLM really smooths over the rough spots.
Where a new language will be severely disadvantaged is in the blank page problem. With a well-known language you can start a new project with just a comment and let Copilot get you over the initial barrier of just having something down on the page. A new language won't work because it needs a fair amount of context before it's able to produce correct code.
If people doggedly continue down that path they're going to stagnate for as long as they assume the AI is always right.
Just because an LLM can can cobble together some output doesn't mean that the output is instantly trustworthy.
Not to mention languages often optimise for different things. If you want to default to JS for everything because the AI can auto-complete that more easily, you're locking into a very specific operating model.
Or maybe you'll teach your new language to the LLM and it will make porting and binding to everything else easy.
It will take 10 to 20 years before there is "Star Trek computer"-level of self-programming, agent-based systems.
I’m a software developer yet still amazed at how many layers of abstraction are tolerated in user space.
Take web development. You’d think at some point the stack of a half dozen frameworks would be abandoned in favor of vanilla JavaScript. But the house of cards continues to grow.
So maybe languages will still stick around when they are no longer the human interface.
Anyone who's ever tried to build anything big with vanilla JS knows why these frameworks exist in the first place. The ones like Vue or Svelte aren't even really that complex to grok anyways, despite all the hyperbole about them.
But hopefully my point still stands. A coder writes instructions for the computer; it amazes me the tolerance for having those instructions go through multiple layers of transformations (sometimes even multiple layers of languages) before hitting the OS.
If the point of a scripting language is to provide a user experience not attainable by a lower-level compiled language (which in turn is not attainable by machine code), wouldn't AI nullify the benefits of both the scripting and the lower-level languages?
Though my intuition tells me that the net effect of layoffs is still negative.
May be obvious, but would rather know if you saw PRs-per-repo decline. Otherwise, I wonder if some very frequent-PR repos (automated or otherwise) had previously skewed numbers. Also curious if early numbers for this half can be obtained/compared.
They need another datasource. There's another classic not reliable system, survey programmers to see what they are using.
Oh, this made me connect some dots. At first I figured increase in GitHub activity during the pandemic was because people were home more and had more time to contribute to open source.
But a lot of people took that time to learn software development. And a lot of courses will have you publish your learning projects to GitHub as a way to learn version control and processes (and also to promulgate awareness of, and analyze their course since there's often templates involved).
Often this is just creating repos and pushing them, and since TFA directly looks at pull requests, perhaps this isn't it, unless these courses regularly include lessons on how to do a pull request (perhaps from yourself to your own repo).
If that was a significant factor, it would make sense why the activity is down.
> They need another datasource.
Agreed here, the source is very imperfect.
Don't assume this "oversight" is an accident or result of stupidity. In one case it was absolutely used to justify a promotion of a yes man.
https://redmonk.com/sogrady/2023/05/16/language-rankings-1-2...
If Stack Overflow becomes more and more a reflection of "historical behavior" rather than "modern behavior", this sort of metric will become a less useful way to judge modern developer usage.
That doesn't strike me as reasonable analysis.
I'm not sure I understand. Granted, they're never going to get perfect data about language rankings. But, why would AI assistants affect the rankings of language in a way that makes their methodology more unsound? They mention a drop-off in SO tag usage, and in Github pull requests: that's an indication of something, but is there any reason to believe it distorts the results so as to give undue prominence to one language more than another?
Intuitively, I can't think of how. "Undue" is the keyword: even if the prominence changed as a result of AI language assistants, wouldn't that be an indication that people are just using those languages more, with the help of AI assistants?
I'm really just wondering why, given the unavoidable squishiness of the methodology, the decrease in magnitude would bother them so much. They probably don't see their work as fundamentally hand-wavey, and maybe I just do, and that's the problem. I guess they're just investigating it, but the results seem to be more or less what you'd expect to conclude ab initio. I could be wrong, often am.
ChatGPT can only provide answers for programming languages and libraries that were included in its training. Granted, unpopular programming languages have online activity, too, but sample size matters.
I don't think software development has started slowing down. So the data is indicating that the way they were measuring is becoming less useful than it was, and that they will need to find new data sources if they are going to continue.
If that really is AI, maybe Copilot and friends will become an interesting source of data on programming language use.
[1] "RedMonk Programming Language Rankings, 2021-2023", https://i.imgur.com/MWhzGC2.jpg.
Individuals since the colors get messy.
Q1, 2023: https://redmonk.com/sogrady/files/2023/05/lang.rank_.q123.wm...
Q3, 2022: https://redmonk.com/sogrady/files/2022/10/lang.rank_.622.png
Q1, 2022: https://redmonk.com/sogrady/files/2022/03/lang-rank-0122-wm....
Q3, 2021: https://redmonk.com/sogrady/files/2021/08/lang.rank_.0621.pn...
Q1, 2021: https://redmonk.com/sogrady/files/2021/03/lang.rank_.0121.wm...
I was making decent progress on Rust and enjoying that, but have reverted back to more python since that is the lingua franca of AI.
So how do we rank that? Rust because of technical merit or python because real world? That's a distinct you won't catch via stack overflow scraping
TIOBE
> Does anyone honestly believe that TypeScript is half as popular as Prolog or D?
I think this is one of the strongest signals that TypeScript is the dominant way JavaScript is written:
https://npmtrends.com/@types/react-vs-react
The type definitions for React are more popular than React itself. Why is that? Even users who aren't choosing to use using TypeScript are benefiting from their IDE installing typings on their behalf, at a rate large enough to exceed the CI/CD systems and users running "npm/yarn/pnpm install" and only installing "react".
It's like this for other packages as well; but the sheer popularity of react makes the point well. Many, many developers are using TypeScript even if indirectly, it's what makes their IDE light up.
The primary "problem" programming language ranking lists solves are: 0. the amelioration of trepidation of the novice as to which tool in the toolbox to "select" without first building something by trial-and-error. 1. The businesses who don't want to know anything but wish to surf fashionable trends in business and technology.
TS is a strict superset of JS. It probably doesn't get as much credit for being safer and superior to JS because it's routinely transpiled out of existence and doesn't offer language features on its own besides gradual, static typing. And there aren't many questions about it because its fairly self-explanatory. The biggest problem with TS is the Cambrian explosion of confusion of procedural configuration files (.ts and .js).
Prolog is used in programming language courses to solve logic brain teasers, formal verification every now and then (although Coq, Isabelle, RedPRL, and Idris exist now) and some niches like TerminusDB and Watson. It was probably used once to solve a formal proof and then modeled by Erlang. It's one of those influential languages amongst language designers like ALGOL 68, Haskell, and LISP that weren't themselves popular. (You can thank ALGOL 68 for the Bourne shell family.)
The other thing is programming language popularity peeing contests are rarely useful because what is practical is productivity and correctness at scale under the financial and effort constraints in reality for a given problem, team, and codebase. Plus, thundering herds of people just in it for the money tend to have resumes that all look the same and completely lack curiosity. Also, recruiters who follow popularity are the ones I can righteously tell to go hell for insulting me by saying things like "X isn't going anywhere" (In one case: X was Go, and t was 2011). While there is "safety in numbers", there are N ways to trim hairy cattle. Sometimes, easier and better ways come along that need to be tried. Curiosity and learning never go out of style except at shops that are dinosaurs marching to the tar pits. Shops that discourage learning in favor of ship-ship-ship yesterday, avoid bettering their employees, and are full of people who have no technical side-hobbies lack engineering craftspeople.
rWars/lWars were sprileg verboten in earlier versions of The Matrix. Now, with all the resources around to waste (WASTE!), such discord can be used to form (or manipulate) popular opinion and herd competitors towards doom (“and then you paid your people double cuz they did it in Rails”)
If you don’t know what languages to use for your apps without reading a popularity survey, an exciting career awaits you in the field of marketing, customer relations, finance or golf caddy. https://youtu.be/za0nyYbp6is
I think Chatgpt/Copilot merely accelerated something that was already known to be in a bad state due to the toxicity of the platform.
If entrants are not artificially inflating "organic" signals via fake content spam (Twitter/X), then the criteria themselves are losing their signal strength (StackOverflow/GitHub).
The diffusion makes it increasingly difficult to understand which channels are important and which correlate to strength in the market.
Unfortunately, these can be more than vanity metrics.
Some VCs or financial markets may use these as methods towards valuation.
I only see js/ts on a huge incline mostly because of backend, and making it easier to hire full stack devs (desirable in slow economies)