The eigenvector of “Why we moved from language X to language Y”
erikbern.com
erikbern.com
edit: A identity matrix wouldn't affect the stationary distribution, but if you had the actual "stay" probabilities they wouldn't all be the same, and thus not an identity matrix at all.
But as you said, the stationary distribution does change if the missing diagonal is not constant. And there is no reason to believe it is constant in the first place. In the end, what the author measures is still different from what he thinks he measures.
Some crowds (Go and nodejs enthusiasts) are notoriously vocal due to the hype.
Finally, it's common practice for companies to have their marketing department to pay "media marketing specialist" to advertise for products (their language) by posting on forums.
Sure, you can't get a lot of questions for an obscure library that on one uses, but the mere fact that tons of people post questions about it doesn't mean it's good or popular. It might mean lots of beginners want to use it because it's been marketed and hyped as easy to use. (Angular comes to mind).
This is literally in the title
>The eigenvector of "Why we moved from language X to language Y"
That there is the same absolute number of people choosing to stay within every language studied.
His analysis would only be correct if the number of people who choose to stay with Java at any moment is equal to the number of people who choose to stay with Rust. That is absurd and hence the eigenvector is meaningless.
Doesn't exactly matter much given that this is just a bit of fun on a blog post though. :P
It's not unreasonable to use that as a proxy for industry trends. I recall reading about manufacturing jobs which, sure, might have lots of factories not changing, but the ones that do change _definitely_ opt for more automation with fewer workers. That's still a trend worth thinking about.
Great point. It's interesting to think about how exactly this could bias the results -- would it bias it in favor of languages that developers tend to initially not choose for their project?
I do see that as very consistent with C, C++, and Java being up there. For new projects, developers love to choose anything but those, but then they find themselves gravitating towards them when the project gets bigger and practical concerns intrude
I think this flaw is even smaller than the issue of using Google statistics to infer transition probabilities. It's just a shitty proxy, at best.
At the end of the day, there's a lot of assumptions going into this analysis. I hope I didn't make it seem more serious than I meant it to be – it's really just a fun project and kind of a joke not to meant taken seriously.
That being said, I think the conclusions are at least "directionally" correct. They might be off by a factor of 2x or 5x or even 10x, but the stationary distribution exhibits an even bigger spread (multiple orders of magnitude) so I suspect the final ranking is still "roughly" correct (with a very liberal definition of "rough")
An interesting thing about this methodology is that it is extremely sensitive to the age of a language. It's possible to switch from an old language to a new language, but not the other way around -- so if you happen to do your measurements after a language has had some uptake but before it's been around for long enough that people have built significant projects on it and subsequently gotten sick of it, the future distribution by this method can only be 100% New Language. (Because sometimes people switch to New Language, but no one ever switches away.)
Actually, to predict the future distribution of language use, you also need to know the rate of people moving from nothing ("I just had a brilliant idea!") to each language. If everyone eventually transitions to Go, but everyone starts in Ruby, then the division of market share between Go and Ruby depends in part on how frequently people start new projects.
And of course, all things trend toward newness, so your objection there seems more about time or human psychology than the methodology of the post.
> I took the stochastic matrix sorted by the future popularity of the language (as predicted by the first eigenvector).
Emphasis in original.
I agree that the analysis makes it abundantly clear people move to older languages, but the question is what new projects are started in, and how many projects represent new versus transitioned projects.
This analysis is interesting, and gives a rough idea of what people are moving from and to when they decide to do that, but not necessarily popularity.
What the author is indexing in the end isn't really predicted overall language use, it's predicted transition target frequency.
> if you happen to do your measurements after a language has had some uptake but before it's been around for long enough that people have built significant projects on it and subsequently gotten sick of it
By this definition, it is not possible to switch from a new language to anything.
It's stated as a binary, but really this defines a continuum of newness, and the metric of the OP is very sensitive to it.
(Although I concluded the Haskell-Elixir equivalency myself based on functional semantics)
The less magic that happens and the more code that is commonly used by all developers, the easier it is for others to read your code and understand what it does. Rob Pike and others have some interesting talks and blog posts on this
That apart from the fact that it is questionable that it can be represented by an operator that is finite and linear.
It's more likely a stochastic process (infinite matrix) with births and deaths.
I would be surprised if it became true. :-)
How many projects are _started_ in a language, and how many _die_?
Given p.e. Java - It may seem that there is a huge flow to go. However, if there are enough new projects started in java, then the number of java projects might still rise faster than the number of go projects.
http://shop.oreilly.com/product/0636920037194.do
The hidden lesson for me was that rewriting the code in <new-language> is not the only option. Another option is to slowly improve the language/runtime itself until you've essentially switched it out underneath the application, which is what happened at Facebook. Meanwhile keep refactoring the code to take advantage. (Granted, this is isn't an option for a small company.)
I work at Facebook and sometimes write www code. When I was interviewing and thought about writing PHP code, I didn't get a warm and cozy feeling, being reminded of terrible PHP code I've seen (and written myself) in the 2000s as a "webdev". Thanks in part to the advances described in the book, the codebase is definitely not like that; it's easily the best large scale codebase I've ever seen (I've seen 3-4).
My thoughts in blog form (written 3 months after I joined):
I noticed it lists 14 results from Haskell to Erlang, which I was skeptical of. When I google "move from Haskell to Erlang" or "switch from Haskell to Erlang" I do find results (such as quora questions, versus questions,lecture notes) but none of those results are the type of article we're looking for.
If they really want to do this, I think they also need to validate that some of those keywords are in the title of each page.
numpy.linalg.eig(x)[0]
numpy.linalg.eigvals(x)[0]
Numerical stability can be hard to get right...The nice thing about the power method is its conceptual simplicity. In cases like this, it's quite hard to screw it up. And it will scale far beyond those numpy functions (not that this is needed for this example.)
Also, did anyone mention PageRank yet?
It's actually the second highest eigenvalue. The highest eigenvalue is always 1 for stochastic matrices.
All this being said, scalability is _obviously_ a non-issue when talking about a matrix of programming languages. All methods are constant time.
[1]: https://en.wikipedia.org/wiki/Matrix_multiplication_algorith... Interesting tidbit: nobody can even prove it's not N^2 :)
And yea, mat mul is not N^3 theoretically, but most implementations are. I've heard that some (mkl maybe) are 2.8, but haven't had someone point code to me. My personal attempts at implementing Strassen were slower than a tuned N^3 implementation, at least for matrices that fit into memory.
m /= m.sum(axis=0)[numpy.newaxis,:]
u = numpy.ones(len(items))
for i in xrange(100):
u = numpy.dot(m, u)
u /= u.sum()
Immediately concerning is the use of a dot product and a sum, which will lose a lot of information when the values being added are of different orders of magnitude. (For example `M + epsilon = M` in many floating point computations).Also, while scalability for this size computation is way overkill, it is precisely a problem where the numerical stability problem gets even worse. Imagine if I have data on some kind of power law, and compute left-to-right `M+epsilon_0+...+epsilon_n`. No matter now large `n` is, for sufficiently different order of magnitude M and epsilon all the information is lost -- `epsilon_0+...+epsilon_n+M` could be an entirely different number. Highly recommend checking out the LAPACK stability guide here http://www.netlib.org/lapack/lug/node72.html for more.
It is commonly used for web scale recommendation problems which likely explains it's usage given the author's prior background. Also underlies pagerank.
AFAIK it is quite stable but converges slowly. From my experience within 50 iterations you will have converged to a stable result.
For larger, and legacy project, maintainability and continuity is the overwhelmingly top priority, being flashy is the least of their concern.
But the codebase numbers in the 100's of millions of LOC, and replacing that with Rust (or anything else for that matter) will come with a number of requirements:
- it really needs to be better in terms of bugs
- it needs to be about as fast or faster
- it would have to come with similar start-up times for the runtime
- you'd need to find funding somewhere or convince people the need is so high they should volunteer their time
For a single individual this is likely not a viable project, even doing just the basics (coreutils) would take you a couple of years at a minimum.
What might be a better way is to reboot unix entirely starting from the kernel and working your way up into userland.
There is plenty there that could use a more modern look at things, after all UNIX really is showing its age and LINUX is a re-implementation of something that was already old when it was started.
Linux doesn't really resemble UNIX anymore. It's starting to look a little like Plan 9, to be honest... give it 20 more years.
Hm. I'm not sure I see the similarities there. plan9: small, elegant, really an improvement on Unix in many respects but unfortunately somewhat theoretical rather than practically oriented.
Linux: bloated, blunt, practically oriented, gets the job done but it never feels like it's the shortest path.
POSIX is defined in terms of C semantics, which includes C unsafety, like managing pointers and the respective length as separate entities, using null terminated strings or casting void* to specific data structures.
Which means the POSIX translation layer would need some sprinkles of unsafe code to be able to comply with the required semantics, thus opening it to the same exploits as C code.
This issue is visible in OS that aren't written in C, but do expose POSIX runtime layers like mainframes.
As mainframes, they usually restrict possible security exploits thanks to running POSIX applications on their own containers or enclaves.
If I remember right, the Fortran one was defined in terms of the C one, but the Ada one was written as if it was an independent specification.
I suppose this was only possible because POSIX misses out a lot of the fiddlier bits of Unix anyway (and the Ada one specified a spawn to use instead of fork+exec).
Let's say how does one make memcpy() safe in Ada? The very first step of converting the pointers + length into an access type must trust the caller did the right thing.
So all the unsafe bits can exist only in the process which is based on them.
You don't need to rewrite, just stop writing extra stuff in C, maybe when you do a really big refactor in C, port it. In the end, we can migrate away from C gradually.
That makes me wonder, is there any chance in hell we get some RUST in the Linux source code?
If you want to fix UNIX security issues related to C, first you need to convince kernel developers across all UNIX variants to move away from C.
Then even a fully modernized UNIX kernel needs to provide support to unsafe POSIX APIs.
But, Linus et al still love C...
Write your own out-of-tree Rust modules if you like, but they'll never be merged in.
Not all UNIXs are closed source.
I think rust opens the door to designing kernels that are small and tight with most of the OS stuff that was traditionally in kernel space moved into user space.
Perhaps someone could start with Minix3 or Darwin rather than Linux or BSD or from scratch. Replace one component at a time in Rust, or D, or Ada...
I'm not sure I get your analogy.
Besides, Real Programmers™ eat mutable state and NULL for breakfast.
source: http://maxedv.com/wp-content/uploads/2011/12/Sourcefire-25-Y...
pwillis@windows:~/Downloads$ wget -q -O allitems.csv https://cve.mitre.org/data/downloads/allitems.csv
pwillis@windows:~/Downloads$ ( for year in `seq 1999 2016` ; do TOTAL=`cat allitems.csv | grep -v RESERVED | grep "^CVE-$year" | wc -l`; BUFF=`cat allitems.csv | grep -v RESERVED | grep "^CVE-$year" | grep -i -e "buffer.*overflow\|overflow.*buffer" | wc -l`; PERCENT=`awk "BEGIN{print $BUFF/$TOTAL*100}" | cut -d. -f1`; echo "Year $year: $TOTAL CVEs, $BUFF buffer overflow related, $PERCENT% total" ; done )
Year 1999: 1573 CVEs, 307 buffer overflow related, 19% total
Year 2000: 1237 CVEs, 250 buffer overflow related, 20% total
Year 2001: 1540 CVEs, 278 buffer overflow related, 18% total
Year 2002: 2370 CVEs, 481 buffer overflow related, 20% total
Year 2003: 1519 CVEs, 346 buffer overflow related, 22% total
Year 2004: 2670 CVEs, 470 buffer overflow related, 17% total
Year 2005: 4686 CVEs, 519 buffer overflow related, 11% total
Year 2006: 7047 CVEs, 612 buffer overflow related, 8% total
Year 2007: 6510 CVEs, 868 buffer overflow related, 13% total
Year 2008: 7034 CVEs, 619 buffer overflow related, 8% total
Year 2009: 4888 CVEs, 589 buffer overflow related, 12% total
Year 2010: 4954 CVEs, 419 buffer overflow related, 8% total
Year 2011: 4441 CVEs, 428 buffer overflow related, 9% total
Year 2012: 5219 CVEs, 399 buffer overflow related, 7% total
Year 2013: 5731 CVEs, 382 buffer overflow related, 6% total
Year 2014: 7494 CVEs, 326 buffer overflow related, 4% total
Year 2015: 6526 CVEs, 357 buffer overflow related, 5% total
Year 2016: 7180 CVEs, 471 buffer overflow related, 6% total
According to this really shitty review of CVEs, buffer overflow is less than 10% (recently less than 6%) of tracked vulnerabilities. That's still a lot, of course.The right thing would be to weight by severity and number of users impacted, and bundle all the memory-safety vulnerabilities together (i.e. buffer overflow, double free, use-after-free, aliasing violation). I'll add it to the big list of blog posts I want to write.
C++ comes to mind for writing browser engines, but those have their fair share (if not more) of security issues as well.
I don't know why so much Internet infrastructure managed to miss that shift. Possibly a case of "ain't broke, don't fix it" - Apache, OpenSSL and what have you haven't changed that much since the early '00s and few people are motivated to write a replacement. Databases mostly existed since then, and newer datastores do tend to use better languages.
Note the types of vulnerabilities and the languages they use. DoS, File Inclusion, XSS, Exec Code, Dir Traversal, Priv Escalation, SQLI, Bypass, CSRF, Info leak, etc.
Many of them (including ones in C) have nothing to do with memory protection. Those that do (null pointer deref, use-after-free, memory leak, buffer overflow, etc) are all trivially protected with small kernel patches that have been around for 17 years, and of course most of these vulns would become trivial with proper mandatory access control. But for reasons that completely escape me, nobody has adopted these basic techniques to prevent small bugs from becoming big holes.
Now balance the common holes in C against all the other bugs in higher level languages with otherwise suitable memory protection, and consider that at least with C there are basic steps that prevent many of these from becoming problems, whereas with other languages you need a hodge-podge of different, more complicated countermeasures. C is actually easier to secure because its bugs are common and not difficult to catch by the kernel.
Honestly, if people spent as much time developing new industry best practices for use of the language as they do complaining about it, this would be a non-issue. But C isn't trendy, so let's all crap on it and pretend it's the only issue so we can have fun reinventing classical bugs with new languages.
I think this is completely wrong. In other languages you just don't have these problems in the first place because you are memory-safe by default. It's not like there's some point-buy system where all languages have to have the same number of opportunities for bugs and not having stupid memory safety bugs means you have more complex subtle bugs instead. You just eliminate a huge proportion of your bugs.
You make the serious mistake of assuming all or the majority of C devs know how powerful the language is, or how to properly use a language like C. All software has bugs, true. It's just that C bugs tend to be a wee bit more dangerous than a lot of other bugs because of the raw power of low level languages like C.
Writing secure code in low level languages is tough, and it requires knowledge of all the pitfalls and nasty corners of the language, which many devs don't have the time, interest or desire to learn.
So maybe that's a flaw in the whole exercise: it's only taking into account remarkable/unusual language switches, the ones people think are worth writing articles about.
1. Movement from Postgres to MySQL
2. Movement from Mariadb to MySQL (and NOT the other way around?!?)
3. Movement from PHP to Java (I remember the sort of people leaving Java for PHP 10 years ago, and I don't think they'd go back, or that PHP people would pick Java as their choice to move to)
I think maybe he has the axes labeled wrong?
And Java is seeing a bit of a resurgence as well as people get fed up with shitty PHP and other dynamically typed languages. Java has some frameworks like Dropwizard and Spring Boot that make it not as terrible anymore.
That wouldn't explain why people jump from the database with better ACID (Postgres) to the one with generally worse ACID (MySQL). Or why people would move from the open source non-Oracle fork (MariaDB) to the Oracle-acquired original project that everyone forked away from (MySQL).
> And Java...
Oh, I agree. Java is awesome these days. I'm just making a disparaging blanket generalization about the people who jumped to shitty PHP to begin with.
If you search "move from Objective-c to Swift", there are actually enough results that it honors the literal string, yielding far fewer results.
Definitely a major methodology error for pairs with a large asymmetry.
Nonetheless I found the post humorous, and I don't think it was held with the conviction some of the top posts seem to think.
objective c to swift: 5216 swift to objective c: 1639
(sorry about the small font size though. had to squint really hard)
The future probability table shows that Swift is to the left of (smaller future probability) of Objective C. The coloring of the chart shows a much darker square for Swift -> Objective C than for Objective C -> Swift.
It seems surprising given your contingency table that your analysis would show that Objective C is going to be the more popular of the two.
Top 5 giving to Go directly: C, Python, Java, Ruby, Scala
Top 5 giving to C: C#, R, Java', C++, Fortran
Top 5 giving to Python: C', Perl, Java', C#, C++
Top 5 giving to Java: C', C++, PHP, Python', C#
Top 5 giving to Ruby: Python', PHP, Perl, Java', (C++ — only 215)
Top 5 giving to Scala: Java', Ruby, (Python', C#, PHP — only 100, 17, 16)
': language also in top 5 givers to Go.
The other top languages take from each other (there is migration in both directions), but currently Go mostly takes here. However it does lose people to Rust — which is actually the strongest go-to language from Go. And C++ does not give to go.
This might point to a discrepancy between Go marketing and reality (efficiency and replacing C++).
It would be great if you could repeat this exercise next year to see how things changed.
(besides: the script is nice and concise!)
https://commandcenter.blogspot.com/2012/06/less-is-exponenti...
Then reading more I realized the results included things like optimizing certain portions of programs into a language for hot areas of code. Though the Rust to COBOL one I should go read. That's nuts.
They are always just "We wanted to do this in a particular way. So we fought the framework till we decided to move to another that does things the way we thought they should be done. Now things are much better but we will fail to mention down the line all the new compromises we have to deal with"
Unfortunately, you'd probably be hard-pressed to find a decent methodology.
But indeed, "we used this language for 10 years without thinking about switching" is a much more interesting metric than "we switched languages 3 months ago".
The only problem is that the labour market is a lagging indicator. Perhaps the best approach is to combine statistics about active switchers (like the OP) with labour market statistics.
Node as an example has been trending in my area for long enough for there to be job listings. Yet if I look at the job listings 90% of them are for JAVA, C# and .NET or PHP. In fact in a region of 1,5 million people where node has been trending for a while, there are 0 Node related job openings this morning.
I don't see it changing until customers request any other kind of deliverables.
Oh and we have zero projects in any sort of public repos.
When it comes down to it, all languages are DSLs. Even LISP/Scheme are DSLs for making DSLs (like Butterick's "Beautiful Racket" earlier today).
Presenting them as per-niche directed graphs is probably less likely to steer newbies (and sadly, not-so-newbies) into another round of "let's redo everything in X!"
SQL Server, Oracle, DB2, Informix.
People really write about the databases you listed, because their users have an entirely different mindset. Sure people may switch from Oracle to DB2, or from SQL Server to Oracle, but some organisations just have "standard databases" that they work with. Switch would be a multi year process, and certainly not something to be advertised, unless it's: "Now with support for SQL Server" in the marketing material.
My employer does enterprise consulting.
grep go | grep -v ago
It's funny, there was this company Yahoo! that was trying to organize the internet to try and fix this...
Edit: Oops, missed the same reply from xudongz
There is nothing that comes close that can compete with the massive success that Smalltalk has been. Its blows my mind how it can be so much better than anything else out there including the usual suspects (Lisp, haskell, blah blah).
But in the end its not about the language , its about the libraries. Hence why Python remains my No1 choice.
In the end however even Smalltalk is terrible outdated. The state of software is in abysmal condition trapped in its own futile efforts of maintaining backward compatibility, KISS and do not reinvent the wheel.
In sort software is doing its best to keep innovation at a minimum and as such pretty much everything sucks big time and is still stuck in stone age.
I once considered becoming a professional coder working in a company doing the usual thing, I am glad I was wise enough not to choose that path. I would have killed myself right now with all this nonsense that makes zero logical sense.
But my hope is in AI, the sooner we get rid of coders, the better. Fingers crossed that is sooner than later. Bring on our robotic overlords.
Saying that I know a lot of people that really love coding and respect it as an art and science, so there is definitely hope.
* I find the template syntax is more sensible than react.
* It's not as total as angular. You can use just the small parts you'd like.
Angular isn't going to help you write your server. That's such a strange statement.
I wonder why that is?
Generally when a language/framework/toolset first hits, it looks magical and fixes loads of problems people are currently experiencing.
It's own crop of problems has yet to emerge (generally these only emerge once a sufficiently large number of projects have been using it for a sufficient length of time).
So at the moment Go is the new thing, and it's surplanting the older new things... come back in 3-4 years and it'll likely be on the losing end to something else.
I'm not dismissing it. Just wanting to define proper context. Else some twithole will start a shit storm over null.
You mean ... Google search results?
I'm not trying to suggest that Google's search engine is intentionally biased toward Google projects, but I think it's reasonable to assume that their own projects wouldn't fall into whatever unintentional blind spots their search engine may have.
You may be reading the axes the wrong way around... look down the node column rather than the node row.
But alas, this post had to use an image.