Most used words in programming languages
anvaka.github.io
anvaka.github.io
if (err !== null) {
return callback(err);
}That's an odd thing to complain about.
if theErrorReturnedByThePreviousFunction != nil {
panic(theErrorReturnedByThePreviousFunction)
}(I am only half trolling).
You could come up with `Either` types, but due to the lack of generic structs you'd have to have a whole lot of those.
Interestingly, Rust has basically the same error handling idiom, but verbosity is considerably reduced by the '?' operator (previously the try! macro).
When I was coding in Go, i would have loved that feature. The mind-numbing error checking in go is hugely annoying.
Ugh. Yes. I really want to love go. No, I really love go. But this is so annoying.
In the short term I have a really rudimentary library I've put up here: https://github.com/asteris-llc/gofpher and a presentation on it here https://github.com/rebeccaskinner/presentations/tree/master/... (should build with pdflatex, I need to get an actual pdf built soon)
Rob Pike has a blog post that looks at a similar, although ideologically different, approach to doing it: https://blog.golang.org/errors-are-values
I think the latter is more go-ish, although also less generic than the monadic approach- it's also more in keeping with the ideology of go, but in practice I still never see it being used much in production code.
err := someOperation(args...)
if err != nil {
return err
}
that's 4 lines, 3 of them are error handlingI now typically deal with this via a Must() function that takes to values and return the non-err value or panics if there was an err.
That means you can call it like con := Must(db.Connect(...)).(db.Connection)
I think as long as panics don't cross library boundaries you're fine. And panics can cross library functions if the function is called MustConnect(). (Which means it'd be neat to have a pre-processor that generated MustX from X if it wasn't already there...)
Made even better by the use of the language's logo as the word clouds shape; the Haskell logo is ">λ=" https://www.haskell.org/static/img/haskell-logo.svg
Unfortunately the symbols will not show up, because I'm ignoring them: https://github.com/anvaka/common-words/blob/master/data-extr...
> Unfortunately the symbols will not show up, because I'm ignoring them
Makes sense. I notice that some funky unicode stuff has still managed to come out quite high, e.g. ⊇ ("superset of or equal to") :)
https://github.com/anvaka/common-words#how-are-word-clouds-r...
In general, word clouds are bad for comparing sizes. For that reason I used plain list in the sidebar on the left (or at the bottom if you are on mobile)
Ah, yes, very good point.
auto sh_ptr = make_shared(....); // Since C++11
or auto uniq_ptr = make_unique(....); // Since C++14
It saves you from having to explicitly name the type parameter and you just need to pass in the parameters of the constructor. It's not as terse as & or *, but it's not that bad.That's an annoying lie, and makes codebases in languages with genuinely strong static type systems (e.g. Haskell) much harder to read than necessary.
Hell C#'s type system is stronger than Go's, yet there is no such prevalence of meaningless variable names (first one comes in at #23, Go has 7 single-letter variable names in the top 23)
For example,
int getAge(Person p);
is clear enough, whereas int getAge(Person person) ;
is redundant.Haskell does type inference, so even if types are static they are not explicit. That's why you still need explicit variable names.
Except neither exposes that pattern (they essentially have only one such variable in their top list, and it's "i"). And I already mentioned C# which is very similar to Java.
> is redundant.
I'm not saying single-letter variables (or no variable at all which is also possible in Haskell) is universally bad, I'm saying it is disturbing to find how absolutely ubiquitous it is in Go. Variable names are sometimes redundant, that is not a universal constant. That redundancy is a factor of the expressiveness of the type system, that's hardly a claim of fame of Go.
> Haskell does type inference, so even if types are static they are not explicit.
Go has local type inference, and leveraging Haskell's global type inference is usually recommended against.
> That's why you still need explicit variable names.
No, it is not.
Also in Java if you use intellij then you'll probably get 'person' as an autocomplete, which actually makes it roughly as easy to type out as 'p', (and a better choice if your entire team has standardized on intellij.)
I'm an embedded C developer. Our shop doesn't let single-letter variables pass in code review.
Even for loop variables I tend to use `idx` these days in languages without smarter constructs. Just slightly better and more readable for no effort.
I take my bikesheds green, thank you.
I strongly disagree with the assertion regarding single letter variables. I would gently correct a junior programmer who tried to do that and I would chew the hell out of a senior programmer if he tried it. In any language.
In short, I don't think there is a particular rule that applies, that said my impression is that a lot of variable names in Go code are typically three letters or at least two, like err, buf, src, dst, ok etc.
EDIT: Ops, I missed that there is a language filter. Ignore my comment :)
It's so not a keyword it doesn't even exist. At least in the Smalltalks I've used, they provided ifTrue: and ifFalse: (and compositions thereof).
I presume I'm missing the sarcasm here...
Either you import a single member, or all with *.
Combine that with IDEs that auto-create the import statement, and there you go.
Java doesn't have the former, and the latter requires an import per symbol (it also has the ability to import every symbol in a namespace, and so does Python, but that's usually discouraged)
That seems to be an example of not using an implicit self.
> A different name for the variable fixes it
There's nothing to fix, it works just fine.
In fact, our coding style demands that we avoid using instance variables in preference of adding `attr_accessor` to explicitly encourage the use of implicit self.
# 1
def length(self):
return math.sqrt(self.x*self.x + self.y*self.y + self.z*self.z)
# 2
def length(self):
return math.sqrt(self.x**2 + self.y**2 + self.z**2)
# 3
def length():
return math.sqrt(x*x + y*y + z*z)
I prefer version 3 by far. Unfortunately, Javascript decided to take the same path with ES6 classes which forces you to use this in the body. Fortunately, it does not force you to use this in the argument list. function length() {
const { x, y, z } = this
return math.sqrt(x * x + y * y + z * z)
}
I think it's pretty short, readable and explicit.EDIT: formatting
struct Foo { x: f64, y: f64, z: f64 }
fn length(Foo {x, y, z}: Foo) -> f64 {
(x*x + y*y + z*z).sqrt()
}
...Doesn't work with `self` because that would be implicit again, though.
Anyway, I think that it is a minor inconvenience that isn't important for oneliners and extremely helpful when reading larger functions.AFAIK it doesn't work on self because the [&[mut ]]self parameter is more or less a keyword determining ownership interaction with the call subject.
The UFCS RFC would have made it sugar for `self: [&[mut ]]Self` (IIRC) but I believe that floundered.
Python's self in the arglist is a C struct pointer sneaking in from the 80's, which is the time when Python has been designed. It could be excused but there should be a deprecation PEP by now. Make it a keyword and let us type it only when we need it.
# 4
def length():
return math.sqrt(.x**2 + .y**2 + .z**2)
Ambiguity and "pseudo-keyword" are both gone![EDITED code for better readability]
def length(s):
return math.sqrt(s.x**2 + s.y**2 + s.z**2)
The beauty with self in Python is that self is not at all magic : it merely indicates that the object instance you're using will be passed as first argument of the class method.Also you can get a custom font with typographic ligatures (e.g. for self, lambda, and so on) in order to make it more visually appealing. For instance (self > 圖) :
def length(圖):
return math.sqrt(圖.x**2 + 圖.y**2 + 圖.z**2)- No static types
- A variable might belong to the scope of a method, object or class, depending on where it was first set and changed afterwards
- Implicit execution of code on import of a module (__init__.py, including parent packages)
- Any object is truthy or falsey, i.e. conditional statements don't require an explicit boolean
y = 100
x = 15
class MyClass(object):
x = 50
def __init__(self, x):
self.x = x
def length():
return x * y
What does that mean? Does MyClass(10).length() raise an AttributeError because MyClass doesn't have an attribute named y? Does it automatically recognize there's a y in the outer scope and use that, or does it call __getattr__ first (i.e. method that gets called when a missing attribute is accessed)? Furthermore, how do I specify that I want to access the class attribute x, or that I want to access the nonlocal x?Blocks like this are an imensely useful way to keep scopes clean:
int someMethod(){
...
{ // some small stuff that doesn't warrant a new method but you don't want it to bleed into the remaining part of the function, e.g.:
float x = readNextFloat();
float y = readNextFloat();
float z = readNextFloat();
float length = sqrt(x*x + y*y + z*z);
cout << length;
}
// do some other things without worrying about potentially initialized variables
...
}So then I expect the same would happen in Python and JavaScript. I don't think it would be better that way.
edit: Also, it is already required for dependent names in templates.
I also think filtering out comments would improve it - especially because so many source files include a copyright statement at the top, and the same licenses (MIT, GPL, Apache, etc) are found repeated in many different files and it distorts the results somewhat.
Presumably "summary" appears quite a lot because C# developers use markup like `<summary>` in their comments, so automated systems can build documentation (I've never used a Microsoft programming language, but a quick search brought me to https://msdn.microsoft.com/en-us/library/z04awywx.aspx ).
In that sense, it's not really a comment anymore: it's one machine-readable language embedded inside another.
That's certainly interesting, to me at least. It tells me about the signal/noise ratio of the language, the prevalence of various forms of documentation (e.g. <summary> is conventional, whilst something like <precondition> is not), etc.
Such terms clearly have an effect on a system's documentation, even if they don't have an effect on the CPU instructions being executed. But I'm a programmer, not a CPU; text files containing source code are my main I/O interface, and they most certainly do contain such markup, and hence I find it interesting to see statistics about. In comparison, I don't step through very much assembly day to day, so I don't really care very much about the compiler output (the part which the comments don't affect). I prefer to reason at the level of the language I'm using, where not only do comments appear, they're very useful!
Yes, and the IDE will auto-generate a doc comment with a <summary> because that's pretty much the most basic doc comment you can get.
> In that sense, it's not really a comment anymore: it's one machine-readable language embedded inside another.
My issue is not that it's a comment, it's that it is essentially worthless as your IDE's basic "add method" intention (or whatever) is going to add it automatically.
> That's certainly interesting, to me at least. It tells me about the signal/noise ratio of the language, the prevalence of various forms of documentation (e.g. <summary> is conventional, whilst something like <precondition> is not), etc.
<summary> is not conventional, it's the primary tag used by the C# documentation system and shown by IntelliSense. <precondition> is not that.
Just because an IDE will write boilerplate automatically, that doesn't mean the boilerplate wasn't written, checked into version control, presented to developers, etc. Even if such boilerplate were added by an IDE, and hidden from developers (e.g. using code folding), it's still there in the language.
In this case, the language is C#, not e.g. some "C#-like" language which gets preprocessed/transpiled by an IDE into C# by scattering boilerplate around.
Whilst tooling can help us live with a language's deficiencies, they don't remove those deficiencies ;)
Which is I prefer a language where I just need to learn one language, and not a separate input language because the language-as-read is to unergonomic to write so a different language needs to be defined for productively writing code.
The explicitness of go's error handling forces you to handle every error specifically. There isn't a chance (for the most part) that you'll get an error from deep in the program that you can't easily handle.
Forcing you to handle them everywhere and anywhere forces error handling to be a part of your architecture.
I am more from this school of designing a system around the reality insteas of trying to patch it everywhere, praying we have enough fabrics to catch it all.
But it would meam rethinking how we build stuff. That was not at all a goal of Go.
I wonder what percentage of `_` usage is value/type discarding in pattern matching and type signatures vs. function application.
Would be nice to somehow ditch `case`:
adt match {
Foo(x) if cond x => ...
Bar(x) => ...
}
pairs.map{ (a,b) =>
...
}- Python developers does not follow Clean Code (ala Uncle Bob) as much as Ruby , because if statement is more frequent than def and return.
- Ruby makes it possible to write in a much more functional style than Python. OR Ruby developers like to develop more in a functional style than Python developers.
- People don't really care about good variable names ("a" is a terrible variable name in scripting languages like JS and Python, still top 11)
- PHP developers might practice "return early" in functions (more return than function keywords) OR their functions just do too much :)
(click them and you can see examples of how each word is being used)
I instinctively think you might be right (at least with your second statement), but what are you using as your metric here?
With a large grain of salt
1. generally, the thing seems to mix words from all context, for instance #6 in Ruby is "should", click on the word and it's mostly comments, same for Rust's #4 "the".
2. "if __name__ == '__main__'" (seriously, TFA counts >600k of those)
3. also unclear what codebases are parsed, lots of django-isms in the conditional examples (if context is None, if request.method == 'POST', if form.is_valid)
See also: Go, where most every function call requires an `if`
> - Ruby makes it possible to write in a much more functional style than Python. OR Ruby developers like to develop more in a functional style than Python developers.
Don't confuse writing more functions/methods and writing in a more functional style, they're very different things.
Comparing word clouds of comments would actually be interesting. But then you hit the issue of "what is a comment" e.g. many language use comments for item documentation, but python uses (doc)strings.
I also wonder what the criteria is for which languages to analyze? There are a few other languages I would like to see, but maybe on github they aren't well represented...
I'm using file extension to differentiate between extensions.
You can request other languages here: https://github.com/anvaka/common-words/issues/4
As long as language's extension is unique, I think I can make a visualization of it.
What did you use to get and analyze code and how much code did you analyze?
I thought about something like this but about variables, methods e.t.c. most used words or even variables name generator based on markov-chain.
Here is more information about it: https://github.com/anvaka/common-words#how
** use the contact form at http://qt.digia.co/contact-us.
furnished to do so, subject to the following conditions:
* This file is part of the LibreOffice project.
// with this library; see the file COPYING3. If not see
So assumed licenses had not been excluded.Having a brief look at the source, I think with the licence marking approach it's still leaving in quite a few lines from each licence (see above for examples).
Super rad project :)
Yeah, autogenerated files/comments feature extremely prominently e.g. the top two items for "should" in Ruby are autogenerated Rails comments.
from django.<path> import <stuff>
for every try block./s