Search: .lenght - Github
github.com
github.com
I have no time to work on something like this myself, but I'm sure a lot of people would find it useful, especially if it acted as a "first defense" before deployment. Curious what other HN'ers think about this.
https://github.com/Miserlou/WhitespaceBot
Feel free to fork it to do whatever you want, that's why I made it.
I use it to create useful indentation guides in Komodo. If the whitespace is stripped away, the indentation guides have gaps where there's a blank line in an indented block.
Maybe Komodo could use a different method to decide where to draw the lines that didn't depend on trailing whitespace. It looks like Sublime Text 2 has a different approach for its indentation guides - maybe the Komodo guys should look at that. But in the meantime I'm using Komodo as it actually works today, so the whitespace on blank lines is important. Let me keep it, please? :-)
I wouldn't mind stripping out trailing whitespace on nonblank lines - that wouldn't affect my precious indentation guides.
But wait a minute, what about Markdown? Two spaces at the end of a line to get a <br>, right? Does the bot skip Markdown files?
Finally, for the folks who have automatic whitespace removal in their editor settings... Please be careful: With this setting, you'll be very likely to make a commit that includes both significant code changes and a mass of whitespace changes in the came commit.
Those kinds of changes should be separated: one commit for the code itself, and a separate commit for the whitespace with a comment like "Whitespace cleanup, no code changes."
This allows people who diff the revision history to diff with whitespace significant most of the time, the only exception being when reviewing a whitespace-only change.
(Edited for friendlier tone...)
banned = ['.git', '.py', '.yaml', '.patch', '.hs', '.occ', '.md', '.markdown', '.mdown']" Remove any trailing whitespace that is in the file
autocmd BufRead,BufWrite * if ! &bin | silent! %s/\s\+$//ge | endif
which has some advantages (described in the README)
This plugin is better than using the builtin vim 'list' command because it
doesn't show an annoying highlight while you are typing in insert mode at the
end of a line. set list
set listchars=trail:•
In other people's code, use `set nolist` to prevent hyperventilation.In general, explore `help list`.
Absence or presence of trailing whitespaces is sometime significant, such as in big ereg in re.VERBOSE mode in Python.
Therefore, it must be visible, and if it is visible it must be removed when it has no use.
so, thanks, but I'll keep my whitespace.
Would be nice if Github provided better documentation and a selection of validation templates to include in new projects. This would better leverage the power of Git and its distributed nature than a bot running on Github.
For some reference material check out:
http://book.git-scm.com/5_git_hooks.html
http://progit.org/book/ch7-3.html
And a couple simple examples: http://mark-story.com/posts/view/using-git-commit-hooks-to-prevent-stupid-mistakes
https://github.com/ReekenX/git-php-syntax-checker LANGUAGE #LENGHT #LENGTH = #LENGHT/#LENGTH
JavaScript 4252 2907459 = 0.0015
C 18981 2902857 = 0.0065
Java 7706 2348900 = 0.0033
Ruby 10789 1690604 = 0.0064
C++ 9458 1315552 = 0.0072
PHP 3116 1167924 = 0.0027
C# 1352 937647 = 0.0014
Python 3662 737292 = 0.0050
Ruby 1232 380484 = 0.0032
Perl 1239 258892 = 0.0048
Objective-C 679 238051 = 0.0029
P.S. There is something wrong with Github's language breakdown algorithm, sometimes it shows same language twice with a different number of hits.I put together a GitHub Illiteracy Index script https://github.com/jond3k/sandbox/tree/master/github-illiter... which you can play around with if you like :D
Whereas C and C++ programs tend to have a lenght operator implemented by the programmer, and from there the error gets snowballed by IDEs and debuggers.
Also fun to search on "functino".
I would guess that most of the spelling errors get propagated through autocomplete. That's how most of the spelling errors in my code get there anyway.
LANGUAGE #LENGHT #LENGTH = #LENGHT/#LENGTH
C# 1352 937647 = 0.0014 <- best
JavaScript 4252 2907459 = 0.0015
PHP 3116 1167924 = 0.0027
Objective-C 679 238051 = 0.0029
Ruby 1232 380484 = 0.0032
Java 7706 2348900 = 0.0033
Perl 1239 258892 = 0.0048
Python 3662 737292 = 0.0050
Ruby 10789 1690604 = 0.0064
C 18981 2902857 = 0.0065
C++ 9458 1315552 = 0.0072 <- worst # Language Illiteracy
1 C 0.02877583
2 Perl 0.01635618
3 Ruby 0.01560477
4 JavaScript 0.01330989
5 Shell 0.01235425
6 Python 0.01046104
7 PHP 0.00910218
8 Java 0.00736395
(For height, length and hierarchy, averaged out)And you thought this would end up being a PHP joke...
https://github.com/jond3k/sandbox/tree/master/github-illiter...
if(somethingThatMightBeAnArray.length){
// do things with array
}
So misspelling of length can be making a lot of code out there behave in an unexpected way."bar".length; // 3
I once worked at a company where a very early piece of code had a typo "properites" instead of "properties". This misspelling became institutionalized, and was used throughout the codebase because it was deemed too expensive to fix. And this was with a static language (with good IDE refactoring support)!
The way javascript (which is what is linked) handles this, as amirhhz described it, leads to silent errors which could turn out a lot worse.
grep -R properites .[1] http://moishelettvin.blogspot.com/2006/11/windows-shutdown-c...
But as to the "usually painless" at Google. So when does that pain happen?
Can you take me through the following scenarios: change a variable name, change a base class name that lots of people extend from, file renames?
How do you go about refactoring? Do everything at once? Breaking it into pieces? Do file rename then variable and base class renames? Or smallest piece at a time?
Once the refactoring is complete how do you communicate to others the changes so when they merge the code in they don't get too messed up? Or worse undo something in the refactoring. (Also follow up is it better to do the big refactoring so there is the one big merge or a bunch of little refactorings and lots of little merges across the spectrum).
I guess the code change isn't the problem. It's making a big change and getting people on the same page is much harder. Especially when their are varying degrees of skill and experience on a project. And it's this stuff that is painful and leads to not wanting to do big refactorings at a lot of shops.
1. Most importantly, the version control head is always the point of reference, and the burden of merging is on people who keep long-lived pending changes. This means that conflicts are resolved as soon as possible by a person who actually knows the context, instead of being postponed until a dreaded merge window. Ultimately, a programmer pursuing refactoring is only responsible for making sure it works on the head, and should announce the change so that others are prepared for merging.
2. There are some huge code bases at Google, but nowhere near the size of Windows. On the other hand, I'm sure that even Windows has to be separated into more or less decoupled components. When I doubted that you work on the same code with thousands of other programmers I was thinking in terms of components, not final products.
3. Cultural aspect shouldn't be disregarded. Code hygiene is encouraged at Google, and some people volunteering their 20% time to help with that. Moreover, there are some custom tools that make global refactorings much easier and safer.
Hope that was helpful.
EDIT: Justly downvoted <strikeout>chastised</strikeout> for attempting humor without understanding.
EDIT: Eh, apologies if this sounded like chastising -- I didn't mean to. As a developer who's been trapped in "Windows-mindset" for many years, I wanted to try to inspire other Windows devs to try to use *nix-based solutions even if their only option is Windows development. Cygwin is in a very good spot right now -- it's achieved so much acceptance that even the most hardened institutions now allow it to be installed.
What's the trade-off by having "undefined" returned instead of having an error reported as soon as the code is loaded?
For core methods like 'length', it seems silly to think that you'd want to redefine it. And indeed, it's usually counterproductive - that's why any experienced JavaScript dev will have coding conventions like "Don't muck with the prototypes of built-in objects."
But at the application layer, this can be really useful. Imagine you're adding a new field to a message deep in the storage system, and then you want to pass that along to a template in the rendered HTML. It's really useful to be able to do this without recompiling & restarting each individual server between the backend and the frontend, and just edit a few template files and have them automatically pick up any changes to backend data formats.
Ditto adding a new database column, if you're using an RDBMS - it's pretty handy to have your model objects instantly reflect the new field, instead of needing to manually add accessors to each of your model classes. Rails and Django are built on this principle.
Also, you have a versioning problem with statically-compiled code in a distributed system. Imagine that you add this new 'lenght' field to a backend message, and add it to the frontend, and they both compile & deploy. Now imagine that a message from an old backend hits a new frontend (it's not possible to upgrade a whole distributed system at once without downtime). What does the new frontend do with it? It needs a piece of data, but the backend had no idea that it had to provide that piece of data. The only thing it can do is return the equivalent of 'undefined'.
In C++/Java code, you usually deal with these by inventing frameworks. Google code, for example, is littered with
if (msg.has_new_field()) {
run_long_complicated_ui_display_routine(msg.new_field());
} else {
fall_back_to_old_behavior(msg.old_field());
}
checks. If you use a more dynamic language like Python, you can use language mechanisms to represent undefined values or fields that are defined at runtime. If you use a static language, you're stuck mimicking them with hashmaps and null.Now the normal (and optimized) route is to find the method on a’s method table and then call that, but if a doesn't have that method then a second method may be called to allow this to be handled. Once you have that sort of mechanism you can make ORM libraries that dynamically examine a schematic at run time and generate accessor methods only as they are needed, decorators, proxies and many other patterns become wonderfully simple, and there are often many more opportunities for meta-programming at run time.
The downside is of course that it becomes harder to find errors when writing or compiling, but tight integration of your development environment with your runtime can help with this.
Of course there are rare cases where "lenght" is a variable and that name is used in every instance but mostly, these are bugs in code that we all use.
Deleted comment
var lenght = 23;
console.log(lenght);
will not cause any troubles. And many of the results returned by the search are of this kind.Come back to that code in a year and try to extend it, stuff will break because you start to use the correct name.