Google's Shell Style Guide
google-styleguide.googlecode.com
google-styleguide.googlecode.com
Thats the only coding style that remotely makes sense to me: https://www.kernel.org/doc/Documentation/CodingStyle
While I'm at it, half the justifications Google gives for the shell guide are inaccurate. Looks like a "overall people are used to this style so we're using it and we'll try to justify it without knowing why"
This is a much better guide:
http://devmanual.gentoo.org/tools-reference/bash/
Example: Gentoo's [[ ]] explanation actually makes sense.
We have these machines that are meant to take the tedium out of our lives, let's USE THEM. Why do we not have tools that format things they way I like them? Who cares how the other team members like it? They can use the tool, too. Who cares how it's stored in your VCS? The tool should just be part of the workflow and format spaces/tabs the way I want them and maybe revert to a canonical format for committing to the repo.
Personally, I prefer tabs. I read code best when it's indented four spaces. But John wants to use two spaces in this code. And Ralph likes three. Holy Hell, don't make me wade through someone else's preferences to understand this codebase! Use tabs and people can set them the way they want. Yes, tabs are broken in many places (e.g. Objective-C methods can get so verbose that they need to wrap; convention is to wrap and align the colons; tabbing as far as possible and then spacing to fill breaks things when Ralph formats his tabs to three spaces; in this case, there is indeed a solution: tab to the start of the original line of code, and space to format from there ... but then the IDE has ideas of its own about what should be spaces and what should be tabs...)
Maybe you can read that code with only two spaces of indent, but I can't. I'm gonna need that reformatted to put a larger visual separation between scope changes or whatever required the indention. Rather than have this debate, a tool should be forged.
Which is the problem. If you think it's possible to make a tool to solve this, great - do it, I'll happily adopt your tool as a precommit hook and use tabs to indent everything (or go on using spaces - your tool would turn them into tabs, right?) But given that this problem has existed since the '70s and no-one has solved it, I'm inclined to think it's impossible. So I'll use spaces in the VCS, which means agreeing an indent width so that our diffs make sense (and I don't really care whether that's two or four, but it does need to be defined and adhered to across the team).
And repeated experiences everywhere point to the fact that code text is already very hard to manage, very sensitive to alterations, misreading, etc. I'd rather remove layers and tools than add more.
Edit: obviously, readability is ask about convention when they are within the realm of sane acceptable conventions. And by everywhere I mean in teams where the loc count is above a few thousands.
I, on my side, claim that guidelines are there for a reason, and should be followed.
> Use tabs and people can set them the way they want.
If members of a team set tabs the way they want, how do they limit line lengths to 80 columns, which is a requirement that's specified in the Linux kernel coding style that was linked to in the comment you replied to?it is still better to have to switch to 2 tabwidth before committing to check line length then to be forced to use 2 spaces forever.
hello( world,
solarsytem,
universe
... );
You can't if you just use tabs - you would have to mix tabs and spaces. hello(
world,
solarsytem,
universe
);
which doesn't require perfect line-up with the first entry, and is still clear to the reader what's going on.this ensures the closing paren immediately and makes it possible to comment out things that are inside the parens.
I'd like to see a diff tool designed around this.
The hard part is creating a tool that can change between styles. Such a tool would need to be made for every language separately which is a pain. I've been wanting to try something like this for years but have never built up the courage.
Mildly surprised to see people arguing for specific tab widths...
http://www.cs.umd.edu/~ben/papers/Miara1983Program.pdf
I'd classify it as an attempt rather than something decisive but whatever floats your boat.
The level of indentation that seems to produce optimal
results in comprehension is between 2 and 4 spaces; as the
number of spaces increase, the comprehension level
decreases.
On the other hand, the sample program hard parts that were indented nine times (!) in places. It's an interesting read, but I'd not try to use it in a discussion.Personally, I don't care; just pick one (hopefully one that matches what most people are using with that language, eg. 2 spaces for Ruby), and get on with it.
Infinity spaces gives less space for code than zero.
Conclusion: use as wide an indentation as possible while still keeping enough space for the code itself.
Tabs have the benefit of having a dynamic size, so I could simply adjust the tab width within my editor. Spaces also do not play as well with editors. Unfortunately, browsers offer no means to change the width of tabs, so it's impossible to properly display code using tabs. This is especially frustrating when reading the source of a page indented using tabs. It would be much easier to read with spaces, but not as easy to read as if I were able to set tab width.
tldr; don't use tabs.
I read that piece many years ago and thought who cares, but I've later changed my editor to never use tabs. He's right.
Tabs are nice in theory, but in practice you end up with garbled code. It's better to train yourself to get used to seeing code bases formatted slightly differently, just one of the things you have to get over, IMHO. Like camel notation version underscores versus dashes.
That's a pretty good point. Indentation is indeed just one small aspect of the overall code style - why would you want to have that configurable and nothing else?
It has always been my own argument in favour of tabs:
- The original argument against tabs was based on Lisp indentation rules.
- Languages with very complex indentation rules like Lisp should use spaces. There's no exception to this rule.
- Languages from the C family should use Tabs, as Lisp indentation rules don't apply to them.
- Nowadays almost all popular languages can be considered to be part of the C syntax family. Python and Lua are notable exceptions. PHP and Objective C files should use tabs.
http://trac.common-lisp.net/mit-cadr/browser/trunk/lisp/zwei...
MIT's Zmacs, the editor of the Lisp Machine indents with tabs.
However, one very old editor using tabs doesn't mean Lisp programmers actually prefer them. It may even be the cause of tabs hate.
Basically your arguments are from a strange parallel world.
pre { tab-size: 2; }
int someCodeDemo ( [tab] int fred,
[tab] int wilma) {
[tab] return fred + wilma;
}
I guess it needs some vertical space between blocks in order to distinguish them, hence why putting { on the next line works.In my opinion, it also looks horrible with Lisp code (where I pretty much consider indentation a solved problem).
Edit: Oh god, how do I format this... Ah, there we go.
The only problem I've noticed is that it muddies the undo stack.
I can understand (grudgingly) 4 spaces for Python
But for C/C++ 4 spaces are awful. Maybe because C lines end up being longer?
Tabs, they exist. Use them
Then the additional issue that everyone maps tabs to a different amount of space characters.
Only spaces can give a proper consistency without relying on external tools.
To me, the problem is consistency. I'm a 4 space man myself, but tabs don't outright offend me, as long as they're everywhere. I'd rather commit edits with tabs than attempt to convert the whole project to spaces, or worse, commit some spaces and some tabs to a single file/project.
This fully breaks the indentation, which is exacerbated when multiple teams working on the same code basis use different editors and tab settings.
So in the end one needs to rely on external tools that go over the source code to normalize the use of tabs.
However, after a character that's not a tab, i.e. any printable character, spaces should be used.
It is the best of both worlds, and can never be misconfigured by changing the number of spaces that a tab represents.
It is also what tab advocates have been saying for years now.
https://docs.google.com/document/d/1HMxg6fv_yig_mVvnxsC8CwTf...
Unfortunately I haven't had time to update the code, and it doesn't work with more recent versions of emacs. Now I'm using vim and all-spaces-all-the-time.
A tab means 8 spaces. If you want your editor to show something different that's up to you, but it's non standard.
If you have a source file with only tabs (and you should, then nobody has to run an re-indent script because they can't read it), you should be inserting tabs. If you have a source file with only spaces, you should map the tab key to insert the correct amount of spaces before editing. If you have a source file with both tabs and spaces for indentation, you should re-indent.
Type this in a bash terminal: $ printf '1234567890\n\tx\n'
For the second paragraph: I couldn't agree more. There's even some "geniuses" that indent python with the 1st level being 4 spaces and the second level as a tab. * sigh *. Really.
printf '1234567890\n1\tx\n12\tx\n123\tx\n1234\tx'
1234567890
1 x
12 x
123 x
1234 x
Tabs are not fixed width, that's the whole point of their existence.This is a non issue when you have nothing to the left of the code (that is, when indenting)
(and the x should be under the 9 btw)
On the real corporation world, it is a whole different thing.
Good luck enforcing coding styles across multi-site projects, with rotating developers and off-shoring subcontractors.
Still, projects like this have something very thorough and that goes much deeper than just tabs vs spaces: http://www.stroustrup.com/JSF-AV-rules.pdf
Therefore, to me, your 'every language' reduces to a list of one.
Then again, i haven't had the chance to define the coding style guide of anything bigger than a hobby project, so i might be lacking pragmatic experience.
Every editor can do it, I'm not complaining about pressing the space bar 4 times.
I'm complaining about the appearance of code indented with 4 spaces.
It's ludicrous that these space-wielding heretics (to borrow from the kernel style guide) are keeping everyone from having that.
In short: configure your editor to display to your tastes and save to whatever style guide you use in your organization. Problem solved.
You might also be surprised by how many source files with indentation spaces have off-by-one indentation errors.
I will add that to my list of arguments against using spaces for indentation. It is so true.
vim has 8 spaces for tabstop and that's what its docs say:
"Note: Setting 'tabstop' to any other value than 8 can make your file appear wrong in many places (e.g., when printing it)."
(Sad that tabs solved that problem decades ago (by having a byte that means "+1 indent" instead of having to build some ascii-art that looks like an indent) but people screwed up the implementations so badly :( )
Failing that I'd rather people just used spaces and no tabs. Using tabs for alignment just means it's going to look fucked up everywhere else. At least spaces are consistent.
And my IDE is called "vim".
Sometimes it's called "ed".
See, most of the languages (I've seen so far) have a well-defined rule of identation. The thing you get when you do "select all, ident file" in your IDE / editor. At this point it doesn't really matter whether you use spaces, or 2-space tabs, or 4-space tabs, or mix, or whatever - there's only one way non-whitespace characters can be positioned, so it will (in theory) look the same on every editor after you reindent it.
The only place where the choice of tabs vs. spaces actually matters is when you want, for some reason, to break the default rules of identation for your language and position something manually. But in this case, there's no debate; tabs are not suited for precise, manual positioning. Only spaces will be guaranteed to make the code look the same everywhere.
Yes, the IDE does this, but indent-length and tabs/spaces is mostly ALWAYS definable within the IDE's settings, and is more related to the IDE than the language you are currently editing.
Still there are some languages where 4 spaces indent is fine. But two is just horrible. Also I like tabs. Convert tabs to spaces in the editor and it works out pretty fine. 2 spaces is just too cluttery.
Note also on the [[ ]] usage... you'll see that [[ ]] won't work for boolean algebra ie. '||' and '&&'
Someone needs to make some kind of GitHub integration that lets you download code using whatever esoteric formatting you prefer, then transform back to some given standard on commit. Then everyone can finally just agree to disagree and get on with life.
# Long commands
command1 \
| command2 \
| command3 \
| command4
However, by ending each line with the pipe the continuation is implicit and the backslashes may be omitted: # Long commands
command1 |
command2 |
command3 |
command4It makes it clearer that the lines following are continuations; especially the last in the pipeline.
The trailing version seems more natural with operators like comma, but less natural with operators like minus, and is far clearer with semicolonless languages. I suspect if I were to go through old code, I'd find both uses, but generally I prefer trailing.
The exception to this is languages like Haskell, though indentation could serve the same purpose.
It's one thing to use bash consistently everywhere but as a heavy multi-machine shell user I've been bitten by incompatible or missing external utilities more often than I care to admit. You might be surprised how many systems aren't using the GNU utilities, have them running in a weird mode or are using ancient versions of them.
Maybe Google is religious about keeping all their environments identical?
https://www.gnu.org/software/bash/manual/html_node/The-Set-B...
Also: set -x is the best debugger in the world; really.
I thought the ./* wildcarding for safety was really cool. they also could have covered find -print0 | xargs -0 for safety.
I'm glad they talked about "$@" being usually always the right thing to do. That's been a hard won learning experience for me in the past...
I know a few "Use the source, Luke" people who will rage at that.
- I can add to my own code, even if it's years old, without re-reading every line every of every method I need
- I can contribute to a code-base built by multiple people without reading every line they've written
Shell scripts are almost the complete opposite end of the spectrum:
Shell script functions are usually only created as a last resort.
Global side effects (creating temporary files, changing global system state, global variables, etc.) are what shell scripts are all about.
There are rare snippets of shell scripting that are different, using local variables and doing some sort of calculation, but that is the exception, not the rule.
While ideally comments would be prolific, poetic, and perfect, some commenting is always better than none and most developers have bad habits of not commenting their code, so pushing them gently in the direction of more, not less, usually works.
Hence why they'd prefer you write Python, not shell.
I would rather put that the comments should be succinct rather being prolific. Better to put some explanation on tricky parts of the code as comments, and have method/function/class behavior as javadoc, pod, pydoc etc.
Shell should only be used for small
utilities or simple wrapper scriptsTools that take the AST and output standardized code for peer review and documentation sounds a little better. It would not deal well with the only human problem really worth having a style guide for - naming things. But at least humans aren't forced to jump through hoops. And the naming thing possibly can be settled with an interface that asks something like - "what do you want everyone to call the 'BitWarper' symbol?", for all named symbols.
But well written software is no place to express individuality - we just want it to work and not make our eyes bleed when we have to fix it! Even better then, just have machines generate and test all the code based on systems of higher order rules and style guides in situations where factory manufactured code is necessary. The outputs should be reasonable if the requirements are well specified (NASA style). Humans can come in after and do the real fun work in optimizing and finding clever hacks (if environment is not mission critical and such liberty can be safely taken).
Having humans program character by character, with their bare hands, while also suppressing creativity, is unnecessary in this Post-Industrial Age.
Yes, it matters.
It really does a lot for readability. I recently summed up my reasons for doing it in all of my projects: http://ubercode.de/blog/80-columns
This also applies for a lot of remote-access tools -- serial console as well as direct, so even if I'm not at the DC, once I'm on those interfaces, something's likely fucked up or headed that way.
Personally I prefer a line limit closer to 100, and 4 space indents in most languages. Some languages end up with a lot more indenting than others owing to structural literals and lambdas, and they benefit more from a smaller indent.
Skinnier prose is easier to read because of the sequential nature of it, which is why I do think comments blocks should be limited to 80-100ish characters. But code? No.
Write your code, run "go fmt", add, commit.
If the justification is to localize variables which should be global, and still will look like "global" to the rest of the program, it's ok.
But this is mainly used without sense, by personal tastes, and usually only makes sense to the author and does not have any real benefice.
I prefer the concept "consistency is to not read unuseful steps", instead of "add unuseful steps for consistency". That's how I think when I read code.
void my_function(int some_parameter, int another_parameter,
int the_third_parameter) {
if (some_parameter != another_parameter) {
here_we_are_in_an_if();
}
function_calls_work_fine_too(some_parameter,
the_third_parameter);
}Though setting your editor up to do that automatically is an incredible pain, and it kind of forces you to use a monospace font.
Anyway, it seems like this system necessitates that the leading white space on a given line must be a mixture of both tab characters (\t) and spaces, unless one sets their editor to insert space characters as tabs (which is standard), and is obviously living in a state of sin.
def foo():
TTbar(asdf, zxcv,
TTSSSSqwer, oiuy)
no matter what anybody uses for tab stop size, the alignment won't get messed up.In Linux they say tabs are 8 characters, which no they are not, so they can have an 80-column line limit. Anybody programming with a less than 8 character tab can't get the column limit right (the number of characters on a line changes with the indentation level).
Tabs are invisible characters that don't have a standard width so are always causing problems like this. They are used because many programmers use editors where they would have to actually press space multiple times to indent/unindent (ie bad editors) and because source control doesn't know when an indent change actually means something vs just being cosmetic.
That can't be an argument for using spaces in code, right? Rather, the terminal should keep track of which parts of its display were generated by a tab character in case the user copies the output?
I don't see how
$((${X} + ${Y}))
... is more recommendable than ... $(( X + Y ))
for example.But well, guide styles are a good thing, and this one can help to many people "not used" to deal with shell scripts to follow some basics.
Edit: and help to people used to it, on working in a team.
I wonder what Google are using Solaris for, and if it's just legacy stuff.
Of course we write style guides and while I like the idea that it's part of the tool, that's rarely the case.
I don't follow the objection about a colleague helping ensure consistency across a team. I'm really not sure why competence comes into that equation either.
Agree with the idea any style guide should be automated. Several CI servers I know of can incorporate style checkers and their reports into their workflow so this can be made really hands off, even to the point of automatically failing code review stage before its been lumped in the review queue.
Don't agree at all with the idea this type of doc should be rejected. In fact it's completely wrong to jump into automation without having "found which way is up" manually first time around.
All these rules make sense for Lisp, and no sense at all for C.
And they did wrote the parser that applies those rules, in Emacs you never indent Lisp code, you press a key and the current s-expr automagically indents better than you could ever dream of doing it.
> It is not necessary to know what language a program is written in when executing it and shell doesn't require an extension so we prefer not to use one for executables.
I disagree with their recommendation against using file extensions for executables, and I'd love to have my mind changed about this.
Using an extension gives you automatic syntax highlighting. It also lets you quickly glean the type of a file when exploring a directory for the first time, which is more helpful than simply knowing whether the file can be executed.
Why does a lack of necessity override those two benefits?
And then the editor doesn't know what to do with that header file without extension. Yes, if you see some Google C++ code they do that
Really, horrible practice.
But apparently they stopped this nonsense
http://stackoverflow.com/questions/301586/what-is-the-differ...
Btw, many editors understand // -- C++ --
Nevertheless Google's c++ style guide (https://code.google.com/p/google-styleguide/source/browse/tr...) in fact says that headers should have the .h suffix.
> Btw, many editors understand // -- C++ --
Maybe, but if I vim /usr/include/c++/4.2.1/iostream doesn't work (only if I set it manually)
When you run a program, your concern should be what it is called. Not how it is written.
A language-specific filename extension puts an implementation detail in userspace. If you run a binary executable, you shouldn't care whether it's written in C, C++, Fortran, or any other language (though source files, being used only by developers and compilers, generally do have extensions).
Worse: if you decide for whatever reason to change the implementation language, you're either forced to track down and change all references to the program name, or to retain the (now incorrect) filename extension for backwards compatibility.
And, as noted, magic(5) or the shebang line should correctly identify the file type and language for syntax highlighting -- if not, your editor is broken. Replace it with a shell script, "editor.sh".
file(1) will tell you the types of files in a directory with far greater accuracy than filename extensions can.
Cron jobs, other scripts, production jobs, etc.
So having to hunt down and rename everything ... is a PITA.
As for knowing the file type at a glance, I'm not sure how often I need this. I'm normally looking for the file by name anyway. If I needed to determine the file-type, I'd write a script to parse the shebang lines of executable files in the current directory and generate a list of the files with their hypothetical extensions (based on a hash/dictionary/whatever). I don't need that very often, so, for me, the trade is worth it.
explore a directory? file -i directory/*
I follow the same convention than Google, just that I use .bash for libraries instead of .sh (as they state the interpreter should be bash, I think they should apply my naming instead of .sh)
It's suddenly a year from now, and your foo.sh tool needs some new features, or is too slow to do the job any more because your requirements have changed.
You decide it's grown too much for Bash and want to move to Python for maintainability, or Go for performance, or C++ to link with some library you need to use.
Now you have to tell your team (and any other teams that have found your tool useful): "We only want to maintain one version, so don't use 'foo.sh' any more. You have to use 'foo.py' (or 'foo.exe' or whatever). Oh, and have fun changing all YOUR scripts and tools that reference 'foo.sh'!"
That's one reason.
Just change it to zip foo.sh + foo.py. Then:
cat foo.sh
#!/bin/bash -e
./foo.py
cat foo.py
# Everything else has moved here.
Much better.
You now have two files to do one thing, and your solution doesn't work (it doesn't pass the arguments from the shell script to the python, or pass the exit code back, and it's making an extra process which can screw up monitoring). You could add the extra code to make it work; but even if you did, you're going to a lot of extra effort to become equal, not better :P
And I'm not sure how a script that calls another script screws up monitoring.
Plus, if someone cares about high level stuff like monitoring but is worried that people's tools might break if he changes script.sh to script.py later on, I think he needs to sort out the lower level stuff first. Like distribution and packaging :)