Stop the Vertical Tab Madness (2010)
prog21.dadgum.com
prog21.dadgum.com
Whitespace is often defined as space, \t, \r, \n, and \v. However many specs, like HTTP, will sometimes exclude \v. Depending on the underline functions products use in their HTTP parsers, you can fingerprint servers, WAFs, proxies, load balancers, whatever when using \v to separate HTTP headers lines or name\value pairs
name
alue
pairs? ;)I'm failing to see why this is an issue? The examples tutorials/references he gave only mention '\v' in tables of possible escapes. What is the downside to mentioning it, people spending 5 minutes googling for what a vertical tab is?
Unless a language removes support for escaped string literals entirely it seems odd to remove support for a particular standard, if mostly unused, escape sequence.
So the harm is that people see vertical tab and without the associated historical context come up with their own ideas of what it does and then waste time.
If you do include it at least show it in such a way that people know it's a historical oddity with really nothing but very obscure uses like the fingerprinting example mentioned elsewhere in the comments.
Talking about old terminals. I worked at a tax office, around 2005, the web was getting trendy, so old AS400 applications were on the way out. I had to use them just before they were replaced by whatever webapp was coming. It was one of the best user experience I had. That old system was so on-point. It's funny because AFAIK terminal code had close to no structure, it parsed fixed patterns on screen buffers, pretty archaic. But, the software was hyper ergonomic, responsive, simple and solid. It did infer/complete lots of fields, find non trivial issues and suggest corrections. I could barely imagine the amount of regression that were about to hit the employees when the html/js version would land, it did make me really sad.
<rant/>
Now we have this shite webapp that only runs properly in IE, it's unusable for any sort of mass-querying, and it takes a good 10 or so clicks on buttons that resize themselves to find out all the info about a server, whereas the terminal program was just "ServerLookup X".
> For instance, many old style terminals could do fields
And often had a next-field key distinct from the tab key.Wikipedia suggest CSV was already around at the time of punched cards, I guess people would prefer commas over some obscure RS code.
Comma-separated values is a data format that pre-dates personal computers by more than a decade: the IBM Fortran (level G) compiler under OS/360 supported them in 1967.
CSV isn't something computer specific, it's basic human grammar, CSV probably just put a tag on a common practice.
SOH + STX to set a document title. FF for page breaks. DLE to allow embedded (binary/uninterpreted) data like images and avoid printing garbage. FS, GS, RS, US to support tables.
I find it interesting that this coincidentally works complementary to syntax like Markdown. :p
http://jimkeener.com/posts/ADF
I did complicate it a little by trying to insert field metadata into the file.
I think the article is right: eliminate \v.
I don't have the time or the initiative to figure out what the old .ppt format does here.
That's a part of the GNU Coding Standards which say:
Please use formfeed characters (control-L) to divide the program into pages at logical places (but not within a function)..
I always found that particularly archaic.
And yes, of course I realize that vertical tab and form feed are distinct characters.
Less obviously, emacs has commands that navigate by logical pages (C-x [, C-x ]). And of course you can adjust the regex that denotes logical pages.
x = 0123
it's time for octal to go away.
Scala recently removed it altogether: https://issues.scala-lang.org/browse/SI-5205
Other languages (Python 3, Ruby) are moving to the 0o syntax.
[0] http://stackoverflow.com/questions/7825055/what-does-the-c-o...
It would be odd to support an escape sequence feature for every ASCII control character _except_ vertical tab, or to support it but leave it out of the docs.
It's just coming along for the ride with the general escape sequences for ASCII control chars features.
The burden for \v is not zero. Every programmer working on the string escapes part of the code has to read and understand the lines that implement it. And it has to be tested and documented and if your documentation comes in multiple languages, translators have to spend time translating text for a completely useless feature.
Writing software is like writing a book of code for other programmers to read. An author wouldn't leave in meaningless chapters in the book it is writing because "what's the harm?" and neither should good programmers.
EDIT: found it, it's called Aspen (http://aspen.io/simplates/). They actually were using form feed (^L or \f), but apparently have switched to an ASCII combination to separate code from presentation.
From the web archive: https://web.archive.org/web/20110412072653/http://aspen.io/p...
There's an old HN discussion about it: https://news.ycombinator.com/item?id=2410221
Whilst it's true that I've never actually used `\v`, I have included it in code before to cover genuine, necessary edge cases... For example:
https://github.com/tom-lord/regexp-examples/blob/master/lib/...
>>> print('Hello,\vworld!');
now :D
Actually I think I'll include a "guess what this control sequence does" slide before discussing them.
Maybe it's useless, but I wouldn't say it's harmful. This is a pretty overzealous rant over nothing: it's almost as useless as the vertical tab.
"In practice, settable tab stops were rather quickly replaced with fixed tab stops, de facto standardized at every multiple of 8 characters horizontally, and every 6 lines vertically (typically one inch vertically)."
You could try testing this.
As for removing it, I'm not convinced that it's worth the effort to; it's basically a single case in an escape-handling switch. The article he links to in the first line can basically be summarised as "I don't understand escaping and want to replace it with something even more complex".
Escaping is amazingly elegant once you realise how general and simple it is, and it's also very important to understand it when designing things like data formats and protocols (length-delimited fields are the best, but it is not always possible.) Ignoring escaping, which is what would otherwise occur, tends to cause rather horrible security issues.
>>> print 'hello\vworld'
hello
world
>>> print 'hello\v\vworld'
hello
world
Now I want to use it a bunch.Edit: The \v character had somehow made it into one of the descriptions for one of the user profiles.
But you should have gotten an error, of course, not the silent truncation you imply.
If you need to salvage the character, your XML library may let you specify it as �b;. That is still a violation, but a lot of libraries seem to let it through: http://www.w3.org/TR/REC-xml/#sec-references (see "Well-formedness constraint"... you are specifically not allowed to use this to do what I'm suggesting here).
Anyways, the moral here is that XML CAN NOT carry arbitrary binary, and EVERY TIME you output something in XML, something in the system needs to run some sort of encoding & illegal-character cleaning pass on the output text. The moral equivalent of "<tag>$content</tag>" in your language is ALWAYS wrong, unless you specifically processed $content into XML character content earlier. This is true even when your really sure $content is "safe". Even if you're right... and statistically speaking, you're not... do it correctly anyhow and call the right encoding function.
It's a hack, sure, having to encode/decode all the time, but if you need to store those characters, it's the only bulletproof way I've found.
Always makes me nostalgic for usenet. Which yes, technically was UUEncode back in usenet days, some slight technical differences from Base64 Encode.
[1]: No, not gripping hand... that's only for when the third choice is the dominant/default/obviously-correct-once-I-say-it choice.