ASCII Delimited Text – Not CSV or TAB delimited text
ronaldduncan.wordpress.com
ronaldduncan.wordpress.com
Everybody hated it. Most text editors don't display anything useful with these characters (either hiding them altogether or showing a useless "uknown" placeholder), and spreadhseet tools don't support the record separator (although they all let you provide a custom entry separator so the "unit" separator can work). Besides the obvious problem that there's no easy way to type the darned things when somebody hand-edits the file.
There are representational glyphs for tab, return, and others (⇥, ↵), and editors can show them in 'show whitespace' modes. There could be representational glyphs for these control characters, too. I'm not sure about the history of these symbols, but I imagine they were initially on keyboards. But if these control characters were on keyboards and had a representation in text then they'd be just as useless as the tab character is today.
Precisely what makes them valuable is their difficulty to type or display.
Again, if they'd caught on you'd imagine there would be some eventual convention in text-editors for what keybind would be used to enter them. Too bad the AltGr key (intended for entering rarely-used glyphs) doesn't appear on pure-English keyboards.
To get around this, I have taken to using xmodmap to turn my right alt key into altgr.
The meaning of the key's abbreviation is not explicitly given in many IBM PC compatible technical reference manuals. However, IBM states that AltGr is an abbreviation for alternate graphic, and Sun keyboards label the key as Alt Graph.
Apparently, AltGr was originally introduced as a means to produce box-drawing characters, also known as pseudographics, in text user interfaces. These characters are, however, much less useful in graphical user interfaces, and rather than alternate graphic the key is today used to produce alternate graphemes.
However, I discovered that Alt + Control = AltGr when I needed to use it at work[1], so it's simply a shortcut, I think.
[1]: We have (I recently switched to an UK keyboard, because it suits me better) keyboards at work, because the Croatian (all slavic languages, to be honest) is horrendously counterproductive for programming. Google the layout, and you'll realise why. An example: You need to press AltGr+B for `{` (if I remember correctly).
What happened was that the Ctrl key became synonymous with "command" after Teletype, so it became more about doing something. Think about Ctrl-x, Ctrl-c, and Ctrl-v as an example, but you still see some relics like Ctrl-d as End of Transmission (EOT) to close a shell or terminal. Alt is like a shift, but it is actually closer to the Fn key on most laptop keyboards. It was an alternative function of that particular key, so where the shift key provided you with an alternate case, Alt was more akin to an entirely different key... it isn't Alt plus an 'a' key, it is Alt-a.
AltGr was like another Alt key. It was originally there to allow you to enter an alternate glyph, especially line drawing characters available in extended ASCII, B0-DF. I thought it was a mapping closer to flipping the most significant bit to 1, but it doesn't exactly overlay the lower ASCII range, so that might be another change that evolved on the way to the modern keyboard.
To your original point, Microsoft Windows will now usually treat the chord Ctrl-Alt as AltGr. I don't know if that is with all layouts, or just those keyboards that lack AltGr. I find that most Linux distributions tend to follow Microsoft's lead and provide similar mappings but now they even repurposed the Win key as Meta or sometimes called Super. So it is likely that Ctrl-Alt is commonly the equivalent of AltGr.
For the propose of this discussion, I think it'd be better if Ctrl could be used to type these text separators, but the way modern operating systems map their modern keyboards, it might be difficult to ever reach consensus on how this should be done.
The clever thing is how some international keyboard layouts use it like some kind of "second shift" for typing character with accents/decorations: like AltGr+a => "ă", AltGr+q => "â", AltGr+s => "ș" etc. ...but not even these keyboard layouts are popular, and usually marked as "alternative" or "programmers' layout for language XYZ", because people are stupid and refuse to learn how to use this and prefer instead a funky layout national language keyboard instead of an US English keyboard with an AltGr that would just solve 99% of special characters problems.
If all the keyboards in the world would just be US English Standard keyboards with an AltGr (most US English keyboards I've seen do have an AltGr!), all latin-alphabet languages with special characters would be easy to type, we polyglots could easily use the same keyboard for typing in multiple languages without having to remember what keys' positions have radically changed on each layout... but people are stupid and refuse to learn even simple key combinations.
Oh, and somebody should shoot the British (and French) for adding that annoying extra key to the right of the left Shift that I always have to disable (and making the Shift much smaller), and for creating extra confusion by branding them as "british international" or "us english business" keyboards.
(I do really miss the days when Mac keyboards had the weird symbol they use for "alt" in menus on the key.)
In the past, I've bought US Mac keyboards just for the #, or switched the keyboard layout in software to US.
Must be annoying to have to press shift to get those symbol chars while you're coding, eh?
Honestly no. It's all muscle memory for me, I don't think about it any more than I think about typing capital letters.
I just think people will criticise their local layout no matter what. I use both a French AZERTY and a Québec QWERTY everyday (at work/home) and I think they are simply equally good for both typing French and for coding. The Québec keyboard (maybe actually Canadian multilingual or something) might have an edge because it more easily allows typing accented letters in uppercase, but on the other hand it doesn't let me type the € sign, so...
Actually, I switch between three layouts on my machine (the third one is for my mother tongue), and I also use AltGr for typographic characters.
Exactly. A naive[1] character-delimited format is only robust when its delimiter isn't on the keyboard. Which means it's practical to read and write only in specialized applications, not in a text editor. Which sort of defeats the point of a character-delimited format anyway. If your use cases are constrained to specialized applications, you may as well just use JSON or something similar, instead of a character-delimited format. (I'm imagining a world where we get to pick our ideal formats, not one where Excel happens to have an almost-working CSV implementation.)
[1] By "naive" I mean one where the format specification describes record and line characters only. Thus the format has no escaping system.
But if it's not on the keyboard, then the user won't type it. So it doesn't get used.
Tab, comma, semicolon are actual characters people DO and need to use inside records. So not only they can be typed into records, they absolutely HAVE to be typed for most textual records.
ASCII record separators are not needed inside records at all. If someone adds them "for how it looks", it's his problem.
One (not using a record separator inside a record) is a matter of choice and doing the right thing. The other (not using comma, semicolon or tab) is a non starter.
Obviously there's a difference.
I wouldn't be surprised if the last time they were on a keyboard, it was a teletype keyboard or a keyboard that punched cards!
Although, actually, were they just typed using the control key from the start?
I think the more fundamental problem with using the ASCII control characters is that there's no way (or rather, no obvious and standard way AFAICS) to use them to define (most) recursive data structures. That's not a problem when you stick to simple CSV-like arrays-of-arrays, but it's enough of a restriction to make one wonder if bringing ASCII back is worth the battles.
Unit Separator is Control-_ (underscore) and Record Separator is Control-^ (caret), for instance.
Most modern text editors won't pass through every control character. Vim lets me type the unit separator, but not the record separator, for instance.
Control-C and Control-D, End of Text and End of Transmission, still have utility in most shells, and it goes back directly to these ASCII control characters.
At least, that worked for me. Interesting that the unit separator does not have the same requirement to precede it with Ctrl-V.
More info in the vim docs:
http://vimdoc.sourceforge.net/htmldoc/insert.html#ins-specia...
You can also use Ctrl-K to enter RFC1345 digraphs, in which case unit separator is 'Ctrl-K US' and record separator is 'Ctrl-K RS'. (See :h Ctrl-K).
Emacs has similar features, as well as a nifty TeX input mode: http://stackoverflow.com/q/6269618
Most modern text editors won't pass through every control character. Vim lets me type the unit separator, but not the record separator, for instance.
Vim has bindings for some control key combos, which is why you can't type them directly. I can type control-_ without issue, but to get an ASCII 30 (RS) to show up, I have to type control-v first (just like when at the shell prompt).
Useful hint: if you look at the output of man ascii (at least on linux with man-pages-3.22) find the control character you want to type in the left column and look in the same row on in the right column to find the letter/key to use with the control key.
For example, to type NAK, it's <control-u>. Vertical tab is <control-k>:
$ echo <control-v><control-k> | od -c
0000000 \v \n
This is useful for control characters that don't have backslash escape expansions.Similarly, Ctrl-L clears the screen in most Unix apps because that generates the "form feed" character, and on a line printer or paper-based teletype terminal, "form feed" means to advance to the next page -- a nice empty clear piece of paper.
There are more details in "man 3 termios", but be warned, the TTY layer is not a pleasant subject for reading.
I'm pretty sure every possible combination of 7-bits is assigned SOME name in ascii, and Ctrl-any-case-insensitive-letter is a defined 7bit value.
This chart says control-s is `DC3 (Device Control, X-OFF)` and control-q is `DC1 (Device Control, X-ON)`. Yup, that's what they do alright.
http://www.unix-manuals.com/refs/misc/ascii-table.html
Don't forget everyone's favorite control-G, BEL, useful sending beeps across the wire on old school IRC, BBSen, and chat services. Or, you know, telegraph machines or something.
On most Unix terminals, these have the effect of pausing and resuming the display when a bunch of data is scrolling by.
Also, the classic control-G is still in C. "\a" is "alert (beep)."
This is slightly different than being able to actually generate those characters as input. If you put a control-d in a file, nothing will stop reading the file when it sees the control-d, it's just another byte. readline knows how to interpret a bunch of control characters. There's also the terminal driver that has interpretations, which you can see with `stty -a`. When the tty layer sees a control-d or a control-c, it closes stdin or generates SIGINT, respectively. But even this can be dependent on if you're on a pty or attached directly to a serial line. There's nothing special about the mapping of control-d to end-of-transmission character, that's just how the layer that sees it is interpreting it. You can, for example, change how control-c is interpreted by `stty intr ^B`, which will make control-b generate SIGINT. And there's a way to put the terminal driver in complete transparent/passthru mode.
And this is in fact how they show up in vim (or at least my fairly uncustomized vim), using ^ to stand for ctrl as was once conventional: `^_` and `^^`
The legacy MARC binary format still used for library data uses ascii 29, 30, and 31 -- although for reasons with probably some bizarre historical definition uses them DIFFERENTLY than defined in ascii.
0x1D == 29 == ascii group seperator == MARC Record separator
0x1E == 30 == ascii record seperator == MARC Field Terminator
0x1F == 31 == ascii unit separator == MARC subfield seperator
¹You can release Control and Shift after typing "U", in which case the character will appear after you type a space. Or, you can hold Control and Shift while typing the code point, in which case the character will appear after you release either modifier key.
[1] http://www.fileformat.info/tip/microsoft/enter_unicode.htm
Emacs makes it really easy to insert arbitrary unicode as well, even if you don't remember the code point. C-x 8 RET s n o w m a n RET will insert a ☃ (snowman).
But now it's like harping on the benefits of HDDVD or Bluray.
What you write is so true. So many large companies use text files to shuttle around data. I worked at one place that used pipe-delimited 10+GB files. It's not sexy, and using awk/sed/cut seems like a hack at first, then you realize that it works and it is the simplest solution to the problem.
You could use ESC (\x1b) to escape itself and either of your delimiter characters, but of course now you've gotten back all that complexity you were trying to avoid by using non-printable characters.
I had to deal with a system using TSV once, with that "feature", and that point made it so that we had to do the escaping at some higher level, with \t and \\.
You trade simplicity of parsing for rigidity of schema.
Something that the page doesn't mention is that CR+LF were originally two separate control codes because the action of returning the print head to the left hand side would take too long with a standard line printer. Therefore, separating the actions into two codes meant that the printer would not miss out any printable characters.
(At least, I read that somewhere on the internet and assumed it was true!)
The DOS line endings are inherited from previous systems, and while they precisely convey carriage and paper movement of a printer it can be a pain to deal with today.
Please add to that list also
HTTP = CRLF
http://www.w3.org/Protocols/rfc2616/rfc2616-sec5.html
However, I still have to see I single webserver that does not accept LF alone instead of CRLF.
A carriage-return operation takes much longer than a single character, or even two or three. It doesn't make sense to issue two characters just to take up time. The printers always had to have some internal buffer memory (and handshaking over the communication lines to say when the buffer is full) in order not to lose any characters.
True, but mildly redundant: "overprinting" was explicitly the purpose of 0x08 backspace (which had nothing, originally, to do with 0x7F deletion.)
Using CR, you'd need N2 + 1 characters.
Overprinted with 0x08: requires 3N characters
Overprinted with CR: requires 2N+1 characters
Last I checked, this still works even on laser printers (at least on a LaserJet), when sending data to it as plain text. It's not actually printing over itself, but it knows to make the repeated characters bold.
CSV on the other hand has a few different variations.
With CSV, you have to escape quotes, commas, and newlines. With tab delimited, you only have to escape tabs and newlines - and that's if they can't be sanitized out to begin with.
My own CSV parsers (I have written a few by now) usually parse that as if the space before (or after) the quote wasn't there. It's nonetheless something to avoid when writing CSV files (Postel's law, etc.).
I love this, it shows just how old the roots of ASCII are.
But your link does have more in-depth explanations (including historical info) for some of the control characters.
Just sayin'
Not to mention that one of the advantages of using a text record format at all is that you can view it using standard text viewers.
$ echo -n "1,2,3|4,5|6|7,8,9,0" | awk 'BEGIN{FS=","; RS="|"} {print NF, $0}'
3 1,2,3
2 4,5
1 6
4 7,8,9,0
In fact, you can also specify the output delimiters as well: $ echo -n "1,2,3|4,5|6|7,8,9,0" | awk 'BEGIN{FS=","; RS="|";OFS="foo";ORS="bar"} {print NF, $0}'
3foo1,2,3bar2foo4,5bar1foo6bar4foo7,8,9,0bar $ echo -n 'a^_1^^b^_2^^c^_3^^' |awk 'BEGIN{FS="^_"; RS="^^"} {print $1": "$2}'
a: 1
b: 2
c: 3It kind of works with standard unix tools.
cut -d$'\37' -f ...
sort -t$'\37' -k ...
join -t$'\37' ...
Those will parse ASCII-31-separated fields. But records are still newline separated, no way to change that AFAIK, short of running everything through tr '\036' '\n'
first. Which defeats the purpose of choosing "weird" delimiters in the first place.(Also note that the $'\..' syntax is bash-specific and doesn't exist in POSIX sh.)
If you need something that's never, ever going to break for lack of escaping, might I suggest doing percent-encoding (aka url encoding) on tabs ("%09"), newlines ("%0a") and percent characters ("%25")? Percent encoding and decoding can be made very fast, is recognizable to most developers, and can be used to escape and unescape anything, including unicode characters. Unlike C-escaping, which doesn't generalize and accommodate these things nearly so well.
So, strip them out of your data if you have to. If you think they need to be preserved or escaped, IMO you're doing something wrong.
They're often my tool of last and only resort when dealing with very large datasets. Sure, you could wait for that dump of all of wikipedia to import into a nice indexed and queryable database, but why not start grepping it immediately? Maybe you want to sort by a key that's textual, but there's satellite data that's non-textual. sort(1) is a pretty amazing program in terms of resource usage; it parallelizes, it makes efficient use of available memory and disk when merge-sorting.
Anyway, there are plenty of examples!
I agree that cut, sort, etc. are good to be familiar with. Someone else[1] linked a "csvquote" utility that pre-chews (and un-chews at the end of the text-processing pipeline) CSV data to make it work better with standard UNIX utilities. Looks neat, so I'll be keeping it in mind next time I'm processing CSV with UNIX utils.
Why? Writing a correct parser is not significantly harder than figuring out how to interface to an existing parser library, and allows cool things like heuristic parsing of malformed files.
OTOH it's shocking how many people can't write a correct CSV generator, even after being explicitly told what they're doing wrong (which is always either "you need to put quotes around the data" or "you need to double any quotes that are part of the data") and given examples.
Here are some surprising valid CSV files:
https://forge.ocamlcore.org/plugins/scmgit/cgi-bin/gitweb.cg...
The test program in the same directory shows the semantic content of each.
Interpreting it as a generalized escape causes two problems. One, if you generate files that way, they will be unreadable by parsers written according to the RFC. Two, if you read files that way, you will silently garble files generated by someone who forgot to escape the quotes that were part of their data.
The RFC declares itself "informational" and says things like there is no formal specification in existence, which allows for a wide variety of interpretations of CSV files. This section documents the format that seems to be followed by most implementations:
There are certainly reasonable arguments about the useful subset of rules.
The problem is that the data you will be given to parse, by some third party agency, will quite likely not have been produced in such a way to get all the corner cases right.
As anyone who routinely has to parse such data is unpleasantly aware of. So now you've got to manually fix data, or have a parser that _doesn't_ get the corner cases 'right' but instead uses heuristics to try to get what was intended out of your particular idiosyncratic and illegal data.
> Then you have a text file format that is trivial to write out and read in, with no restrictions on the text in fields or the need to try and escape characters.
Phrases like that lead to lovely security bugs.
Escaping should always be a consideration. Not thinking about it, thinking "it'll never happen", etc. is what leads to things like HTML and SQL injection vulnerabilities.
If you're inputting or outputting data in any format, always keep in mind things like "what are the delimiters? What if the data in the input/output contains them?"
Doesn't allow for tab-delimited or any-character-delimited text and handles "Quotes, Commas, and Tab" characters in fields.
Ascii table for IBM PC charset (CP437) - Ascii-Codes
They correspond to these Unicode characters
28 FS ∟ 221f right angle
29 GS ↔ 2194 left right arrow
30 RS ▲ 25b2 black up pointing triangle
31 US ▼ 25bc black down pointing triangle
They may not be particularly intuitive symbols for this purpose though.see also: IBM Globalization - Graphic character identifiers: http://www-01.ibm.com/software/globalization/gcgid/gcgid.htm... (then search for a code point, eg U00025bc)
Unicode code converter [ishida >> utilities]: http://rishida.net/tools/conversion/
As used by Lotus 1-2-3 and undoubtedly others before there was an Excel.
Example record:
42,"Hello, world","""Quotes,"" he said.","new
line",x
Now go write a little state machine to parse it... (hint: track odd/even quotes, for starters)1. Go to System Preferences => Keyboard => Input Sources
2. Add Unicode Hex Input as an input source
3. Switch to Unicode Hex Input (assuming you still have the default keyboard shortcuts set up, press Command+Shift+Space)
4. Hold Option and type 001f to get the unit separator
5. Hold Option and type 001e to get the record separator
6. (Hold Option and type a character's code as a 4-digit hex number to get that character)
Sadly, this doesn't seem to work everywhere throughout the OS- I can get control characters to show up in TextMate, but not in Terminal.
FS: Control-\ 0x1c (field sep)
GS: Control-] 0x1d (group sep)
RS: Control-^ 0x1e (record sep)
US: Control-_ 0x1f (unit sep)
(These control key equivalents have always been the canonical keystrokes to generate the codes)But they have to be preceded by a Control-V (like in vi) to be treated as input characters. Control-V is the SYN code (synchronous idle), but has no special meaning in an interactive context, which is presumably why it was chosen.
The full set of control codes (0x00 - 0x1f) and their historical meanings are why Apple added the open/closed Apple keys, eventually the Command key. They wanted a set of keystrokes that were unambiguously distinct from the data stream.
Control-S, e.g., will pause text output in the Terminal (also xterm, etc). This was super useful in the days before scrollback. :) Control-Q to resume (actually flush all the buffered output).
Overloading Control sequences was an unforgivable sin committed by Microsoft.
...if I remember the history correctly, Apple decided that having both open/closed Apple keys was confusing, and having the Apple logo on the keyboard was tacky, so they renamed the key for the Mac, and Susan Kare selected a new glyph, which is a Scandinavian "point of interest" wayfinding symbol.
...as a further aside, Control-N and Control-O are the cause of the bizarre graphical glyphs you sometimes see if you do something silly like cat a binary file. Control-N initiates the character set switch, and Control-O restores it. This can be used to fix your Terminal when things go awry. Most people just close the window, but I hate losing history. :)
0x20 - 0x74, unshifted:
!"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~
0x20 - 0x74, shifted: !"#$%&'()*→←↑↓/▮123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_◆▒bcde°±▒☃┘┐┌└┼⎺⎻─⎼⎽├┤┴┬│≤≥π≠£·
...works in Firefox. YMMV.Terminal-charset-quickfix: at shell, type "echo ^O". To get the literal ^O, use Control-V then Control-O.
Baudot4Life, yo.
Which, BTW, was an indication that your machine needed service. The spring should have been wound tight enough and the track clean & oiled well enough to get the carriage back to the first column in time to not drop any characters. A pneumatic piston ("dash pot") slowed the carriage down as it approached the first column so it wouldn't crash into the stops and get damaged.
For a trivial example, try building an ASCII table using this format, with columns for numeric code, description, and actual character. You'll once again run into the whole escaping problem when you try to write out the row for character 31.
(The problem reminds me of what it's like using "/" as a delimiter when using `sed` to edit file paths.)
I'm wary of anything that solves a problem partially while still remaining vulnerable in the end, because it can discourage properly solving the problem. Rather than play musical chairs with the separator character to try to minimize the chance of a conflict, I'd rather see a sensible encoding/escaping scheme used to eliminate that chance entirely.
For ASCII delimiters, forbidding ASCII delimiters in data is practical.
Sure - you can't, say, nest ASCII tables into one another due to this limitation.
But for simple structure, it doesn't hurt to have ASCII separators in the toolbox.
The only big problem I see is that they're rendered as invisible characters, which will make debugging harder. If we wouldn't have abandoned and forgotten the special ASCII chars, this wouldn't be the case.
If your dev tools show special chars (like mine do), then it's perfectly fine to use them.
In hindsight it's too bad we don't have similar characters that follow a more sexpr-ish layout - say, ListStart, ListEnd, and Delimiter. Then you could tree them endlessly. If you wanted to be really fancy you could add an "assignmentSeparator" character to officially bless key-value-pairs and encompass a nice JSON-ish format, but Lisp pretty-well demonstrates that isn't necessary.
But in hindsight it's just too bad we don't use these control characters at all.
The only reason to use delimiters, ever, is for user-modifiable data (e.g. source code) where you might want to insert or delete characters and have the containing block remain valid.
---
And now, a fun tangent, to prove that how deeply-rooted this confusion is in CS: user-modifiable data was originally the sole use-case for \0-terminated "C strings" in C.
C has two separate types which get conflated nowadays: char arrays, and \0-terminated strings. Most "strings"--as we'd expect to find them in other languages--were, in C, actually char arrays: you knew their length, either because they were string literals and you could sizeof them, or because you had #defined both FOO and FOO_LEN, or because you had just allocated len bytes on the heap for foo, so you could just pass len along with foo. Because you knew their length, you didn't need to use the string.h functions to manipulate them. It was idiomatic (and perfectly-safe) C, when dealing with char arrays, to just iterate through them with a for loop.
The concept of \0-termination, and thus what we think of as "C strings", only applied to string buffers: fixed-size, stack-allocated, uninitialized char arrays. The string.h functions are all meant to be employed to manipulate string buffers, and the \0 is intended to mark where the buffer stops being useful data, and starts being uninitialized garbage.
The strings in string buffers had short lifetimes, and didn't usually outlive the stack frame the buffer was declared in. Generally, you'd declare a string buffer, populate it using some combination of string literals, strcat(3), sprintf(3), and system calls, and then pass the string--still sitting inside the buffer--to a system call like fstat(2) to get what you're really after. That would be the end of the both string buffer's, and the string's, lifetime.
If you ever did want to preserve the contents of a string buffer into something you could pass around, though, this would be idiomatic:
int give_me_a_path_string(char **out)
{
char buf[MAX_PATH];
/* ... */
int len = strlen(buf);
*out = memcpy(malloc(len), buf, len);
return len;
}
Note that, after this function returns, the pointer it has written to doesn't point to a "C string": instead, it's a plain pointer to a heap-allocated array of char, with exactly enough space to hold just those characters. If you want to know how big it is, you look at the return value.So:
• C has "C strings", but they were only intended as buffers.
• C also has "char arrays", which are really what you should think of as C's equivalent to a "string" datatype. char arrays, not "C strings", are the fundamental data structure for representing and persisting strings in C.
• char arrays are less like "C strings" than they are like Pascal strings: they come in two parts, a block of memory N chars wide, and an int containing N. You don't examine the block to determine the length; the length is explicit.
• Pascal (and thus most modern languages with strings) put both the length and the character-block on the heap as a unit. C puts the character-block on the heap, but puts the length on the stack. This is more efficient under C's Unix-rooted assumptions: you need the length on the stack if you want to work with it to immediately shove the string through a pipe.
It's just these worse-is-better text-based protocols like HTTP, created by application developers, that toss all the advantages of length-prefixing away. (And, even then, HTTP bodies are length-prefixed, with the Content-Length header. It's just the headers that aren't.)
My favorite way to deal with this stuff is Consistent Overhead Byte Stuffing:
http://en.wikipedia.org/wiki/Consistent_Overhead_Byte_Stuffi...
In short, you take the data and encode it with a clever scheme that effectively escapes all the zero bytes. The output data contains no zeroes, but results in almost no overhead, with the worst case being an increase of 1/254 over the original size, and the best case being zero increase. (Compare to e.g. backslash escapes of quotes in quoted strings, where the worst case doubles the output size.) You then use the now-eliminated zero byte as your record separator. This lets you stream data (with a small amount of buffering to perform the encoding) while still easily locating the ends of chunks.
I've played around with COBS but never used it in a real product, so this is not entirely the voice of experience here. But it is a nifty system.
My team pretty much length prefixes everything. :)
Data formats that can't handle recursion are never practical. They only work until they don't, at which point they're entrenched and impossible to replace.
Unfortunately I cannot quickly find a better source than:
"The "escape" character (ESC, code 27), for example, was intended originally to allow sending other control characters as literals instead of invoking their meaning."
Due to the volume and novelty of data that we work with, we are often pushed into a corner between human time and machine time. Each data set comprises a new set of concepts, and each is huge. In this corner, sometimes a character-delimited file is the best solution. There is not time to carefully craft a binary format and then document it so it will not be forgotten later, nor is there time to wait for a general-purpose format parser to operate on tens of billions of records. We need a solution that can be designed in 1 minute and be legible by all of our tools without modification.
Typically, I have used tabs in the place of the ASCII separators. This ensures readability without any kind of parsing. Also, this lets me use the default behaviors of well-worn, bug-free tools in the core of the Unix toolchain for basic data processing tasks. Frankly, this is not a bad compromise.
If you are passing messages around a web stack, JSON, XML, and friends are ideal solutions. If you have to occasionally deal with CSV, use a parser. I just want to note that for many tasks in data analysis, it's OK to simply use the dead and dusted convention of mixed delimiters and data.
As these things develop, I will be trying to investigate how to use more modern formats such as binary JSON representations in my work, and I'd be curious what solutions people here suggest for working with very large data (e.g. many trillions of observations).
Ultimately, this is a language problem. If we invent new meta-language to describe data, we're going to use it when creating content. That means the meta-language will be used in regular language. Which means you're going to have to transform it when moving it into or out of that delimited file.
There is no fixed-length encoding you can use to handle meta-information without imposing restrictions on the content. You're always going to end up with escape sequences.
This isn't some clever new discovery. It's begging us to repeat the same mistakes that led to the world adopting printable ASCII delimiters in the first place.
Encodings aside, the principle of having a hierarchy of delimiters can be hugely powerful
https://github.com/dbro/csvquote
will convert all the record/field separators (such as tabs/newlines for TSV) into non-printing characters and then in the end reverse it. Example:
csvquote foobar.csv | cut -d ',' -f 5 | sort | uniq -c | csvquote -u
It's underrated IMO.It is indeed a simple state machine (see https://github.com/dbro/csvquote/blob/master/csvquote.c), and it translates CSV/TSV files into files which follow the spirit of what's described in the original article in this thread.
But instead of using control characters as separators, it uses them INSIDE the quoted fields. This makes it easy to work with the standard UNIX text manipulation tools, which expect tabs and newlines to be the field and record separators.
The motivation for writing the tool was to work with CSV files (usually from Excel) that were hundreds of megabytes. These files came from outside my organization, and often from nontechnical people - so it would have been difficult to get them into a more convenient format. That's the killer feature of the CSV/TSV format: it's readable by the large number of nontechnical information workers, in almost every application they use. I can't think of a file format that is more widely recognized (even if it's not always consistently defined in practice).
It sort of works, but there are known issues which I have listed in the README.
It's led to many a headache.
Reminding everyone that an unused, unloved standard exists is just reminding everyone that the hard part went undone.
The problem is, of course, that we can't see it (no glyph) and we can't "touch" it (no key for it) so people won't use it. Ultimately, we're all still stick-wielding apes.
1. Control characters are not supported in the almost of text editors. 2. Control characters are not human friendly. 3. The text may contain control characters in the field value.
In any formats, we cannot avoid the escape characters, so even I think CSV/TSV format is reasonable.
> CSV breaks depending on the implementation on Quotes, Commas and lines
CSV does not break; the implementation is broken if it doesn't parse CSV properly. With a proper implementation, CSV solves every problem that will arise from this method.
And then someone wrote an open source CSV parsing library that handles edge cases well and everyone forgot these characters existed.
The edge case bugs in ASCII codes could still crop up. It shouldn't, but then, valid SQL shouldn't crop up in a web form either. And when it does, we'll need escape codes just like CSV, only it won't be well tested in all the tools (because it's not going to frequently happen).
It's like all the OSS advocates laughing at Microsoft's idiotic "My Documents" folder. It's not there because they didn't realise how much trouble it would case programmers, it's there because they wanted for force people to deal with edge cases.
Of course, I was one of those odd ducks doing a lot of Classic ASP work with JScript at the time.
The only trouble I ran into was that the library doesn't like getting rid of your quote character[1], and I don't see an easy way around it[2].
That said, I really don't like this format. The entire point of CSV is that you have a serialization of an object list that can be edited by hand. Sure using weird ASCII characters compresses it a bit because you're not putting quotes around everything, but if you're worried about compression you should be using another form of serialization - perhaps just gzip your csv or json.
In Ruby in particular, we have this wonderful module called Marshal[3] that serializes objects to and from bytes with the super handy:
serialized = Marshal.dump(data)
deserialized = Marshal.load(serialized)
deserialized == data # returns true
I cannot think of a single reason to use ASCII Delimited Text over Marshal serialization or CSV.1. ruby/1.9.1/csv.rb:2028:in `init_separators': :quote_char has to be a single character String (ArgumentError)
A tab delimiter is not preferable as it is not visible, and can be problematic to parse via command line tools (ie what do I set as the delimiter character?).
I think that is the whole point of having ASCII delimited text files is to have human readable data in it.
C-v-shift-_ and C-v-shift-^ both work for me.
They print a little strangely, but if you were really dedicated to the idea, you could alias the tools you use to use these by default for their input and output separators.
No, please, put the gun down... let me explain. Sometimes you have a database that's so complex and HUGE that changing tables would be a nightmare, or you just don't have the time. You have a field that you want to shove some serialized data into in a compact way and not have to think about formatting. You could use JSON, you could use tabs or csv, but both of those require a parser.
With these ascii delimiters you can serialize a set of records quickly and shove them into a string, and later extract them and parse them with virtually no logic other than looking for a single character. And because it's a control character, you can strip it out before you input the data, or replace control characters with \x{NNN} or similar, which is still less complex than tab/csv/json parsing.
Granted, the utility of this is extremely limited, probably mainly for embedded environments where you can't add libraries. But if you just need to serialize records with the simplest parsing imaginable, this seems like an adequate solution.
I agree with several other comments that the biggest issue is not being able to represent them in an editor. If you use some form of whitespace, then it is likely to lead to confusion with the whitespace characters you are borrowing (i.e. tab and line feed). If you use special glyphs, then you have to agree on which ones to use, and it still doesn't solve the problem of readability. Without whitespace such as tab and line feed, all the data would be a big unreadable (to humans) blob, and with whitespace, it would lend confusion about what the separator actually is. Someone might insert a tab or a linefeed, intending to make a new field or record, and it wouldn't work. If the editor automatically accepted a tab or linefeed and translated it to US and RS, then there would have to be an additional control to allow the user to actually insert the whitespace characters that this is supposed to enable. :/
This isn't intended to be a data exchange format, it is a serial data storage format. In this way, there may be some valid usages, but modern file systems do not need this sort of representation and it has no real benefit over *SV formats for most use cases. I suppose It could still be used for limited exchange, but since it can't be used storing binary, much less Unicode (except for perhaps UTF-8), other formats are less ambiguous and more capable.
So on Mac OS X Mavericks: http://support.apple.com/kb/PH13867
1. Replace all the commas in the text with the unique non-printing char before converting to CSV.
2. Convert this char back to a comma when processing the CSV for output to be read by humans.
Because commas in text are usually followed by a space, the CSV may still even be readable when using the non-printing char.
I must admit I've never understood why others view CSV as so troublesome vis-a-vis other popular formats.
in: sed 's/,/%2c/g' out: sed 's/%2c/,/g'
I guess I need someone to give me a really hairy dataset for me to understand the depth of the problem with CSV.
Meanwhile, I love CSV for its simplicity.
sed does the job and on almost all UNIX clones it never needs to be installed.
Because it's already there.
Is this really a standard complaint about XML? I thought the main complaint was that it wasn't human-writeable. I wouldn't want to read novels in XML, but I've never had a problem opening up an XML file in a text editor to get at bits of it.
That can be fixed with `xmllint --format`; but I agree that, once you need to bring in external tools, it's not clear that calling it 'human-readable' is really appropriate any more.
http://en.wikipedia.org/wiki/Unicode_control_characters
Seems like the same control characters are present.
Anyone who got used to using ^[ for escape in vi would already be familiar with this approach.
File separator - C-v C-\
Group separator - C-v C-5
Record separator - C-v C-6
Unit separator - C-v C-7
They are all visible characters in both vim and emacs by default.You can see them on the terminal with `cat -v`
It would be nice if more tools were built to take advantage of these characters, but there are some that do.
M-x ucs-insert 1c
M-x ucs-insert 1d
M-x ucs-insert 1e
M-x ucs-insert 1f
for file, group, record, and unit separators. Is there an easier way? od -cIt's not often that the tab delimited format is problematic, at least nothing that a simple string-replace operation can't solve, so it's not worth trying to convince every existing text reader and text processors to recognize these long forgotten record separators correctly instead.
foo!bar!baz!
a!b!c!
d!!!
e!f!!The underlaying mapping formats for specific industries are a pain to parse but everything is easily formatted using stars or pipes as field separators
ST|101
NAM|john|doe
ADR|123 sunset blv|sunrise city|CA
DAT|20140326|birthday
Ah, the joy of simplicity.Just grabbing the first few segments from the example message in the wiki article:
MSH|^~\&|MegaReg|XYZHospC|SuperOE|XYZImgCtr|20060529090131-0500||ADT^A01^ADT_A01|01052901|P|2.5
EVN||200605290901||||200605290900
PID|||56782445^^^UAReg^PI||KLEINSAMPLE^BARRY^Q^JR||19620910|M||2028-9^^HL70005^RA99113^^XYZ|260 GOODWIN CREST DRIVE^^BIRMINGHAM^AL^35209^^M~NICKELL’S PICKLES^10000 W 100TH AVE^BIRMINGHAM^AL^35200^^O
The delimiters are defined at the beginning of the opening MSH (message header) segment. HL7 is zero-indexed, but your zero index is always the segment label, so it's easy for non-technical people to count naturally to get the field identifier without having to explain counting n-1 to them.The one exception to that is the MSH segment. Things get a little screwier there because the first instance of the field delimiter is also counted as a full field in the spec, so it tends to trip people up. So even though "^~\&" above looks like it should be MSH.1, it's actually MSH.2, etc.
The delimiters used in the wiki example are the most common you encounter, but some systems do things differently because reasons. The primary HIS at my hospital uses colons and semicolons, for example (and I want to poke out my eyes with ice picks every time I have to look at the messages coming from it as a result). But since it's all defined right in the message header, it's trivial to convert between delimiters when you need/want to.
Either way, this is how the vast majority of electronic medical records are transmitted today.
Of course xml and json do the same, but more verbose.
ASCII formats like EDI were invented when every byte in transmission counted.