Things Every Hacker Once Knew
catb.org
catb.org
It becomes immediately obvious why, eg, ^[ becomes escape. Or that the alphabet is just 40h + the ordinal position of the letter (or 60h for lower-case). Or that we shift between upper & lower-case with a single bit.
esr's rendering of the table - forcing it to fit hexadecimal as eight groups of 4 bits, rather than four groups of 5 bits, makes the relationship between ^I and tab, or ^[ and escape, nearly invisible.
It's like making the periodic table 16 elements wide because we're partial to hex, and then wondering why no-one can spot the relationships anymore.
The original 1963 version of ASCII covers some of this; a scan is available online. See also The Evolution of Character Codes, 1874-1968 by Eric Fischer, also easily found.
(It's also the only non-control ASCII character that can't be typed on an English keyboard, so it's good for creating WIFi passwords that your kid can't trivially steal.)
Don't count on it. There's a fairly long standing convention in some countries with some keyboard layouts that Control+Backspace is DEL. This is the case for Microsoft Windows' UK Extended layout, for example.
[C:\]inkey Press Control+Backspace %%i & echo %@ascii[%i]
Press Control+Backspace⌂
127
[C:\]
This is also the case for the UK keyboard maps on FreeBSD/TrueOS. (For syscons/vt at least. X11 is a different ballgame, and the nosh user-space virtual terminal subsystem has the DEC VT programmable backspace key mechanism.)Now replace "X" with "Delete".
Another good source on the design of ASCII is Inside ASCII by Bob Bemer, one of the committee members, in three parts in Interface Age May through July 1978.
https://archive.org/details/197805InterfaceAgeV03I05
I do understand that I've probably simplified "how I understand it" vs "how/why it was designed that way". This is pretty much intentional - I try to find patterns to things to help me remember them, rather than to explain any intent.
> on a Vic-20
Which, weirdly, used the long-obsolete ASCII characters of 1963–1967, with '↑' and '←' in place of '^' and '_'.I am not following, can you explain why ^[ becomes escape. Or that the alphabet is just 40h + the ordinal position? Can you elaborate? I feel like I am missing the elegance you are pointing out.
00 11011 is Escape
10 11011 is [
So when we do ctrl+[ for escape (eg, in old ansi 'escape sequences', or in more recent discussions about the vim escape key on the 'touchbar' macbooks) - you're asking for the character 11011 ([) out of the control (00) set.Any time you see \n represented as ^M, it's the same thing - 01101 (M) in the control (00) set is Carriage Return.
Likewise, when you realise that the relationship between upper-case and lower-case is just the same character from sets 10 & 11, it becomes obvious that you can, eg, translate upper case to lower case by just doing a bitwise or against 64 (0100000).
And 40h & 60h .. having a nice round number for the offset mostly just means you can 'read' ascii from binary by only paying attention to the last 5 bits. A is 1 (00001), Z is 26 (11010), leaving us something we can more comfortably manipulate in our heads.
I won't claim any of this is useful. But in the context of understanding why the ascii table looks the way it does, I do find four sets of 32 makes it much simpler in my head. I find it much easier to remember that A=65 (41h) and a=97 (61h) when I'm simply visualizing that A is the 1st character of the uppercase(40h) or lowercase(60h) set.
Thank you!
Physically it might have been as simple as a press-open switch on the original hardware, each bit would be a circuit which the key would connect, the SHIFT and CONTROL keys would force specific circuits open or closed.
the letters in the third column are A = 1, B = 2 etc: 40h + the position in the alphabet.
Awesome to see ^@ as null and laying it out this way makes it easier to see ^L (form-feed, as the article says: control-L will clear your terminal screen), ^G (bell), ^D, ^C etc etc
The 40h offset is 2 columns' worth.
Cool story bro, I know, but I meant to put the file online in response here, but I can't find the source doc anymore >_< Edit: actually, I found an old incomplete version as a Apple Numbers file. If there's interest I can whip it back up into shape and post it as PDF.
[1] For example, when a unix C program outputs "\n", it's the terminal device between the program and the user TTY that translates it into \r\n. You can control that behavior with stty. I know this is something ESR would laugh at being novel to me. On bare-metal, you have no terminal device between you program and the UART output, so you need to add those \r\n yourself.
[2] That's ESC "escape" at decimal 26/hex 1B, and you can generate it in a terminal by pressing the Escape key or Ctrl-[
It's entirely possible that someone reading this thread will be able to source it.
But it's just a text table, right? It should be fairly trivial to reproduce it from a decent picture.
By the way, thanks for clarifying the existence and purpose of the now-in-kernel "terminal device." I've understood the Linux PTS mechanism and am aware of the Unix98 pty thing and all of that, but identifying it like that helps me mentally map it better.
Speaking of DOS, I've never forgotten that BEL is ASCII 7, so ALT+007 (ALT+7?) will happily insert it into things like text editors. I remember it showed up as a •. I'm not quite sure why.
Both shortcuts work in terminal emulators.
(As an aside ^W is much easier to input^Wtype as a "fake backspace" thingamajig^H^H^H^H^H^H^H^H^H^H^Hmnemonic than ^H is.)
If you work on embedded devices you will still encounter serial/RS-232 all the time. Often through USB-to-serial chips, which only adds to the challenge because they are mostly unreliable crap. Then there are about 30 parameters to configure on a TTY. About half do absolutely nothing, a quarter completely breaks the signal giving you silence or line noise, the final quarter only subtly breaks the signal, occasionally corrupting your data.
Still, there is nothing like injecting a bootloader directly to RAM over JTAG, running that to get serial and upload a better bootloader, writing it to flash and finally getting ethernet and TCP/IP up.
Still, there is nothing like injecting a bootloader directly to RAM over JTAG, running that to get serial and upload a better bootloader, writing it to flash and finally getting ethernet and TCP/IP up.
I'm happy to have gotten mostly rid of this. Gone are the days of choosing motherboards based on the LPT support, and praying that the JTAG drivers would work on an OS upgrade.
It's still there, and not going anywhere - the only thing that's past is LPT. Last time I used a Linux-grade Atmel SoC, it had a USB-CDC interface but the chain was still the same: boot from mask ROM, get minimal USB bootloader, load a bootstrap binary to SRAM, use that to initialize external DRAM, then load a flashing applet to DRAM, run it, use that to burn u-boot to flash, and then fire up u-boot's Ethernet & TFTP client to start a kernel from an external server and mount rootfs over NFS. Considering the amount of magic, it worked amazingly well. The whole shebang was packaged into a zip file with a single BAT to double-click and let it do the magic.
As for COM and LPT - FTDI and J-Link changed the embedded landscape forever, and thanks for that.
disclaimer: I'm friends with the BMP manufacturer.
So your sensors can talk across the internet.
From `man socat', which runs on Linux:
(socat PTY,link=$HOME/dev/vmodem0,raw,echo=0,wait-slave
EXEC:'"ssh modemserver.us.org socat - /dev/ttyS0,non-block,raw,
echo=0"')
generates a pseudo terminal device (PTY) on the client that can
be reached under the symbolic link $HOME/dev/vmodem0. An appli-
cation that expects a serial line or modem can be configured to
use $HOME/dev/vmodem0; its traffic will be directed to a modem-
server via ssh where another socat instance links it with
/dev/ttyS0.now kremvax' network stack is bound over my own no need for nat
By "bound over", do you mean tunneled, or "made available alongside"?
I think I'm right in sayig that if one has /net and /net.alt then /net is asked first and then /net.alt
I'm a bit rusty on the details
To make a connection a process open's its /net/tcp/ctrl file, writes a connection string, gets back a response. If that response is a number then it opens /net/tcp/1/ctrl and /net/tcp/1/data and read / writes data to the data and sees out of stream messages on ctrl. To close the connection one closes /net/tcp/1/ctrl
Everything is done via the 9p protocol. So if I can write code on my Arduino that understands 9p and a way to send/rec that data (e.g. serial but even SP1 or JTAG) to a Plan9 machine locally, then the Arduino itself could just open a tcp connection on kremvax.
Once you start layering these things on top of each other it gets a bit mind blowing and you get drunk on power. Then when you are forced back Linux or Windows you realise how dumb they are even with their fancy application software.
And nice. You talk to the kernel via plain text. No ioctls! :D It's like they made an OS designed to make bash happy.
And I can sum up my conclusions of 'modern' Linux in one word: systemd. It's almost worse than Windows now. I'm not surprised this kind of thing doesn't work on Linux :P
...But I'm sad now, that I can't do all of this on Linux.
Hrm. Maybe someone should pull the Plan9 kernel-level stuff out and make a FUSE-based thing (or even a kernel driver!) that emulates all of this on Linux. There's already plan9port for the utilities... okay, they'd need to be modified back to being Plan9-ey again, but it could be really interesting.
Does anyone know the reasons why USB was not made backward compatible with RS-232? It would only take a very short negotiation to determine if both endpoints support USB.
To expand a bit, the typical fully compliant RS232 setup has a special level translation IC (for example the MAX232) to deal with the huge voltage range and convert it to something more low voltage digital logic friendly. And that IC usually requires power supply levels that the typical cheap USB device would not have, which means it would have to add more circuits to generate them, adding even more cost to the hardware.
And the USB committee probably figured to be not fully compliant with RS232 would just confuse people, and damage hardware, so it was better to be not compliant at all.
But...
> And the USB committee probably figured to be not fully compliant with RS232 would just confuse people, and damage hardware, so it was better to be not compliant at all.
...you are sadly right.
Man that job was frustrating, but it sure was a lot of fun!
It's a mess at the moment, because there are unauthorized clones of both FTDI and Prolific, and both companies release drivers that purposefully don't work (or worse...brick them) on the clones. But, there's not really a way for the end buyer to know for sure they are buying the real thing.
I use them because they'll go down to 45 baud for antique Teletype machines. They're popular for Arduino applications, and there are lots of cheap breakout boards with 0.100 pins for Arduino interfacing.
I guess on the surface the big thing I really like is device differentiation. Do CP2102Ns have unique serial numbers, or can that free utility burn in info I can use to differentiate?
Going a bit deeper, can I bitbang with it?
Thus, you can force the host machine to demand a device-specific driver if you need to. By default, it appears to the OS as a USB to serial port device. Linux and Windows recognize it as such, without special drivers. Linux mounts it starting at /dev/usb0; Windows mounts it starting at COM3.
No bit-banging, though; it doesn't have the hardware.
(For an interesting side story, search for "ftdigate". At some point FTDI decided they're fed up with copycats and released a new driver pack (through automatic Windows Update) that bricked counterfeit chips. This led to a lot of angry people and number of amusing situations, including someone jokingly submitting a patch to Linux to do the same.)
And with the emergence of USB-CDC, it's not nearly as necessary as it used to be, since most modern OSes support that now.
We would.. but since they already cheated, we had to switch suppliers.
... whaaa? They're probably the biggest company in that market, closely followed by Prolific (PL2303).
> ...and their chips tended to implement CDC in variously broken ways
No, their chips implement a proprietary protocol. Not CDC at all.
What I can add for everybody who feels the same disappointment as ESR: It's very common for a growing community that three things happen.
A) The number of people with just a little knowledge over the holy grail of your community increases.
B) The popular communication is taken over by great communicators who care more about their publicity than your holy grail.
C) This gives the impression that the number of really cool people decreases. And that is depressing to old timers. But it's in fact often not true. Actually most often the number of cool people increases too! It's just that their voices are drowned in all the spam of what I like to call the "Party People" (see B).
So yes, you can actually cheer. It's harder to find the other dudes, but there are more of them! Trust me, I'm not the oldest guys here but I've seen some communities grow and die till now, and it's nearly always like that.
D) the B)s use various "social" tactics to tar and feather A)s that get in their way...
I'm not suggesting the article is a, "Gosh, Millenials!" conversation. I just get a warm tingle when reminded that I have absolutely no clue how to do something people did just a generation ago, and I don't need to. It's success!
The videos may look a bit dated now, but the content is amazing and Burke is terriffic.
https://en.wikipedia.org/wiki/James_Burke_(science_historian...
It once really was The Learning Channel
The best part is perhaps Burke himself, and his very very British presentation style.
At least that's what I comfort myself with as I gaze over my stash of DB25 connectors and Z80-SIO chips...
---
For those curious like myself (heavily elided for smaller wall of text):
The Z80-SIO (Serial Input/Output) ... basic function is a serial-to-parallel, parallel-to-serial converter/controller ... configurable ... "personality" ... optimized for a given serial data communications application.
The Z80-SIO ... asynchronous and synchronous byte-oriented ... IBM Bisync ... synchronous bit-oriented protocols ... HDLC and IBM SDLC ... virtually any other serial protocol for applications other than data communications (cassette or floppy disk interfaces, for example).
The Z80-SIO can generate and check CRC codes in any synchronous mode and can be programmed to check data integrity in various modes. The device also has facilities for modem controls in both channels, in applications where these controls are not needed, the modem controls can be used for general-purpose I/O.
What's really interesting is this bit:
• 0-550K bits/second with 2.5 MHz system clock rate
• 0-880K bits/second with 4.0 MHz system clock rate
110Kbps at 4MHz. That's almost the 115200 baud we're all familar with. At 4MHz! (2.5MHz yields 68750bytes/sec, or 67.13Kbps.)
https://archive.org/download/Zilog_Z80-SIO_Technical_Manual/...
Also - rather amusingly, I discovered that the MK68564 datasheets ripped off the Z80-SIO's intro text verbatim. Are these compatibles or something completely different? https://www.digchip.com/datasheets/parts/datasheet/456/MK685...
It'd be like if we moved from using horses to using giant machines made out of glued-together living horses, and then all the horse-tenders died.
Abstracting, yes, but I don't know about building-upon. The thing is, a lot of this stuff is still sitting around beneath the covers, and someone needs to understand it.
Even worse, sometimes there's stuff that's abstracted over that's important, e.g. if the Excel or 1-2-3 teams had known about Field/Group/Record/Unit Separators, would they have ever come up with CSV?
Or the fact that SYN can be used to self-syncronise, due to its bit pattern …
Neither team had anything to do with CSV which originates in Fortran's list-directed I/O. And there is no field separator (which would be redundant with unit separator), FS is the file separator, these codes were intended for data and databases over sequential IO, in modern parlance a group is a table.
I have no idea how my car works. I mean, I more or less understand the principles underlying the internal combustion engine, but I wouldn't be able to service one, much less assemble one. But I don't need to. Typically, the only indication made available to me that something is wrong is a single bit of information ("check engine light"), but that is enough. You don't have to be a "car person" to make effective use of a car. I get in, I go, and well over 99% of the time that's the end of the story.
Compare this with computers. When something goes wrong, it's usually vital that you (or your users) relay the precise error message (and God help you if there isn't one). You generally have to be a "computer person" to some degree to make effective use of a computer. If you are unconvinced by this comparison, contrast how often your family asks you to perform [computer task] versus how often you would approach a mechanic family member to perform [car task]; Contrast how often you hear "I can't do this, I'm not a car person" versus "I can't do this, I'm not a computer person".
I consider swaths of modern hackers who simply don't know about much of ASCII as evidence of babysteps towards computers maturing as a technology.
First, we use computers for so many more things than cars. The average user does really well with basic tasks like checking their email and simple word processing. This would be daily driving in your car analogy. Occasionally things blow up, but that isn't too different from a major problem with a car. However users are constantly trying new things with computers, new programs, websites, and tasks. Car-owners who are constantly trying new things with their cars have as many problems, if not more, than the average computer user. The difference is that the people who use the full range of their car's capabilities are deeply interested in their vehicles.
Second, abstractions like the check engine light are far from perfect. How do you know whether the light signals imminent failure or a minor inconvenience? What additional information is needed for the mechanic to diagnose the problem? I recently chased down a problem in my car that caused the check engine light to come on with a code that was physically impossible. It took a few weeks of careful experimentation and instrumentation before I was able to figure out what it thought was going on. This was a case where I absolutely needed more than a cursory knowledge of how my car works.
I also think that a hacker should be similar to a amateur mechanic: although their car might be fuel-injected, they have a cursory knowledge of how a carburetor works. They may have an automatic transmission, but they understand what a clutch is. Compare that to many developers who have never set foot outside their niche; They have never used a radically different programming language or a different OS. They've never taken the time to dig into the layers beneath the one they use. I would argue that is a weakness. How will you ever debug a problem when it inevitably occurs in the layers beneath you?
I find this weird. As I proceeded through Comp sci in high school, going from Pascal to C to assembler, I was always troubled by my lack of understanding, "but why does it work?" That anxiety finally disappeared in college when I learned how logic gates are constructed and went through the exercise of implementing multiply as a series of logical operations.
Similarly, I find it strange that someone would be comfortable driving a car without fairly deep knowledge of how it functions and how to repair it. I don't understand how you're not plagued with anxiety.
I think this is more about social norms/conventions than anything. I would never ask a family member who happens to be a surgeon to remove my gall bladder for me, or a family member who happens to be a mechanic to replace my clutch over the weekend. But for some crazy reason, it's perfectly acceptable to ask your "computer person" family members to spend hours removing the 1200 malware infections you got by installing that cute puppy toolbar you downloaded. The complexity of the tasks doesn't have anything to do with it.
Nobody else has argued on this point yet so I'll throw my 2¢ in. (An aside: I had to look up the codepoint for ¢ - 2A - so I could use it. I haven't memorized ASCII yet, let alone Unicode.)
In my opinion, things like the first 31 characters of ASCII, line discipline, NIC PHY AUIs, the difference between RS-232 (point-to-point) vs RS-485 (current-loop), how to make your Classic Mac show a photo of the developer team (hit the Interrupt key then input "G 41D89A"), or how to play notes on period VT100s (set the keyboard repeat rate really high); we've moved into an era where Web stacks reveal hard-to-diagnose bugs in nearly-40-year-old runtimes (Erlang), Apple will give you $200k if you extract the Secure Boot ROM out of your iPhone (in one person's case via a bespoke tool that attached to the board and talked PCIe), the UEFI in Intel NUCs is such a close match for the open-source BSP that Intel releases that it's a lot easier than everyone would like for you to make UEFI modules that step on things that shouldn't be step-on-able and let you fall through holes into SMM (Ring -2), and few people care that sudo on macOS doesn't really give you root-level privileges anymore.
We've just replaced all the old idiosyncrasies with a bunch of more modern idio[syn]crasies. You're right that technology has matured, but this has unfortunately meant that a lot of the innocence we took for granted has been lost. Things aren't an absolute disaster, but it's more political now, and we have to keep on our toes. Computers aren't universally somewhere we can go to to have fun; we have to work to find the fun now.
(Also, I just did that thing I often do with forums - I expect my reply to appear at the end of the thread, so I go to click on the reply button at the end. But that would reply to the comment at the end of the thread, not yours. This is a problem endemic to forum UI and not a HN issue. Not all aspects of computers have matured yet, not by a long shot.)
I guess it depends on where one live, as i see that happen all the time (either family, friends, or neighbors).
The ACARS[0] protocol I work with every day starts each transmission with an SOH, then some header data, then an STX to start the payload, then ends with either an ETX or an ETB depending on whether the original payload had to be fragmented into multiple transmissions or fits entirely into one.
These codes aren't archaic and obsolete in the embedded avionics world.
[0] ACARS: "Aircraft Communications Addressing and Reporting System" - see ARINC specification 618[1]
Why is DEL's bit value 0xff (or 0255)? Because there was a gadget out there for editing paper tape. Yes. You could delete a character by punching out the rest of the holes in the tape frame. I used it once. It was ridiculous.
Was there anything that created these lace cards as part of normal operation? I'm guessing not, considering the ramifications.
What about programs that did this when they encountered bugs?
Because back then, people would write on wooden tablets covered with wax.
So.. next time you see a kludge and think of "historical reasons", consider that "historical" goes back much farther than 20th century. :)
The intent was that, when reading tape, DEL is ignored, because it's a position that was punched over with DEL, i.e. deleted.
Also, when punching tape, a key will punch its holes and (naturally) advance to the next character position. For DEL, that means you erase the character under the cursor, and then the next character is under the cursor. That is, it's a ‘forward’ delete. (I think it was DEC's VT2x0 series terminals that screwed that up for everyone.)
It's kind of appropriate that a typical TECO command line looks a lot like transmission line noise. :)
The archive tapes were saturated with insecticide, so bugs would not be inclined to chew up your stored info.
There were also separate mechanical duplicators, plus multi-layer tape so that ordinary terminals could make more than one copy in real time.
For RS-232 electrical reliability, it's hard to beat a design intended to allow any pins to be connected to + or - 25 volts or ground in any combination without doing any damage to the equipment at either end.
Plus not restricted to the minuscule cable-length specifications of USB, and to reach the rest of the conected (by phone) world, the same EIA open-source digital protocol was just modulated/demodulated to analog audio upon send/recieve.
Remember RS-232 was always expected to be at least building-wide if not site-wide, depending on the size of the site.
Ordinary data communication at relatiely slow speeds has benefits that might as well be taken advantage of when they are needed.
Of course no code was absolutely required for any of these processes, but you could still seamlessly share ASCII files between Apple, Commodore, DOS PC's etc. using native COM port commands.
To get more speed between two points, on early PC's you could get software to multiplex more than one COM port to handle a single data stream over multiple signal pairs.
When needed, this would require and tie up multiple phone lines to reach off-site but it worked, plus it was the same technique on your own local copper but then it was more feasible to be always on.
Many buildings were originally equipped with top-quality AT&T/Bell copper pairs each dedicated to a separate signal for each (prospective) phone line to each office through its on-site relay box. At the time many of these pairs were rapidly becoming idle with the arrival of the modern office multiline phone which ran on fewer pairs, or had its own dedicated wiring installed at deployment.
With Windows 9x, COM port multiplexing was built into Windows, and with the arrival of the 115Kbaud UART's you could theoretically get 460Kbaud between offices by running four 3-conductor DB9 cables from the phone access plate on the nearest office wall, and using a PC having connectors for the full 4 COM ports which had become standard on motherboards.
DB series are the size of old 'parallel' ports. DE are the common width, like DE15 for vga.
>Standard RS-232 as defined in 1962 used a roughly D-shaped shell with 25 physical pins (DB-25), way more than the physical protocol actually required (you can support a minimal version with just three wires, and this was actually common). Twenty years later, after the IBM PC-AT introduced it in 1984, most manufacturers switched to using a smaller DB-9 connector (which is technically a DE-9 but almost nobody ever called it that)
>Almost nobody ever called it that
BTW, I use Form Feed (CTRL+L) character in my code to divide sections, and have configured Emacs to display them as a buffer-wide horizontal line.
PostgreSQL has an interesting approach to this problem that I've found really straight forward and allows me to express text as text without getting into strange characters. What they've done is allowed using a character sequence for quoting rather than relying on a single character. They start with a character sequence that is unlikely to appear in actual text: $$, it's called dollar quoting. Beyond just $$, you can insert a word between the $$ to allow for nesting. Better explained in the docs:
https://www.postgresql.org/docs/current/static/sql-syntax-le...
What the key here is that I am able to express string literals in PostgreSQL code (SQL & PL/pgSQL) using all of the normal text characters without escaping and the $$ quoting hasn't come with any additional cognitive load like complex escaping can (and before dollar quoting, PostgreSQL had nighmareish escaping issues). I wish other languages had this basic approach.
my $str1 = qq!This is "my" string.!;
my $str2 = qq(Auto-use of matching pairs);
$str2 =~ qr{/url/match/made/easy};
The work I did with Perl included a LOT of url manipulation, so that qr{} syntax was really helpful in avoiding ugly /\/url\/match\/made\/hard/ style escaping.(Of course, this may also be one of the reasons that programmers in its broad language family have a pronounced tendency to shoehorn too many problems into complex string manipulation, but I suppose no capability comes without its psychological costs.)
[1]: http://cvs.schmorp.de/vt102/vt102 (note - contains VT100 ROM as binary data, but opens in browser as text)
print '-', substr(<<EOT, 0, -1), '!\n';
Hello, World
EOT
Prints: -Hello, World!
iirc sh-shells also have that. str1 = f"""This is "my" string."""
str2 = """Auto-use of matching pairs"""
str3 = r"""/url/match/made/easy"""Your question about self-referential documents and linking I don't understand; maybe an example. The PostgreSQL dollar sign quoting feature is simply a way to use single quotes (important in SQL) without having as many escaping issues. So instead of:
SELECT 'That''s all folks!';
You could write: SELECT $$That's all folks!$$;
or SELECT $BUGS$That's all folks!$BUGS$
And where it starts to save you in PostgreSQL is with something like (PL/pgSQL): DO
$anon$
BEGIN
PERFORM $quote$That's all folks!$quote$;
PERFORM 'Without special quoting';
END;
$anon$;
Note: this code produces nothing, it just should run without error (I ran it on PostgreSQL 9.4). In PL/pgSQL, the body of the procedural code is simply a string literal... but that means any SQL related single quoting would have to be escaped if we used single quotes. So using normal single quotes the previous code example would look something like: DO
'
BEGIN
PERFORM ''That''''s all folks!'';
PERFORM ''Without special quoting'';
END;
';
And it gets worse as you get into less trivial scenarios... which is why I suspect this dollar quoting system was created to begin with.There are a couple ways to handle depending on the scenario. If I were dealing with a static text under my control, say a direct insert of the text, I would either just enclose it all in traditional ' characters or come up with some unique quote text between the $$.
If I'm dealing with arbitrary text coming from, say a blogging website, I would either handle traditional SQL escaping in my input sanitizing code (or thereabouts) since I have to do that anyway ($$ is great for handwritten code where escaping introduces cognitive load, but not necessarily important for machine generated code) or I might create an inserting PL/pgSQL function with the article text as a parameter... that will get escaped without my having to do anything assuming I simply insert the text directly from the parameter.
NULL is actually in-band, not out-of-band, and in fact it illustrates the issues with in-band communication you mention. That's what, presumably, ESC was for: a way to signal that the following character was raw and did not hold its normal meaning.
You can still devise a pretty good protocol with straight ASCII over a wire, using SYN to synchronise the signal, the separator characters to separate data values and ESC to escape the following character (like '\' is used in many programming languages).
Yes, ESC is a code extension mechanism: it means that the following character(s) are not to be interpreted according to their plain ASCII meaning, but some other pre-arranged meaning. Ultimately a shared alternate meaning for terminal control was standardized as ISO 6429 aka ECMA-48 aka “ANSI”. Free reading: https://www.ecma-international.org/publications/standards/Ec...
That this gave us keyboards with an Escape key that GUIs would repurpose to mean ‘get me out of here’ is a coincidence. (Plain ASCII had Cancel = 0x18 = ^X for that.)
MIT culture for historical non-ASCII reasons also referred to Escape as ‘Altmode’, which is ultimately how EMACS and xterms ended up with their Alt-key/ESC-prefix clusterfjord.
ESC is used for introducing C1 control sequences
Also line printers using standard paper were 132 columns across and 66 lines down, which was 11 inches at 6 LPI. This matched the US portrait paper height and allowed for about 10 character per inch plus margins for the tractor feed and perforations.
("relatively harmless" because the paper wasn't actually printed on or otherwise damaged -- the operator just had to refold it and move it back to the input side).
They actually display fairly well in vim
while (my $item = <>) {
...
}
By default this will read lines from STDIN, but if you set $INPUT_RECORD_SEPARATOR to the RS character, it'll read a whole record at a time. You can also set $LIST_SEPARATOR to the FS character, which you can use with the split function to divide your record into fields, and the join function will use it automatically to turn your list of fields back into a record.I used these characters and Perl features in the early 2000s for managing a data processing workflow. The data was well-suited to being processed as a stream of text, and by using these separator characters I was able to avoid the overhead of quoting and escaping which made the processing MUCH more efficient.
Oh, and Ruby copied Perl's `$/` and `$,` for record and field separators. (The "English" module provides $INPUT_RECORD_SEPARATOR and $FIELD_SEPARATOR.)
I've seen formfeeds used in source files like that before; it's a simple way of getting some primitive wider navigation (Emacs has built in commands for navigating by page), but I don't know why you'd limit yourself to them when there are much smarter options available now.
* ASCII codes for those single and double box characters, so I could draw a fancy GUI on those old IBM text monitors
* Escape codes for HP Laserjets and Epson printers for bold, compressed character sizes etc.
* Batch file commands
* Essential commands for CONFIG.SYS
* Hayes modem control codes
* Wordstar dot format commands
* WordPerfect and DisplayWriter function keys
* dBaseII commands for creating, updating and manupulating records
I wish they would all move out of my head and leave room for me to learn some new stuff quicker!
ATD*99***1#
Even if the underlying technology has completely changed, the interface has not.> USSD code running...
then
> Connection problem or invalid MMI code.
So what was my phone just trying to do?
I'm curious what a "USSD code" is, and what kinds of codes I could input into the dialpad that do interesting things.
(I'm already aware of standard telco features like call forwarding, and I know iOS and Android both have their own set of "easter eggs" accessible via the dialer. I'm talking specifically about non-secret codes that talk to the baseband and stuff like that, if this sort of thing exists.)
Either way, the power led is useless, so better to hook it up to the speaker header ;)
Presumably not ASCII - they'll be from some extended IBM character set (CP437?), I would think.
<esc>&l1O to switch to landscape :)
I'm just about to write some code to parse HP PJL so you never know when ...
* Lotus 1-2-3 "slash" commands
Sixty years later and we're still bounded by legacy. As a result of shortage the general-use squawk codes are namespaced into each national ATC region, so aircraft have to constantly change squawks even in supposedly contiguous regions such as Eurocontrol. Squawk 4463 means a different thing in UK airspace than in French.
Ironically, military aircraft still support Mode 3 in order to integrate with civilian ATC, who call it Mode-A, but all their special don't-shoot-me-I'm-friendly stuff is handled by more modern encrypted protocols.
[0] https://en.wikipedia.org/wiki/Automatic_dependent_surveillan...
[1] http://www.rtl-sdr.com/adsb-aircraft-radar-with-rtl-sdr/
But DOS was developed on/for systems with CRT displays.
It doesn't really bother me, but every now and then this strikes me as peculiar.
Dealing with raw vs cooked (where LF is automatically translated to CRLF) ttys in UNIX is also a giant pain when you have to do it, so in a way it's not surprising that they decided to leave that out. The original DOS kernel was very minimal compared to UNIX even of the same era. Of course, it turns out that having to write CRLF into files is also a pain - Windows has binary and text mode files instead of raw and cooked mode ttys - and one that you encounter much more often.
* http://jdebp.eu./Softwares/nosh/italics-in-manuals.html
I wrote a better manual page for ul(1) that explains some of this. ul is basically a TTY-37 to your-terminal-type converter, and it implements a lot of the effects that one would see on a real teletype. Unfortunately, the old manual hasn't progressed much beyond the original 1970s one and doesn't explain a lot of the functionality that the program actually has.
On the other hand, there are far fewer use cases for LF without CR, certainly nothing that isn't better done using ANSI codes.
LF without CR is something that one would do on a typewriter for typing tabulated data or mathematical formulae. It's just a way to go "down" but stay at the horizontal position you were previously at.
https://www.google.com/search?q=python+universal+newlines+mo...
Edit: added the PEP (it's from 2002) and excerpt from it:
https://www.python.org/dev/peps/pep-0278/
This PEP discusses a way in which Python can support I/O on files which have a newline format that is not the native format on the platform, so that Python on each platform can read and import files with CR (Macintosh), LF (Unix) or CR LF (Windows) line endings.
That seems to be the reason for a lot of odd design choices. ;-)
Aw man… I'm only 36, but now I feel old for growing up in a time where a typewriter was still common enough to run into (even if they were rapidly being displaced by personal computers).
They still exist in the wild though as a hipster accessory — they probably do well on Instagram too I suppose.
My grandmother had an old manual typewriter, which I had to use once to type up some homework when I was in High School.
I do not miss them.
It boggles my mind why that hasn't been replaced by a PDF form. Perhaps IT being siloed in another building keeps such legacy going.
Meaning that your typewritten document will be accepted as evidence during a lawsuit or similar, while a PDF of same may not.
The slide pertaining to ASCII is here:
https://speakerdeck.com/alblue/a-brief-history-of-unicode?sl...
Oh wait, this article is 'man ascii' & 'man kermit'.
> Oct Dec Hex Char Oct Dec Hex Char
> ────────────────────────────────────────────────────────────────────────
> 000 0 00 NUL '\0' 100 64 40 @
> 001 1 01 SOH (start of heading) 101 65 41 A
> 002 2 02 STX (start of text) 102 66 42 B
> 003 3 03 ETX (end of text) 103 67 43 C
> 004 4 04 EOT (end of transmission) 104 68 44 D
> 005 5 05 ENQ (enquiry) 105 69 45 E
Holding Ctrl set bit 6 to '0', bit 7 to '1', and bit 8 to '0'. 'C' and 'c' differ by bit 6 only ('1' for 'c').From now on every time I Ctrl-d I want to think the voice of the Master Control Program.
It's been a while though, maybe he says both.
Kind of interesting how remnants of culture wars of 35+ years ago linger today, and from the perspective of 2017 how anybody could have gotten annoyed at people who eat egg-and-cheese pies.
Plus, people used them manually (control-S/control-Q) on systems to stop output scrolling by, and restart it when they've read what's on the screen, before built-in pagination filters (e.g., more(1) or less(1)) became common. (Specially back in the DECsystem-10/-20 days.)
It also doesn't help that most modern web servers also include logic to handle a single LF character to terminate lines anyway.
I have ported Linux to a custom ARM board many years ago. Started with a boot loader written in assembly and writing a single char into serial port for a debug console. It takes a single line of assembly or C. Infinitely easier than USB. From there on, I was able to unwind the whole system, develop USB drivers, TCP tunnels, etc.
I'd like to see how looks editing and running BCPL/C programs using ed on such a terminal.
My conclusion is that I could probably put up with using one of these, but that I'd feel quite cramped if it was my main workstation.
The mechanics of `ed' suddenly make a lot more sense now.
The dialing bit at the end was awesome. Reminds me of relay-based lift motor rooms! (There are videos of those on YouTube.)
Not quite true - early adopters like DEC kept using the 1963 version for a very long time, which prompted others to follow them. When the Smalltalk group decided to replace their own characters for ASCII in Smalltalk-80 to be compatible with the rest of the world, it was the 1963 version that they used.
Due to this, since I use the Celeste program in Squeak Smalltalk to read my email, I see a left arrow whenever someone wrote an underscore. The other difference is that I have an up arrow instead of ^. But it did adopt the vertical bar and tilde from 1967 ASCII, so it was a mix.
To the OP, I'm very curious why you're using Celeste. I take it you've been using it for years?
It is not something I would give to a "normal" person to use, but it is more than good enough to me. The main problem is that it does a reasonable job of showing incoming emails (though attachments show up as links at the end of the text) but when editing an email to send it shows the raw headers and MIME for any attachment.
What do you mean by "design Smalltalk computers"? That sounds really interesting. Do you mean you configure them to autoboot Linux into Squeak or similar? What are they used for? (...I'm guessing education-type environments...?)
I can understand what you mean. I tried to get it working, as I said, and while I got the main window open I had no idea how to configure it (and I must admit I don't have much incentive to.)
EDIT: Also, the most significant bit was used for 3/4 of the instruction space as a flag for a byte vs. word operation (or add vs subtract) so having it alone in its own octal digit (0/1) made perfect sense. For example 01ssdd was MOV (word) and 11ssdd was MOVB. 06ssdd was ADD and 16ssdd was SUB.
That's not quite right (the ACK part isn't right at all). See <https://en.wikipedia.org/wiki/Enquiry_character>.
Unix, on the other hand, was made similar to Multics which had the clever idea of on-the-fly replacing a line feed with whatever the printer required. So text files needed only the LF, and a CR was added automatically if the printer required it. This had its upsides and downsides, but the major upside was that you could print a text file on two different systems and reliably have it come out the same way!
This is one of the reasons that in e.g. C, the '\n' character is somewhat awkwardly defined as "a one-byte number that will move the cursor to the start of the next line", when on some operating systems (Windows) this will end up actually being the two-byte CR+LF sequence (well, when output is in text mode...). Even around the time of C's development this was already an issue and newline translation magic was already required, the ancestors of Linux and the ancestors of Windows just happened to put that magic in a different place.
Even if it isn't (and cannot) be literally true, I think Eliezer Yudkowski was on to something when he invented the "Merlin Interdict" in his fantasy work: Harry Potter and the Methods of Rationality.
Every highly systemized art form requires a living tradition to pass its most advanced achievements. Documentation, no matter how extensive, cannot convey all the subtleties, mostly because the experts are not 100% aware of all the little details that set them appart from the merely competent.
Some human endevours are best learned as a long process of deliverate practice under a wise mentor, who can help you steer the path and point your attention back to one apparently irrelevant detail or another. Self study is posible and effective, but it will only take you so far. And to rediscover some lost art from first principles requires a level of geniality on par with the first creators of the art in the first place.
By reading the article instead of lamenting lamentations ;)
Some of that stuff piques my interest, but then it can be difficult to get started understanding it relative to other stuff that's more widely covered on the web. I guess I am just a little pampered when it's comparatively much easier to learn about a web framework. But that's probably part of the old school hacker ethos as well.
These were used for serial data sources, not just network but punch cards or drums or magnetic tapes. GS, RS and US were intended for databases on serial data sources, "group" is a modern-day table. Lammert Bies has more: https://www.lammertbies.nl/comm/info/ascii-characters.html
You can repurpose the final two (record and unit) for CSV, but that's not their original role, and you'll have to make sure they're never opened via user-controlled anything as these control codes are non-printable.
I suppose the Excel or 1-2-3 team either didn't know about the separator characters, or they did but then-current text editors weren't useful enough to do anything with them (and user-editability was important). It's a real shame: with a more integrated system, they could have added functionality to the editors to understand ASCII record encoding, and CSV need never have existed.
Maybe a simple viewer would be nice.
An annotated history of some character codes or ASCII: American Standard Code for Information Infiltration
http://worldpowersystems.com/J/codes/
It describes how we got here in excruciating detail, starting with Morse. Military communication systems had great influence. Before the ASCII, as we know it today, there was an "ASCII-1963" that was a bit different.
The long and winding path: Morse Baudot Murray ITA2 FIELDATA ASCII-1963 ASCII-1967
Aren't SO and SI used by the ISO-2022-JP character encoding?
Of course Unix didn't interpret them; the terminal did.
The k,l,m,n,o,p,q,r,s,t,u,v,w keys then become useful for drawing lines
If you have a linux desktop, switch to an actual console (Ctl-Alt-F1), then:
echo -e \\x0E
And type a bit of lowercase characters between k and w. To get the terminal back to normal:
echo -e \\x0F
Probably not what you used to use, but "ctty com2" (and a couple of Returns directed at the VT100) will make it "go." I'm not 100 on what the keymapping is when you're in the setup screen (virtual Setup key at bottom-left).
The `]` character itself will not be printed, since `\b` will delete it from the visual line, and this effectively creates a side-channel for communicating "invisible" information within the regular character stream.
It's a pain to work with, though, since it makes things like `strlen` behave in very non-intuitive ways. Just imagine a string becoming longer when you delete the BS character. That's no fun.
For example, there are lots of old papers from Xerox PARC, Burroughs, AT&T and many others freely accessible on the Internet.
Some of them, we have to thank to the laborious work from people that bothered to digitize their original form, produced by plain typewriters.
Yet, I doubt many youngsters bother to read them.
Then again, I have read a few papers, it's just highly inconvenient. And in what kind of setting would you take the time to really read these anyway? Work setting? No, just get your work done. At home? I don't mind reading a bit during my own time, but I don't want to spend hours upon hours trying to understand something from a very different context. Academic? Yeah sure why not. But I'm not in academia!
For example https://www.youtube.com/watch?v=MikoF6KZjm0 is in a museum, as is https://www.youtube.com/watch?v=NEbMksxQAgs. You can go and check them out and poke them (to some extent).
While watching videos is really awesome, actually being able to watch these things in action provides a level of context that is impossible to convey digitally.
As for when to read them. I read them while commuting to clients, train and plane travels.
However, a small percentage of people will care, and it is that small group which will carry on this knowledge. Bemoaning that the "youngsters" don't bother is like complaining that most people don't bother learning Old English--there are some people who do learn it, presumably because they find it interesting, and those are the people that everyone else relies on for their occasional Old-English-reading needs.
What I tried to talk about was the other (if related) often repeated lament: that youngsters "forget" (in reality, never did know for starters) and reinvent stuff that was common knowledge/tech not long ago. I see this as something else than papers: I assume papers were more of a "bleeding edge" at the time they were written, and not all of them did spread to become well known (not to mention commonly understood) even at their time. For me (a tech mid-age?), the first realization of the trend came with some recent stories (a year or two ago?) about "a new app" for "woohoo, communication between mobile phones over audio!" I.e. modems, reinvented (not to mention the irony of phones now being more digital than analog, though still modem-ing over radio waves at lowest level... but I digress...).
That's the thing I believe old masters (but also commoners, like me! or you, whichever category you find yourself in) could try to popularize, so it doesn't get forgotten. And that's what I found brilliant in the original article. Re-pollination of "age"-old ideas, once common and obvious, now forgotten because out of use. In a lightweight, easily consumable and approachable tale.
That's the way of a wise old sage, telling young ones some lightweight and amusing stories by the campfire, but secretly hiding in them the good stuff he knew, or that he learnt the hard way in his own time. Engaging, amazing, and playing on curiosity, for the benefit of the new ones; not complaining and deriding. Playing with how to make the young ones care; tricking them into caring. That takes more effort, true; but then it may become a good proof whether one really cares about this stuff being remembered.
[edit] By the way, thanks for writing down your argument, especially because it made me flesh out some of my vague thoughts better, and explore them further even for my own understanding.
It's an invaluable source of insight into why various weird things in Windows are the way they are.
Deleted comment
I say "presumably" because cmd.exe is a fair bit of a rats' nest due to its organic development over the years. One of its biggest issues is that you must use cmd.exe in order to read/write stdio while a program is running; there's no API to do it any other way. (So you have to open a hidden window then use a bunch of fancy tricks to poll it every so often per second. Yup.)
However, this is being tracked and will (at last!) be fixed (apparently soonish): https://github.com/Microsoft/BashOnWindows/issues/111#issuec... (that's the relevant comment, the whole thread is interesting - there's actually a carefully-shot photo of the whiteboard all of this is on!)
(As for standard handles, you can redirect them without needing cmd, but there's some fiddly stuff involved. I don't remember the exact details, but it's something like: redirect parent process's handles to appropriate pipes, spawn the child, retract the original redirection, access pipes as appropriate. The API implies that you can just supply the pipes directly, but it doesn't work that way, because Windows is stupid. Which everybody knows. (But what not everybody knows: it's still not as stupid as POSIX.))
I finally get it: cmd.exe must be running for the *ConsoleOutput API calls to work.
I don't have the specifics for POSIX pipe redirection, but I do know that you use dup2() to redirect everything appropriately. Oh... as for actually launching a terminal, yeah, that's insane. Or rather it has an insane history.
I was shocked how literally the ASCII codes were followed by the PBX. It sent an "ENQ" before each command, we had to send an ACK back, and then it sent us STX/ETX-delimited records.
I'm 32 and working today, in 2017. I hope I make stuff that lasts this long.
The problem is, frameworks and libraries a) are both easy and fun to write, b) are really hard to comprehensively test with full architectural coverage, and c) suggest some form of standardized behavior. While (b) creates technical debt, the main issue is (c) vs (a): we're flooded with "do it this way!" from a thousand groups, even in situations where the developer(s) didn't really intend for that to be their predominant statement.
With all this noise and chaos, it almost feels awkward to stick to old, icky, widely-hated legacy standards in the face of all this innovation. Or at least that's what it's felt like to me. Objectively thinking about it and comparing everything, though, nothing's perfect, but what's been around for a while has the combined benefit of a) having a fairly widespread mindshare, and b) having known solution patterns for a wide range of issues.
I guess what I'm trying to say is that building software on top of tried-and-tested methodologies is likely to produce long-lived results (which is fairly logical).
I'm also reminded of http://qdb.us/53151 :D
- I say "or better than" because large corporations often have standardized internal Web style guidelines and rendering toolkits/templating engines, and making sweeping changes in those is harder than for websites that are little more than a landing page and some documentation - so the bigger an enterprise is, the likelier it is that its website might look mildly dated.
> 56 kilobits per second just before the technology was effectively wiped out by wide-area Internet around the end of the 1990s, which brought in speeds of a megabit per second and more (200 times faster)
That should read 20 times faster?
probably refers to the first widely installed (with coaxial cables) 10 MBit Ethernet networks...
For example we ship a certain device that can be configured fully both from CLI, from web GUI and from Windows application. Customers internally have teams that are 100% polarized - some teams use GUI stuff only and some teams use CLI only. Both refuse to switch on principle. ANd in our case I think in 5-10 years CLI will die. It is just a mess - you can do fast configuration with userfriendly extras in CLI but only on small scale. It just doesn't allow userfriendly editing of kilometer long configs, especially on 100s and 1000s of devices. So if we'll continue improve GUI configuration CLI will eventually die I think. Same could happen in general IT systems, provided there is better alternative (or forced alternative).
I do have to agree, GUIs can be easier to use if they're well-designed. It's also arguably easier to build GUIs than CLIs in some situations, particularly where you don't need something to be fully Turing-complete.