The format of strings in early (pre-C) Unix
utcc.utoronto.ca
utcc.utoronto.ca
bec 1f / branch if no error
jsr r5,error / error in file name
<Input not found\n\0>; .even
sys exit
jsr calls a subroutine passing the return address in register 5. The routine error interprets the return address as a pointer to the string.r5 is incremented in a loop, outputing one character at a time. When the null is found, it's time to return.
The instructions used to return from "error:" aren't shown but there is a subtlety here, I think.
".even" after the string constant assures that the next instruction, "sys exit", to which "error:" is supposed to return, is aligned on an even address.
By implication, the return sequence in "error:" just be sure to increment r5, if r5 is odd. I am guessing something like the pseudo-code:
inc r5
and r5, fffe
ret r5
http://minnie.tuhs.org/cgi-bin/utree.pl?file=V1/sh.s
error:
...
inc r5 / inc r5 to point to return
bic $1,r5 / make it even"How I do per-address blocklists with Exim"
https://utcc.utoronto.ca/~cks/space/blog/sysadmin/EximPerUse...
I run Exim, and I'm also a huge believer in blocking spam at the SMTP level, and also do some things globally that should perhaps be per-user. I'm eagerly interested in everything this fellow has to say.
Maybe I should just make HTML 5 games instead...
"In BCPL, the first packed byte contains the number of characters in the string; in B, there is no count and strings are terminated by a special character, which B spelled ‘*e’. This change was made partially to avoid the limitation on the length of a string caused by holding the count in an 8- or 9-bit slot, and partly because maintaining the count seemed, in our experience, less convenient than using a terminator.
Pascal strings were significantly size-limited due to memory constraints (an implementation detail), so there was a reason to prefer null termination, but in essence we're still slaves to an obsolete instruction set.
All Pascal dialects like Apple Pascal and Turbo Pascal had pointer arithmetic.
Should I list all the compiler specific features that everyday C programmer uses but aren't defined in ANSI C, e.g. inline assembly?
No professional Pascal compiler was a pure ISO Pascal, that error was left for schools teaching bad Pascal.
Professional Pascal compilers were way more powerful than C, safer and in the days of K&R C chaos, just as portable.
The PDP-11 features that make this possible are (1) post increment addressing and (2) MOVE instructions set the condition codes.
As long as you are handling good data, it's clearly more efficient.
not in x86
People who go ape-shit over 0-terminated strings would probably be even more upset if we still did that.
For example, a string copy 'instruction' that copies a 0 terminated string from one location to another looks like this:
lea esi, source_string
lea edi, dest_string
repnz movsb ; copies esi -> edi until *esi = 0repnz movsb is undocumented and undefined, MOVS doesnt modify flags, not sure how would that work
While the rep/repe/repnz can be used in conjunction with movs(b/w) it will not terminate on the z flag in the eflags register (so it won't stop if *esi == 0) but rather when ecx = 0.
So ecx needs to contain the size of the string you want to move, and rep movsb does not use the zero flag.
Taking all three together, you end up with the conception of a "string" as a serialized bitstring encoding a sequence of four lexical types: a NUL type (like the EOF "character" in a STL stream), an ASCII control-code type (or a set of individual control codes as types, if you like), a set of UTF-8 "beginning of rune" types for each possible rune length, a "byte continuing rune" type, and an ASCII-printable type. (You then feed this stream into another lexer to put the rune-segment-tokens together into rune-tokens.)
In the end, it's not a surprise that all of these components were effectively from a single coherent design, thought up by Ken Thompson. It's a bit annoying that each part ended up introduced as part of a separate project, though: NULs with Unix, gets() with C, and runes with Plan 9.
One of the pleasant things about Go's string support, I think, it that was an opportunity for Ken to express the entirety of his string semantics as a single ADT type. That part of the compiler is quite lovely.
You have two choices, counted or terminated.
Counted places a complexity burden at the lowest level of coding.
With terminated you still have the option of implementing strings with structs or arrays with counts or anything.
And people did of course. Many many different implementations of safe strings exist in C; the fact that none have won out vindicates the decision to use sentinel termination.
If only Dennis had had the foresight to nip that one in the bud...
This would naturally want malloc to know the type and count, eg: char[] x = malloc(char[], 100). That means no opportunity to screw it up (let the compiler turn that into the sizeof math to pass to the actual allocator).
If bounds checking is a performance bottleneck you could turn it off with compiler flags; that's not a valid argument against it.
But hey... all the various buffer overflows, RCEs, and various exploits are totally worth the minor performance gains /sarcasm.
The article says it doesn't know where ZTSs came into Unix, but I thought I read somewhere, many moons ago, that it was a pretty simple choice: prefixed strings used one more byte of memory than a prefixed string for a string longer than 255 characters, and on the systems at the time, that mattered a lot.
Just think about how something like strsep (strtok to you older folks) would work in a system like this.
And even then — one byte of memory could be nothing compared to CPU overhead. Or maybe not — RAM was insanely expensive these days.
So in other words, I don't really understand why you are arguing for replacing a standard - one that works well for its purposes, mind you - with another when this has in fact already happened. And even less I understand why you are trying to frame a good and sound engineering decision as somehow a mistake?
The String object was a character array (presumably terminated by a NULL) which was prefixed by a size_t (it may have been assumed that String inherited another precursor class).
Still, I do like the idea of working within a bounded set. This sounds much more like a language feature where you could work within an array of known size. Though most of the time when I care about easy String handling I'm better off writing it in a higher level language.
You can encode the byte length of the string minus the rounded up size (known by malloc) using the same trick that OCaml uses: https://rwmj.wordpress.com/2009/08/05/ocaml-internals-part-2... Using this trick you can get the byte length of the string from the base pointer in O(1). Such strings can also contain \0 characters.
Thus you don't need to store the string length explicitly as a separate field, and you can even make string pointers within the string work, and it's backwards compatible with legacy functions that expect a null-terminated string (albeit it won't work if you pass a string containing \0 to one of those legacy functions, but that is to be expected).
The prevailing view of how to do things was based on tape (magnetic or paper). In that context, it's really inconvenient to know how long your tape is before you write it and really convenient to just keep reading it until you hit a marker or the hardware says "sorry, end of tape". See also the "end of file" character, ctrl-D.
Judging history in hindsight is uncharitable.
Burroughs B5000 in 1961, with memory safe systems programming in an Algol dialect.
https://en.wikipedia.org/wiki/Burroughs_large_systems
Similar examples can be provided for systems older or of similar age to PDP-11.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.162...
Security was hard then and continues to be hard now. It would also be nice to have a cost comparison between the various alternatives.
1. If string is 0-15 chars long, str_t is a small string, aka char[16]. First byte is length, payload starts in 2nd byte, tail is always zeroed. Empty string is all zeros always.
2. If string is 16-248, str_t is a half string. The whole payload exist in a single contiguous buffer. First bytes is length. A pointer to the rest of the payload exist within the remaining 15 bytes, starting at the first alligned address. This way you can cast half string as small string when passing as value without knowing the size at compile time (and stil recover the rest of the payload). Gap between the length and the pointer is reserved and always zeroed.
3. For lengths 249 and higher. Each value 249-255 is not the lenght, but an encoded way to represent one or more styles of bigger strings. Not need to hold the payload in a single contiguous buffer. Remaining 15 bytes will hold a pointer to the data structure in the same fashion as half string. Reserved bytes are used to represent length if possible. If not, length is required to exist at a fixed position in the first buffer accesible when dereferencing the pointer.
If you get it right and hide all of the complexity from the user then it's not quite as bad, although undoubtedly confusing the first few times someone inspects it in their debugger.
In Visual Studio debugger you can write scripts to pretty-print string variables in the "watch" window, I would guess other good debuggers can do this too.
The goodness of this approach is that if you just need to pass around the string, a shallow copy of the first 16 byte block is enough. It might or might not contain the whole of the array, but if you need to know, it means you need to go through all the corner cases, no strcat(small_buffer, unsanitized_user_input), thank you very much!
If you think that it is a bad idea to do:
str_t* string = (str_t*)malloc(string_length);
then half of the goal is been achieved right there.The other half is to have this nailed down in the language definition in sufficient detail so that implementations do not differ widely and you can poke inside it with assembly when you have to.