Simple Dynamic Strings library for C, compatible with null-terminated strings
github.com
github.com
1. Yep, more than "strings" SDS may be consider a library for dynamic buffers, especially from people coming from C++ or higher level languages. However I think that for C, it makes sense to provide a very low level thing like that.
2. In practice, if you see how SDS is used (extensively) inside Redis, it normally models things where you would do realloc magics, pointer math, and so forth: from client output buffers, to actual strings to accumulate error messages, to the Redis String object itself.
3. Probably an UTF-32 layer could be implemented on top of that, but I would think the two modules completely separated from the POV of the implementation. Like having an additional utf32.c file that provides additional interfaces on top of SDS.
4. The peculiar approach used by SDS header-before-pointer creates some problem with Valgrind and similar tools, however Valgrind will just report "possibly lost", that are messages mostly safe to ignore. However the advantage of using SDS strings directly with C function libraries is very handy.
SDS strings are not perfect and tend to be extremely optimized for Redis, if you look at the API, they allow to do things that are very low level, like pre-allocating the internal buffers to improve performances when you know you are going to read a big chunk of data from a socket or alike. However while imperfect and very tuned, these kind of libraries show how much you can easily improve C, with little work, and how many unsafe things in C are about lack of abstractions.
The value of SDS is "less is more" in the simple API they provide that pretends that SDSs are just plain-strings++. You can see this in a few features: plain C pointers interface, and the policy of always terminate the string. They need more love anyway, and to be less specialized for Redis, adding more useful APIs. Maybe at some point I'll find the time.
How do you handle this?
In a similar system I work on, we have a list of functions (like strlen, Strcat), where we carefully vet every occurrence in the code base. It's annoying but the only way to stop subtle bugs we've found.
C has always badly needed built-in strings (and arrays with size info, generally).
To save a byte, C designers committed The Most Expensive One-byte Mistake https://queue.acm.org/detail.cfm?id=2010365
Maybe not C + Knuth, but some way of portably advancing the language over time.
UTF-32 is, in all but extreme niche cases, a supremely wasteful and inefficient way to store and process strings, even in light of the cost of variable-length codepoints. The longest UTF-8 sequence is now four bytes anyway, so UTF-8 is never ever less compact than UTF-32. The only conceivable practical use of UTF-32 is a case where the upper 11 bits are used to tag characters with metadata, but in that case you'll have to mask those to display them anyway.
In practice, if your application must be aware of codepoints, it most likely should be aware of extended grapheme clusters as well, and those have an arbitrary number of codepoints in them.
UTF-16 is closer to being useful in some locales, but the marginal benefit (one byte less per BMP CJK codepoint) is, in my view, outweighed by the incompatibility with ASCII. Even a small number of ASCII strings in a given transmission will wipe out any advantage UTF-16 could have in CJK locales, and most transmissions these days have enough ASCII to do so. For example, today's Japanese Wikipedia main page is 102,620 bytes in UTF-8, and 182,480 bytes in UTF-16, despite most of the visible graphemes being CJK characters (which require three bytes in UTF-8). The Korean Wikipedia main page has larger text, and less of it, so the difference is even more stark (81,072 bytes in UTF-8, and 146,630 bytes in UTF-16). Even after aggressive gzip compression (with zopfli), text-heavy UTF-16 Japanese and Korean HTML documents are about 10% larger than their UTF-8 equivalent; and this is before you've transmitted any CSS or JavaScript.
When you're iterating over characters, with UTF8 branch prediction mispredicts all the time. Characters in UTF-8 have random length from 1 to 3 bytes. This might be good for bandwidth, but the tradeoff is slower processing, all modern CPUs have branch prediction hardware, and deep pipelines.
With UTF-16 CPUs predict these branches all the time, surrogate pairs are extremely rare, the rest of languages are 2 bytes/character.
While it's technically possible to optimize UTF-8 with clever programming like manual SIMD, that code gonna be much more complex than just a while or for loop, i.e. more expensive to develop and harder to support.
I think that's the reason why all platforms with rich GUI ecosystem (Windows, OSX, iOS, Android, JavaScript) use UTF-16 almost exclusively.
Considering this is C, there is no way to prevent this easily since you can't express a type that is one-way convertible (i.e. sds -> char* ok, char* -> sds not ok).
I do have to wonder though if avoiding a couple of .str's is really worth the risk. You definitely have to always keep in mind to never accidentally mix a char* with an sds, for the advantage of slightly cleaner looking code.
typedef struct sds_s {
char data[0];
} sds;
Which makes them essentially equivalent, but not from a type perspective. Then to handle conversions, you add explicit functions a la char* sds_cstr(sds *str);
Then you have full type checking to help you (and an additional advantage if your data format changes).Why not just use the member explicitly? This way, converting a STS string into a legacy C string could be as simple as writing “.cstr” (or if you just be as terse as possible, it could be defined as “.s” as another poster suggests).
In this case, compromise of increased code verbosity is extremely minor at worst (just a few characters). At best, this extra explicitness can actually be seen as a good thing for code readability (not to mention the huge benefits we’re discussing of type safety).
So the question then is: Why doesn’t SDS do this? The actual library uses a regular C typedef (which is unsafe for the reasons described above).
char* sds_cstr(sds *str){
return &str[0]; //you could probably just cast too
}
The other posters solution is actually not equivalent in this way (his sds is convertable to char* instead of sds*).Otherwise (if we want to retain the zero-cost guarantee of converting to a C string), directly accessing the member variable is: functionally equivalent, simpler, more concise, more readable, more explicit.
With that said, I agree that the additional verbosity is a significant drawback, and I can understand going with the slightly-less-safe option for that reason alone.
However: If we open the discussion up to C++, then this is pretty easy to solve without any real syntactic or performance compromise: we can make a class/struct which defines implicit conversions to C strings, and disables any implicit conversions from C strings.
typedef struct {
char *s;
} sds_t;
will produce an error if you try to pass a char* instead.
http://codepad.org/CmrskN9n(And as other commenters have noted, wrapping the char* in a struct addresses the immediate footgun in question.)
legacy_fn(my_text.s);
Versus the current: legacy_fn(my_text);
I don’t think saving two characters per legacy function call is even remotely worth the loss of static type safety (which risks serious memory corruption and/or security holes, which are entirely preventable at compile-time in this way).In fact, I even find the explicitness more pleasantly and clearly readable: I like being able to know at-a-glance when types are changing, especially in a language as unsafe as C.
Lastly, if you can tolerate just using some of C++‘s features, you can define a no-compromise solution: A type (still represented by a single pointer under-the-hood) that will implicitly convert (with zero runtime cost) into a C string but not vice versa.
typedef struct sds {
char *s;
} sds;
Now you get type safety: you can't accidentally pass a plain 'ole C string to an sds function. But you do need to access the 's' field to get the C string for passing to non-sds functions.https://sourceforge.net/p/joe-editor/mercurial/ci/default/tr...
I also have NULL terminated arrays of strings, like arguments lists:
https://sourceforge.net/p/joe-editor/mercurial/ci/default/tr...
It's set up so that you could make a dynamic arrays of any types by copying the header and source files, but changing a few constants and providing comparison and duplication functions for the elements involved. It's like manual template instantiation.
Of course this is prone to memory leaks, same as sds. But there is branch here:
https://sourceforge.net/p/joe-editor/mercurial/ci/coroutine/...
In this version, all strings are allocated on an obstack. Space for temporary strings is automatically reclaimed when you return to the top level.
You can mark a string as permanent, then it will not be reclaimed, and instead has to be explicitly freed. So you can still have memory leaks but less likely, and also you could have accidental automatic freeing, but a lot of explicit frees are eliminated.
C++ strings are better... except that I have my own library for them also because I hate not being able to return NULL to indicate a failure. This is called the semi-predicate capability of C strings, and is easy to have in C++ with a different library.
When I have a major refactoring project, I typically will try to compile it with C++ first to help catch all the weird wobbly bits like this.
[EDIT: I now see in the comments other people with a similar opinion, so maybe not as unpopular as I would've thought!]
Is this a complete drop-in replacement?? If so it could occasionally make my life much easier!
[0]: https://www.fefe.de/libowfat/ [1]: https://www.fefe.de/gatling/
aStr("string"); then aMap aFold aConcat, etc, but aAppend(&array,'\0') needs to be done manually if passing to standard str functions. This is so long* or whatever arrays can have the same interface
Often the former will need just a single shared allocation (per thread). That works if you're only ever constructing a single string at a time.
Often the latter won't be dealing with memory management at all. You will often represent it by a plain char pointer and optionally a length field. Management-wise, at most you will need to construct it initially by making an immutable string from the mutable one, and a call to a release function when the string is no longer needed.
The fact that the string isn't modified in the immutable phase makes it trivial to pass it around as an ordinary char pointer (plus optional length), which is as convenient as it gets.
In rare cases you might want to make a hashmap for string interning. I've used interning in a past compiler project, where it was used to collapse identifier strings.
That will probably cover more than 95% of all your string needs. And it's very easy to be vastly more efficient than any from-the-shelf GC'ed string type with this. All you need to invest is to write a little near-boilerplatey code, but it's far from "hard".
I'm not convinced of the advantage in returning a char* to the user.
I see how it's convenient to be able to use all the built-in and other functions that accept a char* as an argument, like printf() or your favorite logging library.
BUT, you can only use the ones that treat the string as read-only. (You can't use strncat(), etc.) And you have no protection against messing this up.
Seems like a better trade-off would be to have a user-visible type that isn't char, then a user-visible function that converts that to char when you need it.
So instead of this:
/* sds is just a typedef to char* */
sds mystring = sdsnew("Hello World!");
printf("%s\n", mystring);
sdsfree(mystring);
You'd have a function like this: /* get read-only C-style string from an sds */
const char *sdsC(const sds s);
And code like this: /* sds is its own distinct type, not another name for char* */
sds mystring = sdsnew("Hello World!");
printf("%s\n", sdsC(mystring));
sdsfree(mystring);
Yes, it's more keystrokes, but surely the safety is worth it considering it is only needed when bridging a compatibility gap. (Also, possibly the sds functions could be a tiny bit more efficient if they aren't always doing conversions on their arguments.)(I do like the idea of putting header and characters into one struct. That's probably good for efficiency compared to a struct that points to a buffer and gives the system a layer of pointer indirection to go through.)
> Normally dynamic string libraries for C are implemented using a structure that defines the string. The structure has a pointer field that is managed by the string function, so it looks like this:
struct yourAverageStringLibrary {
char *buf;
size_t len;
... possibly more fields here ...
};
> SDS strings as already mentioned don't follow this schema, and are instead a single allocation with a prefix that lives before the address actually returned for the string.Joking aside, I assume this is char set agnostic?
(Note: This has not been touched since 2000, use [2] instead)
[1] https://www.fefe.de/libowfat/Should have been done 30 years ago.
I'm also missing stack allocation support, needed for fast short strings. It should be even included in sdsnew, for len < 128.
People seem to think that 1 codepoint = 1 integer frees you from thinking Unicode is hard. But 1 codepoint is not 1 glyph. You have combining characters, zero-width joiners [used also in emojis], RTL markers, Han unification, probably more. So you can't really think of a Unicode string as a random-access, one-glyph-per-unit type of thing in any encoding.
So I hope when people say "UTF-32 support" they mean "decode UTF-8".
The stuff you mention does make Unicode very messy though.
The alloca() function allocates size bytes of space in the stack frame of the caller. This temporary space is automatically freed when the function that called alloca() returns to its caller.
That would not work, since the code calling alloca() is the library, and you want the to allocate on the stack frame of the function calling the library. I don't think that is possible, unless you make the library function into a macro (perhaps using ?: on the allocation size).