Str: Yet another string library for C language
github.com
github.com
Every library like this is incompatible and most make slightly odd choices like in this library ownership of the string is denoted by a bit in the info/size field. Not that that is a bad choice or anything, but it is one reason someone might decline to use it and decide to write their own.
The lack of namespacing in C doesn't help, this library chooses str_ as its prefix, which is a bit likely to collide with other libraries. It also makes it harder to try to write libraries that allow for the string library to be switched out.
It looks like it'll be getting strdup and strndup.
I bet the security industry agrees with my definition.
It builds character and you learn something from it.
It is a rite of passage.
However, don't add to the pile of dependency hell that is already plaguing many open source projects. If you feel uneasy with how C strings work, consider switching programming languages instead. You will probably have an easier time and there will be less unmaintained incompatible string libraries rotting around on github.
If you create small libraries that don't produce shared objects intended to stand on their own with a stable API/ABI, but are simply headers or at most produce a .a from source fully intended to become vendored in-tree, you're not contributing to "dependency hell".
When we refer to "dependency hell", AIUI, it's in reference to unresolvable runtime dependencies creating hell for end-users.
...until you someone exploits the bugs in it.
Everyone who did the exercises in K&R should be able to write their own string library, probably with less bugs than the standard one. However, I really feel it's much better for everyone to use proven code like bstring.
Fewer bugs than what “standard one”?
That's not how most people use the word "bug".
BUGS: Never use gets().
That's awful and I love it.
0-terminated strings not only have proven to be a rich source of bugs, they're remarkably inefficient as well [1]. Doing better was a major focus of the initial design of D.
[1] This is because of constantly scanning to get the length (which also necessitates reloading the string contents into the memory cache), and having to make copies of strings instead of just slicing them.
C programmers are expected to make the best choice based on the situation. The various choices trade off memory usage, CPU usage, source code readability, and program correctness.
Which is problematic for thread safety and depending on the source of the string (constant) may not be possible.
If available, strdupa() would be a fine way to get a suitable local copy of the string. Commonly though, the programmer knows that there will not be threads and can make the string non-constant.
I wonder how much time was wasted in early computing (maybe not wasted really) because of the fear of incompatibility that is getting smaller and smaller as computing platforms coalesce into standardized-ish things.
Watching the M1 roll out and how it doesn't seem to care much that x86 is a thing and gets along with its life has been fascinating.
It doesn't have the same breadth of features as, say, Python's string class, but it's ok.
String literals also have an extra 0 appended, making it transparently easy to still pass strings to C functions like printf.
I don't know if that is a typo given you normally call them "fat pointers", but they are "pretty hot and tempting".
1: https://www.boost.org/doc/libs/develop/libs/nowide/doc/html/...
- I'm pretty sure public symbols starting with "str" are reserved by the standard.
- Declaring function arguments as const is pretty silly for value types, imo.
It is also very well documented. And all you need to embed it in your project is : sds.c sds.h sdsalloc.h
The source code is small and every C99 compiler should deal with it without issues.
sds a = sdsnew("hell");
sds b = a;
a = sdscat(a, "o"); // this invalidates b
Masqueraded pointers are inherently linear (or affine if you are pedantic). Any length-changing updates to such pointers can potentially reallocate them, so any value can't be "updated" more than once; values should be consumed and returned by many operations. No typical C types behave like this: primitive values or structs can be updated by assignments and pointers can be updated by dereference. C doesn't support linear types and, while normal pointers do need care, masqueraded pointers need much more care to use correctly. Yes, you can replicate the same bug with normal pointers by replacing the third like to `free(a);`, but you wouldn't expect a bug for non-destructive operations. (Put in the other way, masqueraded pointers make many otherwise non-destructive operations destructive.)While technically not a string library, this and the strict-aliasing issue for type-generic routines prompted me to write my own small extensible vector library [1] years ago.
https://github.com/maxim2266/str/blob/f4e84657b23977ab3c5cd7...
seems unlikely to matter if you have a bunch of strings flying around...
two features I'd love to see implemented:
- wrapping thread safe tokenization using strtok_r so it's pleasant to tokenize a string
- sprintf-like formatting
anything that improves string handling in C is doing God's work
We had "int errno" as global state. We fixed it, in a compatible way, to be thread-safe. Platforms without threads can still implement it the old way if desired.
The same kind of compatible fix could have been done with the strtok() function. There was no need to introduce another function.
Simply: the internal state of strtok() shall be distinct for each thread. (which is trivial if the platform only supports a single thread)
All the UTF-8 codepoints can be held inside an 8bit char, which is what this library seems to use under the covers.
You might need to add a couple UTF-specific methods if you want number of graphemes rather than number of bytes, but there's nothing to stop you placing UTF8 data inside a char buffer.
I mentioned UTF-8 specifically, because the UTF-8 encoding actually does specify this particular feature:
> UTF-8 is capable of encoding all 1,112,064 valid character code points in Unicode using one to four one-byte (8-bit) code units. Code points with lower numerical values, which tend to occur more frequently, are encoded using fewer bytes. [0]
UTF-8 is a variable length encoding, some characters need 8 bits, other characters need 32 bits per your own quote.
And it is designed so that it fits in one-byte units. That was a central goal of the encoding.
So... If you have a char buffer, like we have been talking about, you can toss any valid UTF-8 sequence inside it.
Up to you to make sense of the data in singles, pairs, fours, and before that to declare the buffer as the appropriate multiplier plus 1.
If you naively put the bytes of UTF-16 or UTF-32 encodings into a buffer, they might contain NUL (zero) byte values. Which, for C strings, means "end of string". UTF-8 makes sure this doesn't happen, which makes it compatible with existing C string functions.
I loved this. “This ain’t your fancy schmancy Tesla, it’s granddaddy’s old Ford pickup”.
> Disclaimer: This is the good old C language, not C++ or Rust, so nothing can be enforced on the language level, and certain discipline is required to make sure there is no corrupt or leaked memory resulting from using this library.
Avoiding the _Generic keyword is difficult. One might try using the sizeof operator.
This impacts sort and comparisons mostly. But without cmp you cannot search in strings.
No, they can't? These two UTF-8 byte sequences in a C char pointer,
c3 a9 00
65 cc 81 00
Represent the same string, but do not compare equal with strcmp.And it's not just that; you've noted how strtok will break down. strchr() can't be used w/ a non-ASCII needle, there is no support for code units, etc.
Lexicographical sorting over UTF-8 strings is actually the same as lexicographical sorting over the corresponding Unicode code point sequence.
Indeed.