C++ String Conversion: Exploring std:from_chars in C++17 to C++26
cppstories.com
cppstories.com
The API looks like it's following best-practice, with a specific result type that also contains specific error information and yet, that's not enough and you still end up with edge-cases where things look like they're fine when they aren't.
from_chars is supposed to be a low lever routine, it must be possible to use it correctly, but it is not necessarily the most ergonomic.
The whole reason for returning a result enum would be so I don't forget to also check somewhere else whether something could possibly have gone wrong.
Especially when there is a specific result type.
What’s coming out is a std::from_chars_result: a status code plus an indicator how much of the data was consumed.
What to name this depends on how you see this function. As part of a function on number types, from_chars is a good names. As part of a function on strings, to_int/to_long/etc are good names. As freestanding functions, chars_to_int (ugly, IMO), parse_int (with parse sort-of implying taking a string) are options.
I can see why they went for from_chars. Implementations will be more dependent on the output type than on character pointers, it’s more likely there will be more integral types in the future than that there will be a different way to specify character sequences, and it means adding a single function name.
const std::string str { "12345678901234" };
int value = 0;
std::from_chars(str.data(),str.data() + str.size(), value);
On the third line: Why can I just pass in `value` like this? Shouldn't I use `&value` to pass in the output variable as a reference?A non-const reference is just as clear a signal that the parameter may be modified as a non-const pointer. If there's no modification const ref should be used.
You could choose to textually "tag" passing by mutable ref by passing `&foo` but this can rub people the wrong way, just like chaining pointer outvars with `&out_foo`.
So you can dip value into the function call and the function can assign to it, as it could to data pointed to by a pointer.
int x = 0;
int &r = x;
r += 1;
assert(x == 1);
int y;
r = y; // won't compile
void inc(int &r) { r += 1; }
int x = 0;
inc(x);
assert(x == 1);
The equivalent using pointers, like in C: int x = 0;
int *p = &x;
*p += 1;
assert(x == 1);
int y;
p = &y;
void inc(int *p) { *p += 1; }
int x = 0;
inc(&x);
assert(x == 1); r = y; // won't compile
This will compile. It will be effectively the same as "x = y". The pointer equivalent is *p = y".My apologies, as it's been a while since I've used C++.
You should think of `from_chars` as a function that accepts the outputs of `to_chars`, not as a general text understander.
std::u8string_view chars{ u8"۱۲۳٤" };
int value;
enum { digit_base = 10 };
auto [ptr, ec] = std::from_chars(
chars.data(), chars.data() + chars.size(), value, digit_base);
return (ec == std::errc{}) ? value : -1;
will fail to compile due to pointer incompatibility.E.g. "۱۲۳٤" isn't as far as I know a valid number to put in a json file, http content-length or CSS property. Maybe it'd be ok in a CSV but realistically have you ever seen a big data csv encoded in non-ascii numbers?
Of course nobody would design an application like that.
Let me ask you this: How much data have you processed which comes from human user input in Arabic-speaking countries?
For example: With Python Python breaking changes are more common, and everyone complains about how much they have to go and fix every time something changes.
Damned if you do, damned if you don't. Either have good backwards compatibility but sloppy older parts of the language - or have a 'cleaner' overall language at the cost of crushing the community every time you want to change anything.
"The Design and Evolution of C++" gives the impression that even back in the 80s, major concessions were being made in the name of compatibility. At that time it was with C; now it's with previous versions of C++.
Overloads are already confusing, if they can't be used generically there is really no point in reusing the name.
Maybe it was for ::reasons::.
You can even parse a whole line of a csv file with multiple numbers in one call "sscanf(str, "%d;%d;%d\n, &d1, &d2, &d3)"
Another person already correctly pointed out that the author was listing a bunch of functions for number and string conversions in general.
The article doesn't change that. It also lists scanf...
That sounds like marketing BS, especially when most likely these functions just call into or are implemented nearly identically to the old C functions which are already going to "offers the best possible performance".
I did some benchmarks, and the new routines are blazing fast![...]around 4.5x faster than stoi, 2.2x faster than atoi and almost 50x faster than istringstream
Are you sure that wasn't because the compiler decided to optimise away the function directly? I can believe it being faster than istringstream, since that has a ton of additional overhead.
After all, the source is here if you want to look into the horse's mouth:
https://raw.githubusercontent.com/gcc-mirror/gcc/master/libs...
Not surprisingly, under all those layers of abstraction-hell, there's just a regular accumulation loop.
MSFT's one is totally standards compliant and it is a very different beast: https://github.com/microsoft/STL/blob/main/stl/inc/charconv
Apart from various nuts and bolts optimizations (eg not using locales, better cache friendless, etc...) it also uses a novel algorithm which is an order of magnitude quicker for many floating points tasks (https://github.com/ulfjack/ryu).
If you actually want to learn about this, then watch the video I linked earlier.
https://github.com/fastfloat/fast_float
For more historical context:
https://lemire.me/blog/2020/03/10/fast-float-parsing-in-prac...
- Afaik stoi and friends depend on the locale, so it's not hard to believe this introduced additional overhead. The implicit locale dependency is also often very surprising.
- std::stoi only accepts std::string as input, so you're forced to allocate a string to use it. std::from_chars does not.
- from/to_chars don't throw. As far as I know this won't affect performance if it doesn't happen, it does mean you can use these functions in environments where exceptions are disabled.
Lastly, having a simple and narrowly specified conversion routines allows one to create a small sub-set of C++ standard library fit for constrained environments like embedded systems.
"Extraordinary claims require extraordinary evidence."
The standards committee's purpose is to justify their own existence by coming up with new stuff all the time. Of course they're going to try to spin it as better in some way.
It compiles from sources, can be better in-lined, benefits from dead code elimination when you don't use unusual radix. It also don't do locale based things.
Removed `std::atoi` from the benchmarks since it was performing so poorly; not a contender. Should be easy to verify.
Rough results (last column is #iterations):
BM_fast_int<std::int64_t>/10 1961 ns 1958 ns 355081
BM_fast_int<std::int64_t>/100 2973 ns 2969 ns 233953
BM_fast_int<std::int64_t>/1000 3636 ns 3631 ns 186585
BM_fast_int<std::int64_t>/10000 4314 ns 4309 ns 161831
BM_fast_int<std::int64_t>/100000 5184 ns 5179 ns 136308
BM_fast_int<std::int64_t>/1000000 5867 ns 5859 ns 119398
BM_fast_int_swar<std::int64_t>/10 2235 ns 2232 ns 316949
BM_fast_int_swar<std::int64_t>/100 3446 ns 3441 ns 206437
BM_fast_int_swar<std::int64_t>/1000 3561 ns 3556 ns 197795
BM_fast_int_swar<std::int64_t>/10000 3650 ns 3646 ns 188613
BM_fast_int_swar<std::int64_t>/100000 4248 ns 4243 ns 165313
BM_fast_int_swar<std::int64_t>/1000000 4979 ns 4973 ns 140722
BM_atoi<std::int64_t>/10 10248 ns 10234 ns 69021
BM_atoi<std::int64_t>/100 10996 ns 10985 ns 63810
BM_atoi<std::int64_t>/1000 12238 ns 12225 ns 56556
BM_atoi<std::int64_t>/10000 13606 ns 13589 ns 51645
BM_atoi<std::int64_t>/100000 14984 ns 14964 ns 47046
BM_atoi<std::int64_t>/1000000 16226 ns 16206 ns 43279
BM_from_chars<std::int64_t>/10 2162 ns 2160 ns 302880
BM_from_chars<std::int64_t>/100 2410 ns 2407 ns 282778
BM_from_chars<std::int64_t>/1000 3309 ns 3306 ns 208070
BM_from_chars<std::int64_t>/10000 5034 ns 5028 ns 100000
BM_from_chars<std::int64_t>/100000 6282 ns 6275 ns 107023
BM_from_chars<std::int64_t>/1000000 7267 ns 7259 ns 96114
BM_fast_float<std::int64_t>/10 2670 ns 2666 ns 262721
BM_fast_float<std::int64_t>/100 3547 ns 3542 ns 196704
BM_fast_float<std::int64_t>/1000 4643 ns 4638 ns 154391
BM_fast_float<std::int64_t>/10000 5056 ns 5050 ns 132722
BM_fast_float<std::int64_t>/100000 6207 ns 6200 ns 111565
BM_fast_float<std::int64_t>/1000000 7113 ns 7105 ns 98847Your dismissive answer sounds so much like the one of a typical old-C style programmer that underestimate by 2 order of magnitude what compiler inlining can do.
Abstraction, genericity and inlining on a function like from_chars is currently exactly what you want.
For small size functions inlining is almost always preferable because (1) the prefetcher actually love that and (2) a cache-miss due to a mis predicted jump is way more costly than anything a bit of bloat will ever cost you.