Curious lack of sprintf scaling
aras-p.info
aras-p.info
I'm probably dumb but one of things that bugs me with various libraries is when someone has made the decision to do something high level at a low level. For example, localizing inside sprintf,
An another example might be an unzip library (read a zip file). Ideally the library should be small IMO. The simplest might you you pass it a bucket of bytes. If you want to make it flexible then you pass it some abstract interface (or 1-2 functions + void* userdata) so you can supply a "read(byteOffset, length)". You can then provide, outside of the library, streamed files, stream networking, etc...
But, bad libraries (bad IMO) will instead provide like 12 overrides "unzip(void* bytes), unzip(const char* filename), unzip(socket), unzip(url)" and end up including the world in their library. This kind of "try to do everything" is extremely common in npm libraries :( I don't need your library to include command line parsing! If you want to make a tool, make a library, then make a separate tool that uses that library. Keep the 2 separated so users of the library don't need dependencies that only the tool needs. (probably the most common npm example but there are lots of others)
Really surprised something as low-level as sprintf needs locale. Even streams I'd expect maybe a Date object would but not the stream itself.
https://www.npmjs.com/package/is-thirteen https://www.npmjs.com/package/is-not-thirteen
Decimal places in floating point numbers is the primary standard thing (some locales use `.` as a decimal seperator and `,` as a thousands seperator... some locales swap interpretations!) Nonstandard extensions may also do things like accept unicode strings, wide unicode strings, etc. which may need to be re-encoded to whatever the ambient locale-specified narrow encoding is (UTF16 => windows-1251?).
In the modern era of sprintf being exiled to some low level internal thing, I agree it makes little sense - causing more bugs than it fixes - but in the days of using it for a lot of heavy lifting of user-facing data, it was an understandable target for this kind of treatment.
Standards committees should not carelessly sit down and devise things that destroy performance by a factor of ten or more, especially when it is entirely unnecessary and where they do not keep or provide simple, fast alternatives for the most common cases where the format is known and has nothing to do with a preferred locale. What they did instead was a massive energy and time wasting imposition of unhelpful semantics on functions that did not have or need any such thing instead of a tailored let's only do this in special cases with special functions approach.
inline string
to_string(float val)
{
const int n = gnu_cxx::numeric_traits<float>::max_exponent10 + 20;
return gnu_cxx::to_xstring<string>(&std::vsnprintf, n, "%f", val);
}
it will be worse in pretty much every metric than fmt::format("{}", val);The C89 standards committee wasn't "careless" here - it was a product of the times. Of course, when modern APIs and standards committees duplicate sprintf's issues, it's far more fustrating - even when justified in the light of "backwards compatability".
Unfortunately, it takes experience for people to realize that they're not such a good idea. A better design is the component approach. As you suggest, a zip library should not be doing file I/O.
I've seen audio libraries where you can only provide a file path for playback. Guess you're screwed if you wanted to get a Vorbis file out of a tarball, decode it silently faster than real-time, and send it over the network...
In case this comes off as sarcastic, it is not intended so.
Europe got decimal digits with their own characters via the arabs, true, but we use shapes derived from the original source, the Devanagari digits, not the shapes the Arabs developed due to their writing system/technology. South Asian languages other than Urdu are LTR.
You can see the parallels here (clipped from Wikipedia):
० 0 १ 1 २ 2 ३ 3 ४ 4 ५ 5 ६ 6 ७ 7 ८ 8 ९ 9
Luckily none of this stuff cares about locales.
This is bad for not only performance, but correctness. If you are writing a JSON serializer and you use sprintf() to format float/double values, someone can make you produce incorrect output by setting the locale to something that uses "," as a decimal separator.
And there is no way to turn this off!
This is one moment where the difference between an application language and a system language seems clear. An application language could maybe assume that string formatting is for showing to a user. A systems language should assume that you are implementing protocols, and that mere user preference should not change your output.
Install linux on your laptop, seeing the source code is worth it.
That turns it off pretty well. Even if other libs call setlocale (looking at you, gtk), they'll get your version.
like so: https://github.com/smcameron/space-nerds-in-space/blob/maste...
It took surprisingly long to track down, partially due to this all being dependent on various ambient Information (env vars, browser config, arguments passed, what exact function is called, etc).
I'm more surprised that sprintf is seen as low level. It formats and outputs strings, just by that token the problem space is enormous. Localization is just a small part.
People are fast to forget the good old days where sending specific user strings to people would crash their browser or os. "Plain" text is not simple.
Why should the country I'm in and the language that I'm currently speaking determine if I want periods or commas as thousands separators? Or if I want dates to be presented in one of dozens of insane formats instead of RFC 3339/ISO 8061?
What are we doing right now to each other besides banging out some symbols that have a shared meaning?
1. The fact that I grew up with those formats holds little weight for me. At this point, most of the literature I read and content I consume doesn't come from my own country. And why should it? when I have access to books, movies, websites from all over the world, and my country makes up for a tiny fraction of that. With the massification of international remote work this tendency will only increase.
2. Even when working with people from my own country, I can agree with them to use formats different from the ones we grew up with. In fact, everytime we type a floating point literal in any programming language, we do it without following the rules we grew up with.
3. In the 21st century, in the context of globalization, massification of the Internet and widespread access to computing, to keep doing these things differently by country/language makes no sense, specially when there are already good international standards that we can follow. It's not that hard to learn, either. We even accept English as the de facto language of programming and software development, and that's a full natural language that we have to spend years learning.
4. In my opinion, yyyy-mm-dd and dot as decimal separator are simply better for practical reasons. However, if we collectively decide that other formats are the "international standard" then I'd follow whatever rule we decide, as long as it's not too bad.
Can you imagine if instead of using SI units, each country used their own special units of measurement? We are definitely better off with SI, standarization makes everything so much easier. I can talk to someone from Japan about meters and kilograms and they will understand without any ambiguity.
I could resign myself to Americans* being allowed to pick the global date format to save a few cycles on an operation, but frankly I don't want to. It's not that important to me, and not having aspects, albeit minor aspects, of my culture steamrolled in the name of pointless efficiently is at least a little bit important to me.
Differences create friction yes, but I'm alright with that. I would prefer a little friction to grey uniformity. You're free to think differently, but you're not the speaker for everybody else.
* I should clarify that most Americans don't seem to want this either, this isn't a jab, it's just that if we did standardise everything their choices would probably be the ones which won out.
So dates in the US are not RFC 3339, they're MM/DD/YYYY drunk-endian. And units are all US Customary, not SI. If you want something else, you need a locale other than en.US, and if you want other things (like spell check) to match US conventions you end up needing to create a custom locale.
Because these have no relation to how a particular piece of content should be formatted. The fundamental mistake made by many of these legacy APIs with respect to localization is the assumption that the locale should be determined based on some property of the user, which is reflected in some OS or application-wide setting that applies to all content. This only works as a rough approximation in the small minority of cases where the user is a member of a cultural sphere / bubble where exposure to multiple languages is a rare exception, like the united states. In the rest of the world, it is an everyday occurrance for multiple languages to exist side by side, including within individual web pages, documents, spreadsheets, and other content, which is why the only reasonable solution is the one where a locale is a property of a piece of content in its most atomic form, not of the user.
Maybe I'll have to write my own locale file. That doesn't sound fun.
My Android phone is set to a language that doesn't match the country I'm in, and it hyphenates all the local phone numbers wrong. This is beyond stupid. The phone knows which country the numbers belong to. I would expect it to format each number according to the conventions of the country they belong to, not according to the language I've selected. I'm not doing anything fancy like RTL, either.
The part that doesn't make sense in C and other languages is that locales are forced on you as global state. And not just by default, they don't even provide locale-free alternatives, you have to modify the global state and reset it after every use, which is cumbersome, slow and error prone. C++ added the locale-free `std::to_chars()`, so things are slightly improving at least, but it's still an ugly and largely unnecessary mess.
> obviously you should not use sprintf, you should use C++ iostreams
Friends don't let friends use iostreams.
I live in a country where the elders of the language decided that we don't format floating point numbers with a decimal point but use a decimal komma instead.
So PI is not 3.14 but 3,14 instead.
Just imagine what pain you have to go through if you want to parse a .csv file written by an application that "tried to do everything right and use locale".
How would you write 3,141.592?
if it's money, you use the . as the decimal separator everywhere. if it's just a number, in the French parts, it's , in the German parts the .
The thousands separator is ' everywhere
Well, there are some exceptions, like actually caring about the original datetime, but they are rare.
Microsoft Excel will helpfully use semicolons as field separators in this case, making parsing CSVs that came out of it in an unknown locale situation even more fun!
Agreed. Had this discussion enough times that I put my thoughts in an article. <iostream> is just spectacularly bad.
https://www.moria.us/articles/iostream-is-hopelessly-broken/
printf("The object is ");
print_object(object, stdout);
printf("\n");
Does fmt support printing a custom type without breaking the format string and without using a temporary string?https://fmt.dev/latest/api.html#formatting-user-defined-type...
It may look like a lot to implement, but it worked flawlessly for me several years ago for a simple case, just by copying and pasting the example code there.
fmt has a “buffer” class which is used as an interface between formatters and their two primary use cases, which are formatting to strings and formatting to files. A buffer is an abstract base class which exposes a pointer to a region of memory and has a virtual function to flush the output (sort of). When you call fmt::format or fmt::print, you’re getting std::back_insert_iterator for a buffer.
I would describe this as “surgical usage of a virtual function” because it is used exactly in a place where it is not called often (you mostly write to a buffer, and flush it less often) and its use reduces the number of template instantiations in your project. That said, the ergonomics for custom formatters is not great.
I wonder should we resubmit this on its own to HN?
Global state considered harmful.
Why not a write lock that allows multiple concurrent readers?
sprintf_l can be as slow as before if you pass in the same global locale object.
What matters is to make the non-_l functions faster and more predictable, which is what you achieve by taking locale out of the equation.
Why would you ever mutate a locale object? Is that the common way to change locales in C? Wouldn't it make more sense to have locale objects be roughly immutable? It doesn't seem like they should have any real reason to change very often in a typical use-case. I would think any given person only has a small (1-3 or so) number of locale's they use on any regular basis.
Are locale objects being mutated really common enough that you need a mutex to protect against accidentally rendering something in the wrong locale?
EDIT: Nevermind, sprintf_l isn't part of the C standard, so really they could be implemented however the authors chose.
I ended up using nanoprintf — it's a single header file and in the public domain.
Given that this talks about a problem in the Microsoft standard library, wouldn't the usual internet advice be "use clang, and also llvm's libc++"? If clang just compiles the same slow code as msvc, it won't magically make it fast.
I18N was one of the few things POSIX didn't do well. Have an environment variable change the sort command causes lots grief.
Nows a good time to rant about Apple's entirely opaque bug tracking process. Every bug you file is private. The only way to know if a bug has already been filed is to file a new issue and see if they close it as a duplicate or not.
(We had to create a set of locale-resistant and consistent versions of shell tools for a system that made heavy use of shell tools to process large amounts of data across thousands of machines. All you needed was one misconfigured locale on one machine and the result would be chaos).
You pay the cost every time rather than when you actually care about it. And when you really don't want locale to interfere it still comes back to haunt you if you don't pay special attention to it. (Remember how Python had locale-dependent XML-RPC that made sure two machines with different ways of formatting floats behaved?).
Locale is bad design. Very bad design.
Sounds like there's plenty of opportunity for library-level improvements here (as well as application-level workarounds), but certainly the sprintf_l(..., locale, ...) being slow is the most surprising to me, and likely the easiest to fix.
> Given that this is an Apple operating system, we might know it has a snprintf_l function which takes an explicit locale, and hope that this would make it scale. Just pass NULL which means “use C locale”:
And then the chart and discussion following it?
They were all correct. But yeah, scary how primitives can result in such poor performance. FWIW, they ended up using {fmt}: https://developer.blender.org/D13998
The point parent poster os making, is that previously, unless set_locale was being called by blender, the resulting export of a blender object was locale dependent. The change from sprintf(..) to sprintf(..., NULL), would then actually change the behavior (not just performance) of the program.
https://github.com/rurban/safeclib/blob/master/src/str/vsnpr...
WSL2 is not that bad, but installing a real Linux in a VM takes about 10 minutes, so this seems a weird thing to say.
The author provides several graphs in which what appears to be the total execution time stays constant as the number of threads varies, and this is described as "good scaling".
How do I know what's being graphed is the total execution time?
> Converting two million numbers into strings takes 100 milliseconds when one CPU core is doing it. When all eight “performance” cores are doing it, it takes 1.8 seconds, or 18 times as long.
This corresponds to a curve where "one core" takes the value 100 and "8 cores" takes the value 1866.
But isn't constant execution time as we increase from one thread to eight threads terrible scaling? What's happening here?
Plotting 1/runtime makes interpreting the actual meaning of a single point much harder. So that is also out of the question.
I took this approach to mean "threads don't affect eachother's performance". Which is easily seen to be equivalent to perfect scaling.
Why? If the whole can be less than the sum of the parts, it can also be greater than the sum of the parts. Maybe two threads can do double the work in 150% of the time. But that would make for a funny definition of "perfect".
It can't be. That's not possible with CPU cores. It can be less but it can't be more.
Here's a proof: you can always timeshare two threads on a single core. If two threads can do 2x the work in 1.5x the time, then you can run that same code on one timeshared core to do 1x the work in 0.75x the time. Thus we have the concept of perfect scaling where double the cores can do, at most, double the work; you can't do better than that because whatever technique you used to achieve it can still be applied back to the single core.
(I'm sure other proofs exist, but the above should be sufficient to show why you can't beat perfect linear scaling.)
The reverse is not true. Just add in a mutex and your perfect scaling is ruined, because some of the CPUs have to wait. The more CPUs you add, the greater the amount of wasted CPU time because the mutex causes one part of the computation to run on a single core.
How often does the locale change in practice? Almost never.
ETA: It's also quite common in engineering blogs for languages, libraries or frameworks, an entry detailing how in the new version they have made a performance improvement by making the fast case faster, or the predictor more accurate, and then removed option 2 from the decision tree, so that we get a bigger benefit from the happy path and the average case, and as a benefit the system is now simpler as well.
Bjarne has always used the expression to mean it doesn't cost more than if the same code has been manually written by hand.
An Assembly written version of iostreams, with the same architecture, would perform the same.