Rust vs. C++ Formatting
brevzin.github.io
brevzin.github.io
Also, I do agree that C++'s approach to formatting the date was more user friendly/ergonomic, but on the other hand it's not something I've struggled with. I also feel like there is a more succinct version (I don't know if it would have the same performance characteristics and how much that mattered when writing the post).
They also later mentioned overall consistency and coherence in the Rust ecosystem, and I value that more over a couple of hand crafted formatters in application code
For example suppose we've got 64MB of RAM, and we've got the iterator 0..i32::MAX
If we try to collect() that, we're going to have a bad time, but while that's a lot of numbers, we certainly could write it to a file, we probably do have gigabytes of disk space, or we could send it over the network - the Internet certainly has storage for these numbers.
So at best you could only do this with Clone'able Iterators, but even then it might not be ideal for formatting to create "wasteful" clones. So maybe you'd only do it for individual Iterator types like Range or std::slice::Iter for which cloning is cheap. But even .map()ping such a trivially-Cloneable Iterator would yield a non-Cloneable or non-trivially-Cloneable Iterator, so at that point you have to wonder if it's worth making a small subset of Iterators have these "useful" Debug impls.
Regarding the mutability and collect issues you raised, I think I just explained poorly. What I meant is that the author could've done something like this https://play.rust-lang.org/?version=stable&mode=debug&editio...
However, debug printing "on the fly" (without collection first) would definitely be more involved with interior mutability and require a wrapper type, for the reasons you outlined, but not impossible.
For projects that are user facing and so need to do localization, you do have a few options, including both gettext and ICU at opposite ends of the scale of "How difficult is your problem/ how much work do you want to do?"
Gettext is just going to let you substitute translations like "X {1} Y {2}" in one language versus "YorX: {2}or{1}" in another language, whereas ICU understands how to render the Japanese form of today's date and which is the correct plural form for 38 of this thing in Russian (but it's still on you to prepare all the plural variants for each type of thing you want to quantify, ICU just tells you which of them to use)
gettext seems pretty difficult to use in Rust - there's no obvious way to apply a runtime format string "My name is {1}".
https://blog.hackeriet.no/rust-and-translation-files/ describes using gettext from Rust but not how to handle format strings.
Maybe if you bake in assumptions about the deployed environment at compile time you can catch errors earlier for your specific environment, but all environments? No. That’s not the responsibility of the standard library and given the language lacks a single specific runtime target- definitely not the responsibility of the language/compiler. This is instead a classic use case for tests and opinionated (optional/third party) libraries instead.
Format strings, for example- are essentially compile time macros in rust in order to avoid all the fun kinds of string handling bugs that C’s implementation allowed- but that means no runtime customization without also defining were alternate formats live (and how they’re verified, etc). Not supporting multiple languages is a feature, not a bug. If you need multi language support, you need to a library to define those semantics in a way that fits or use case.
However, you don't have to use format! to format stuff, you would presumably want an API which better reflects the runtime errors you can now encounter, such as wrong number of arguments, wrong order of arguments, incompatible formats.
You are probably aware, but there is a currently ongoing project to move the format_args macro further into the compiler (it's a builtin macro right now, but not doing much that a proc macro can't do), to do its work during AST->HIR lowering. On top of that, some optimizations are proposed that would be impossible to implement on macros today.
For some of those optimizations, a primitive to allow proc macros to expand macros of their own would make it possible to have with a pure-macro solution. Even just the ability for macros to say that they want their input to be expanded instead of receiving pre-expanded input would be enough. These primitives are not available today, but are possible future extensions.
The design of Rust's standard library is to be as minimal as possible. Full localization support would be too much for a tiny standard library such as Rust's. Contrast this to Java or Go.
In the library ecosystem, there are full implementations of fluent available, and the compiler is being translated. It's not ergonomic to use yet, meaning you have to put the strings into a separate .ftl file (maybe it will never be), but it's quite powerful and developed by a lot of experts on the topic.
Edit: It's almost like the whole world got a lot of work done with the tools they already had.
The "type-safe" means "type-checked" by the compiler for correctness to help prevent bugs. It doesn't mean "safety-as-in-not-dangerous".
On x86 what will happen is the code will compile, but the function is going to read its argument from a floating point register instead of an integer register as it should. This:
1. Is a bug, since a completely unrelated garbage value is going to be printed.
2. Leaks the value of a register, which may be a security issue.
There are still other common issues which can easily turn into vulnerabilities, leaking private process memory, when people pass untrusted strings as format strings with the intention of printing them raw.
So you want a safe print to prevent trivial bugs in general, and security vulnerabilities in particular.
char buf[10];
const char* foo = "wrong?";
int res = snprintf(buf, 20, "What could possibly go %d", foo);
Will compile and do... something...error: format '%d' expects argument of type 'int', but argument 4 has type 'const char*' [-Werror=format=]
You only get into trouble when you use runtime format strings (like passing a user string as first argument to printf)
It also doesn't work if your code isn't a textbook example of being wrong. [0] is _slightly_ more contrived but still suffers all of the exact same problems, despite all of the information being available at compile time.
Type-unsafeness in general also just allows for hard-to-find bugs, since only certain data at runtime will introduce undefined behavior.
Which is just the great https://fmt.dev/latest/index.html that even c++11 projects can use.
std::println is more or less what you would obviously build for a modern language and it's notable because C++ could have provided something pretty similar even in C++ 98, and something eerily similar in C++ 11 but it chose not to.
What it didn't inherit from C was a way to write variadic functions with variadic types, so that had to be home grown.
I don't considers this to be "proper" variadic arguments, because a functions argument has to have a type. and these, as far as I'm aware of don't have one. This is about a powerfull as passing a void**. This is essentially memcopying multiple differently typed into a char* buffer and then passing that buffer. You can than correctly copies them back you have pretty much the same behaviour. Both methodss obviously lacks important aspects of the language abstraction of a function parameter and i don't what that feature can bring to the table that the previous techniques don't.
You can get it indirectly using _Generic().
The fmtlib code uses a C-style macro to try to handle format errors at compile time, which detects whether it has enough consteval and no-ops if it does not. On a modern C++ with enough consteval it does exactly what you'd expect it to do, but on older compilers it does nothing.
The result is that fmtlib on an older compiler gets you Exceptions, at runtime, for a format that was invalid at compile time, and if you upgrade the compiler it magically switches to compile time errors.
Edited to add quote from fmtlib docs:
"Compile-time checks are enabled by default on compilers that support C++20 consteval. On older compilers you can use the FMT_STRING macro defined in fmt/format.h instead. It requires C++14 and is a no-op in C++11."
C++ can obviously do more with boolean types because it's got a more advanced type system, like disallowing values other than true and false to be assigned in code that only uses defined behavior. Boolean itself can take up whatever size the compiler decides it does, same idea as C. But, template types can special case boolean to use bitfields and allow for more space-efficient operations.
When i want to print a string i don't want to worry about the security implications of that. With printf i have to. [0]
And i certainly don't want a turing complete contraption. [1] Also looking at log4j.
And even if everything is correct, it's has to parse a string at runtime. I consider that alone unaesthetic.
>Edit: It's almost like the whole world got a lot of work done with the tools they already had.
The best metaphor i know for this attitude is "stacking chairs to reach to moon". If you don't care about the limits of the tech you will be stuck within it.
I'm time and time again amused how anti intellectual and outright hostile to technological progress the programming profession is. programmers, out of all of them.
Did you propose/implement/release something better than printf?
> I'm time and time again amused how anti intellectual and outright hostile to technological progress the programming profession is. programmers, out of all of them.
Perfect is the enemy of good. Some people talk about getting work done, some people get the actual work done and move on.
In my experience, people with this motto generally produce code which frustrates the whole team.
Being a perfectionist is toxic in its own way, though.
There needs to be a balance. I think that balance is to think and plan a few steps ahead (not too much, as it's counter productive) before hitting the keyboard. I know this sounds a bit like a "d'oh, of course" but it really—and unfortunately—isn't something that people practice; they just think they do.
This is what the article is about? Things much better that printf are a dime a dozed and available since 20 years.
>Some people talk about getting work done,
Like this article does? While you busy arguing that you could do the same thing, but much worse?
How hard was that to implement? Seriously no reason it couldn't have been part of C89. Why wasn't it? Because the compiler writers and the C++ standards committee have no personal use for it. It took 40 years of waiting and five years to get it just barely past the standards committee. If you think no one would strenuously oppose a feature like embed you'd be wrong.
Those guys also have no interest in printf type functions. And improving printf would be a lot more work than implementing #embed.
ld -r -b binary -o foo_txt.o foo.txt
foo_txt.o then has these symbols:extern const char _binary_foo_txt_start[]; extern const char _binary_foo_txt_end[]; extern const void *_binary_foo_txt_size;
So you need to write your own declarations (it doesn't generate a header file).
_binary_foo_txt_size is weird and has to be used as: (size_t)&_binary_foo_txt_size
Or use (size_t)(_binary_foo_txt_end - _binary_foo_txt_start) instead.
These people's "actual work" often ends up causing endless streams of security vulnerabilities and bugs too.
Most of the same people you are referring to don't seem to believe that security vulnerabilities exist or are important enough to care about for some reason, but in the real world these are very important issues.
On the other hand, we have people that apparently wouldn't make a program if they are not guaranteed (by another human being) that it will be safe.
If those people generating bugs and vulnerabilities would had to sit tight waiting for someone to make a safe language to do anything, today the world would be 40 years or more behind.
(safe languages that, sarcastically, were created using all those unsafe tools and insfrastructure)
Also in this real world a trillion of printf are being output right now, and will be for a long long time. Is the world falling apart?
You can also list all the printf CVEs but... how many println! are being output?
Sure, but we're talking about printf here. printf is manifestly mediocre.
I guess 'perfect is the enemy of mediocre' doesn't have quite the same ring.
Technically, it doesn’t have to do that. If a program includes the header declaring printf using the <> header defined in the standard and then calls printf the compiler is allowed to assume that the printf that the program will be linked to will behave according to the standard, and need not compile a call to printf. It can generate code that behaves identically.
A simple example is gcc converting a printf with a constant string to a puts call (https://stackoverflow.com/questions/25816659/can-printf-get-...)
This feels a little defensive, but also pretty out of line with the philosophy of the C++ standards committee. The committee has been aggressively stapling every new leg they could find to that dog for decades. They just chose not to staple this particular leg on until now.
Your comment doesn't bear any resemblance with reality. C++ started with a spartan standard library and only recently did it standardized it's file system API.
Compare that with what, say, POCO already offers. Or Boost. Or java/C#/Python/etc.
What exactly led you to believe that absurdity?
The fact that the standards committee simply chose to just add every feature every other language has.
Again, this take is outright wrong and totally clueless. I mean, the summary of each change introduced by any of the C++ standards is freely available. C++20's most compelling features beyond concepts and modules were small improvements over existing features like lambda captures and template resolutions, or new atributes.
What compells you to make such nonsensical claims?
Literally all the additions between C++03 and C++20.
Here's a small list since you'd rather attack me than review the changelogs, it seems.
- range-based for loops.
- enum class.
- digit separators and binary literals.
- consteval/constexpr/constinit.
- std::move
- std::forward
- std::variant based mock pattern matching.
- lambdas.
- structured binding declarations.
- an ABI for garbage collection.
- coroutines.
- concepts.
C++20 even added a three-way comparison operator.
This is just a random selection.
Second, I'm increasingly associating that brit 70s show "keeping up appearances" with rust specifically that overly blushed face of the woman protagonist.
ADDENDUM:
Even iostreams back in the 1980's was a huge deal at the time, as it showed that it was possible to implement type-safe io in a statically typed, compiled language using language features just like any other library, instead of having it be special purpose.
> For many languages, print is special and the implementation is provided as part of the language implementation.
Sure. Printing variables to strings is definitely worth language-level features imho. It's a bad thing if a language (C++) requires users to come up with extremely complex libraries (fmt/std::format) because the language lacks the features to make such a common operation simple and reliable.
C style printf is a dumpster fire. C++ iostreams are unuseably slow. Modern languages definitely solve this particular problem much better!
https://nim-lang.org/docs/iterators.html#fieldPairs.i%2CS%2C...
Using C-style printf sucks balls. It's extremely error prone and it doesn't support complex types. Using Rust's print system is delightful. I can't make a type error and arbitrarily complex types can be printed with roughly zero effort via {:?} and #[derive(Debug)].
Requiring a macro is no better than what C does. Most of those languages have a less awful macro system than C's, but that's not an actual language advancement. The state of printing is still awful in the languages you mention.
https://github.com/ziglang/zig/blob/master/lib/std/fmt.zig#L...
(Some people even try to write entire programs in dynamic languages, as crazy as that is.)
Regarding "might as well have a makefile rule," it's not really fair -- in practical terms having a seemless way to evaluate functions in your code at compile time is way different than doing codegen. Like, in practice, how would you write a type-safe format function using codegen? You could do it, but it would be very, very gnarly. Very different.
Even Go -- which famously for a while tried to advocate codegen as an alternative to generics (w/ stuff like stringer to codegen print functions for human names for constants) -- resorts to dynamic types & runtime errors for formatting.
So let's see:
formatting approaches we've seen so far:
– dynamic w/ runtime errors (e.g. Go, C#, C, ...)
– static via baked into compiler support
– static via special-case restricted compile-time evaluated language (e.g. rust macros)
- static via full-lang compile-time eval (e.g. zig)
I think you said all those are awful, so, I'm curious -- what's the better approach you have in mind?
I do think there's space for a "write a complex value literal by writing a string in a DSL and embedding that into the source" feature. But that shouldn't be specific to format strings, and it shouldn't be by running arbitrary code at compile time; rather it should be a language feature. I haven't seen a version of that that I'm really happy with yet, but Haskell's OverloadedStrings or Scala's StringContexts are some baby steps in the right direction.
My haskell knowledge is minimal, but looking at formatting and fmt library examples, what stands out is you're operating in the syntax of the language, so it's clumsier to use (imo) than a string with interpolations. (It seems not-dissimilar to C++ streams, to me.)
I guess your second paragraph gets at that. I think you're arguing for, "build it into the language, but flexibly." Fair enough!
I disagree in the strongest, most emphatic way possible.
As a user of programming languages the difference between printing in C and printing in Rust is night and day. Whether that is achieved with a powerful macro system or with the core language does not matter to me in the slightest. It carries zero weight.
Writing code that prints variables in C is painful and bad. Writing code that prints variables in Rust is easy and good. If you'd like to say that "easy and good" is the same as "painful and bad" then I disagree.
When std::format ships the gap between C++ and Rust will shrink dramatically. However Rust will still have an advantage with derive macros. C++'s lack of reflection continues to be a major pain point. Maybe in C++29.
I don't think you can fully separate "language" from "macros". If something can be implemented easily in a macro then there is less reason to bake it into a new feature in the language. I don't think it's a good idea to add language features solely so you can say it's a language feature and not a macro feature. YMMV.
String interpolation is part of the language in Swift.
what's the point of not having print provided as part of the language implementation, though?
Also, it's a fairly simple to decide upon dividing line between the language and the library - if it can be implemented in a handful of assembly instructions (or, ideally, 1 instruction) it's part of the language. If it can't and needs a "function" (either in C, or implemented in assembly) to work, or needs to talk to some other part of the computer (like a kernel, or a BIOS) it's part of the library.
Like malloc()/free() - again, not part of the core language spec, but part of the standard library, which can be omitted in some ("freestanding", as opposed to "hosted") implementations.
Remember, C was created in the '70s. Memory was measured in kilobytes. Tens of kilobytes if you were lucky. Even through the '80s, 1 megabyte was a lot.
world = "earth"
print(f"Hello {world}")