HNHacker News
TopNewBestAskShowJobs

vitaut

1,365 karma · joined May 12, 2014

Carbon-based open sourcerer, code necromancer and a former alien. Author of C++20 std::format and http://github.com/fmtlib/fmt. Opinions are not mine.
submissionscomments
vitaut··on Go 1.27
The core of newer methods like yy, xjb and zmij is remarkably simple: https://vitaut.net/posts/2026/yy-dtoa/. Shortest uscale is basically Schubfach or, rather, it's variant called Teju Jagua and has 2-3 wide multiplications compared to 1 for newer methods.

The complexity is optional and comes from squeezing the last few nanoseconds =).

vitaut··on Tesla is recalling its cheaper Cybertruck because the wheels might fall off
Interviewer: Mr. Musk, I understand the wheels fell off the Cybertruck.

Musk: Well, that’s not very typical. Most vehicles are designed so the wheels don’t fall off.

Interviewer: But these ones did.

Musk: Well obviously. That’s why we recalled them. But wheel retention remains a very high priority at Tesla.

Interviewer: What caused it?

Musk: A minor component interaction that generated maximum freedom.

Interviewer: Freedom?

Musk: For the wheel.

vitaut··on The hidden compile-time cost of C++26 reflection
The binary bloat is also caused by unnecessary inlining and the linker eliminates most of it (but it's still annoying e.g. for godbolt). {fmt} supports a superset of std::format and std::print features including localization. stringstream's bloat is unrelated and mostly caused by large per-call binary code from concatenation-based API.
vitaut··on The hidden compile-time cost of C++26 reflection
std::print author here. Indeed, std::print shouldn't be expensive to compile, it's just a thin wrapper around a single type-erased function. The only reason why it is expensive in libstdc++ is that the type-erased function is inlined which goes against the proposed design but unfortunately can't be enforced via the standard wording and remains a Quality of Implementation (QoI) issue.

Fortunately, libstdc++ is fixing this: https://gcc.gnu.org/pipermail/gcc-patches/2026-March/710275..... There is still more work to optimize the includes but it's a good start.

vitaut··on C++ Modules Are Here to Stay
Modules have been working reasonably well in clang for a while now but MSVC support is indeed buggy.
vitaut··on C++ Modules Are Here to Stay
This style is used in {fmt} and is great for documentation, especially on smaller screens: https://fmt.dev/12.0/api/#format_to_n
vitaut··on C++ Modules Are Here to Stay
We did see build time improvements from deploying modules at Meta.
vitaut··on Banned C++ features in Chromium
The main effect of this is that some of the conversions between char and char8_t are inefficient.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
I was impressed how fast the Rust folks adopted this! Kudos to David Tolnay and others.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
Note that ~3-6ns is on modern desktop CPUs where extra few kB matter less. On microcontrollers it will be larger in absolute terms but I would expect the relative difference to also be moderate.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
I don't have exact numbers but from measuring perf changes per commit it seemed that most improvements came from "printing" (e.g. switching to BCD and SIMD, branchless exponent output) and microoptimizations rather than algorithmic improvements.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
If you compress the table (see my earlier comment) and use plain Schubfach then you can get really small binary size and decent perf. IIRC Dragonbox with the compressed table was ~30% slower which is a reasonable price to pay and still faster than most algorithms including Ryu.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
It is possible to compress the table using the technique from Dragonbox (https://github.com/fmtlib/fmt/blob/8b8fccdad40decf68687ec038...) at the cost of some perf. It's on my TODO list for zmij.
vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
Note that it has the same table of powers of 10: https://github.com/rsc/fpfmt/blob/main/bench/uscalec/pow10.h
vitaut··on Banned C++ features in Chromium
Somewhat notable is that `char8_t` is banned with very reasonable motivation that applies to most codebases:

> Use char and unprefixed character literals. Non-UTF-8 encodings are rare enough in Chromium that the value of distinguishing them at the type level is low, and char8_t* is not interconvertible with char* (what ~all Chromium, STL, and platform-specific APIs use), so using u8 prefixes would obligate us to insert casts everywhere. If you want to declare at a type level that a block of data is string-like and not an arbitrary binary blob, prefer std::string[_view] over char*.

vitaut··on Floating-Point Printing and Parsing Can Be Simple and Fast
The shortest double-to-string algorithm is basically Schubfach or, rather, it's variation Tejú Jaguá with digit output from Dragonbox. Schubfach is a beautiful algorithm: I implemented and wrote about it in https://vitaut.net/posts/2025/smallest-dtoa/. However, in terms of performance you can do much better nowadays. For example, https://github.com/vitaut/zmij does 1 instead of 2-3 costly 128x64-bit multiplications in the common case and has much more efficient digit output.
vitaut··on Is Rust faster than C?
Other examples are CTRE (https://github.com/hanickadot/compile-time-regular-expressio...) and format string compilation (https://fmt.dev/12.0/api/#compile-api). The closest C counterpart is re2c which also requires external tooling.
vitaut··on Is Rust faster than C?
It's easier to write faster code in a language with compile-time facilities such as C++ or Rust than in C. For example, doing this sort of platform-specific optimization in C is a nightmare https://github.com/vitaut/zmij/blob/91f07497a3f6e2fb3a9f999a... (likely impossible without an external pass to generate multiple lookup tables).
vitaut··on Zmij: Faster floating point double-to-string conversion
Please note that there is some error in your port:

Error: roundtrip fail 4.9406564584124654e-324 -> '5.e-309' -> 4.9999999999999995e-309

Error: roundtrip fail 6.6302941479442929e-310 -> '6.6302941479443e-309' -> 6.6302941479442979e-309

Error: roundtrip fail -1.9153028533493997e-310 -> '-1.9153028533494e-309' -> -1.9153028533493997e-309

Error: roundtrip fail -2.5783653320086361e-312 -> '-2.57836533201e-309' -> -2.5783653320099997e-309

vitaut··on Zmij: Faster floating point double-to-string conversion
Yeah, that's what I meant.
vitaut··on Zmij: Faster floating point double-to-string conversion
Should be fixed now.
vitaut··on Zmij: Faster floating point double-to-string conversion
My bad, you are right. The small integer optimization should be switched to a different output method (or disabled since it doesn't provide much value). Thanks for catching this!
vitaut··on Zmij: Faster floating point double-to-string conversion
I started a section to list implementations in other languages: https://github.com/vitaut/zmij?tab=readme-ov-file#other-lang.... Once yours is complete feel free to submit a PR to add it there.
vitaut··on Zmij: Faster floating point double-to-string conversion
It converts 1.0 to "1.e-01" which reminds me to remove the trailing decimal point =). dtoa-benchmark tests that the algorithm produces valid results on its dataset.
vitaut··on Zmij: Faster floating point double-to-string conversion
Cool, please share once it is complete.

C++ also provides countl_zero: https://en.cppreference.com/w/cpp/numeric/countl_zero.html. We currently use our own for maximum portability.

I considered computing the table at compile time (you can do it in C++ using constexpr) but decided against it not to add compile-time overhead, however small. The table never changes so I'd rather not make users pay for recomputing it every time.

vitaut··on Zmij: Faster floating point double-to-string conversion
It depends on the input distribution, specifically exponents. It is also possible to compress the table at the cost of additional computation using the method from Dragonbox.
vitaut··on Zmij: Faster floating point double-to-string conversion
I am pretty sure Dragonbox is smaller than Ryu in terms of code size because it can compress the tables.
vitaut··on Zmij: Faster floating point double-to-string conversion
Think about things like logging and all the uses of printf which are not parsed back. But I agree that parsing is extremely common, just not the same level.
vitaut··on Zmij: Faster floating point double-to-string conversion
> Unlike formatting, correct parsing involves high precision arithmetic.

Formatting also requires high precision arithmetic unless you disallow user-specified precision. That's why {fmt} still has an implementation of Dragon4 as a fallback for such silly cases.

vitaut··on Zmij: Faster floating point double-to-string conversion
Thank you! The simplicity is mostly thanks to Schubfach although I did simplify it a bit more. Unfortunately the paper makes it appear somewhat complex because of all the talk about generic bases and Java workarounds.
Page 1 of 5Next →