Parsing can become accidentally quadratic because of sscanf
github.com
github.com
> sscanf() converts the string you pass in to an _IO_FILE* to make the string look like a "file". This is so the same internal _IO_vfscanf() can be used for both a string and a FILE*. However, as part of that conversion, done in a _IO_str_init_static_internal() function, it calls __rawmemchr (ptr, '\0'); essentially a strlen() call, on your input string.
Looks correct. Call flow:
https://github.com/bminor/glibc/blob/21c3f4b5368686ade28d90d...
https://github.com/bminor/glibc/blob/21c3f4b5368686ade28d90d...
https://github.com/bminor/glibc/blob/21c3f4b5368686ade28d90d...
https://github.com/bminor/glibc/blob/83908b3a1ea51e3aa7ff422...
> It looks like the problem is that reading floats is done using c4::from_chars -> c4::atof -> sscanf which ultimately calls _IO_str_init_static_internal in the glibc implementation. As far as I understand, this essentially causes strlen to be called on the string that rapidyaml passes to it, essentially causing strlen to be called O(n^2) times.
As such, this doesn't follow. Just because sscanf takes the length of the entire source string doesn't mean will get N^2 behavior, unless you do silly things with sscanf, like extract a small token from the beginning of a large string, and then skip past it to do the same thing again after that small token: essentially, build a lexical analyzer in which the pattern matcher consists of calls to scanf. This is not an intended use case for sscanf. No self-respecting language parser designer would do such a thing in a non-throwaway, production program.
sscanf has undefined behaviors for out-of-range numeric values, so you can't use it for just blindly extracting numeric tokens at all, regardless of any other issues.
sscanf should only be used on a string that your program has either prepared itself, or carefully delimited by some other means.
For instance, suppose we have a string which holds a date (could be an input to the program). We can test whether it's in the YYYY-MM-DD format, and get those values out:
int year, month, day;
char junk;
if (sscanf(input, "%4d-%2d-%2d%c", &year, &month, &day, &junk) == 3)
{
/* we got it */
}
The purpose of junk is to test that there is no trailing garbage character. If there is, it will be scanned and the return value will be 4. If we don't care about that, we can just omit that argument and conversion specifier.This will accept inputs like 01-1-1. Worse, it will scan leading signs!
If you wan to insist on nothing but digits, and exact digit counts, you can break it up as string tokens first:
char yearstr[5], monthstr[3], daystr[3];
char junk;
if (sscanf(input, "%4[0-9]-%2[0-9]-%2[0-9]%c", yearstr, monthstr, daystr, &junk) == 3)
{
int year = atoi(yearstr);
int month = atoi(monthstr);
int day = atoi(daystr);
}
atoi is safe here, by the way, because we know that the integers are decimal strings no more than four digits.This is the sort of thing where sscanf comes in very handy. If you know how to use it, it solves these sorts of things nicely, and the performance is probably okay. At least not degenerately bad.
Complete sample with all the validation:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
int main(int argc, char **argv)
{
char yearstr[5] = "", monthstr[3] = "", daystr[3] = "";
char junk;
if (argv[0] && argv[1] &&
sscanf(argv[1], "%4[0-9]-%2[0-9]-%2[0-9]%c", yearstr, monthstr, daystr, &junk) == 3 &&
yearstr[3] && monthstr[1] && daystr[1])
{
int year = atoi(yearstr);
int month = atoi(monthstr);
int day = atoi(daystr);
printf("%d:%d:%d\n", year, month, day);
return 0;
}
return EXIT_FAILURE;
}
Some tests: $ ./date asdf
$ ./date 1-2-3
$ ./date 2020-01-1
$ ./date 20-01-01
$ ./date 2020-01-01
2020:1:1
$ ./date 2020-01-01z
$ ./date 2020-01-12
2020:1:12So in other words, it's completely useless?
> This is not an intended use case for sscanf.
Then what is?
My guess at the historical intention of *scanf is that it was meant for parsing tab-delimited and fixed-width records from old-school data files.
CSVs are such a pain if you have freeform text (in which you have to handle newlines and escaped quotes).
It seems like using FS/GS/RS/US from the C0 control codes would be divine (provided there was implicit support for viewing/editing them in a simple text editor). I get that it's a non-printable control character, and thus not traditionally rendered... but that latter point could have been done ages ago as a standard convention in editors (enter VIM et. al., as they will indeed allow you to see them).
* I find dead keys irritating beyond reason, so I do not count them as an option
TSV also works more reliably than CSV, because most people don't put tab characters in the data in these kinds of records. Tab is even the default field delimiter for cut. But everybody uses CSV, because again it's easier to reason about the above. (shrug)
most parts
Why can't you be normal? https://commons.wikimedia.org/wiki/File:DecimalSeparator.svg
Rob Pike politely declined the patch, commenting "I might prefer to leave it unaddressed for the moment. Proper locale treatment is coming (I hope soon) and this seems like the wrong place to start insinuating special locale handling into the standard library."
Three years later another Go team member commented that "Date localization is definitely still in the planning."
We're in year 8 now. The issue is still open. Rob Pike is still hoping.
№ of countries ≠ № of people though. ⅘ of the 5 most populous countries¹ use decimal points, and together they alone² have ~42½% of the world’s people. Do all decimal point countries together have ≤50%? (Edit: Is it really not “normal” to do things the Anglophone way?)
¹According to https://en.wikipedia.org/wiki/List_of_countries_and_dependen...
²Not counting territories
https://github.com/golang/go/issues/6189
https://github.com/golang/go/issues/26002
In the entire TXR project, I used sscanf in one function: a floating to integer conversion routine. The context is that integers can be bignums, so we can't just rely on ordinary conversion, except in the fixnum range.
http://www.kylheku.com/cgit/txr/tree/arith.c?h=txr-252#n2912
The approach taken is to sprintf the double value into a generous buffer, with a generous precision, using the %g format which will automatically use the e notation if necessary.
We detect whether the e notation was produced, and other characteristics, and then proceed accordingly.
I see that an impossible case is being tested for; since values with a small magnitude are handled via the fixnum route, we will never hit the case where we can short-circuit to zero.
Can anyone see a bug here, by the way? :)
But we are not continuing to scan into the junk here to get more tokens here; we don't have a loop that will cause O(N*N) behavior.
E.g. this could be user input from a form field where they are expected to type a date, and we want to reject that if there is trailing junk. At that point, the user gets an error, and has to fix the input. We don't look for anything in the trailing junk.
It's not about "scanning into junk". The problem would happen if you used scanf to extract integers from a string of `' '.join([1]*1000000)`. Because you would make 1,000,000 scanf calls, and _each_ of those calls would call strlen on the string, quadratically scanning over the string to repeatedly find the end.
In fact, every sscanf is O(size of the string) even when the string is well formed. This is not a well behaved function.
char * start = data;
char * next = strchr(data,'\n');
while(next)
{
*next=0;
sscanf(start, ...)
start=next+1;
char * next = strchr(start,'\n');
}the sum of all scanned bytes is quadratic to the length of the huge input string.
?
edit: oh the autocorrect...
> extract a small token from the beginning of a large string, and then skip past it to do the same thing again after that small token: essentially, build a lexical analyzer in which the pattern matcher consists of calls to scanf.
Scanning floats with sscanf means undefined behavior on a datum like 1.0E999 against a %lf (double), where the exponent is out of range. IEEEE double goes to E308, if I recall.
C99 says "Unless assignment suppression was indicated by a[n asterisk], the result of the conversion is placed in the object pointed to by the first argument following the format argument that has not already received a conversion result. If this object does not have an appropriate type, or if the result of the conversion cannot be represented in the object, the behavior is undefined."
I'm seeing that on glibc, 1E999 scans to an Inf, but that is not a requirement.
This problem could be solved with sscanf, too, while maintaining the degenerate quadratic behavior. sscanf patterns could be used to destructure a floating-point number textually into pieces like the whole part, fractional part and exponent after the E, which could then be validated separately, like that the E part is between -308 and 308 and so forth. Someone could still do this in a loop over a string with millions of floats. I'd say though that you're way past the point of needing a real lexical analyzer, and not multiple passes through scanf.
Do I need a full fledged parser for that? No. Does it have to traverse the whole string every time? No, because I’m being very clear about the maximum length of the thing I’m trying to parse.
You cannot not conclude that it’s a bug.
A mechanism for turning a string in memory into a string should not be taking its length.
GBLIC has an extension called fmemopen. The API is:
FILE *fmemopen(void *buf, size_t size, const char *mode);
This interface smells like it might have the flaw that if someone wants "buf" to come from a null-terminated string, they must measure that string in order to supply the size parameter.I.e. the underlying mechanism is binary, and does not stop at null characters.
There needs to be a separate mechanism with a different underlying stream implementation:
FILE *fstropen(const char *str); // mode implied: "r"
This doesn't require a size.That fmemopen extension points to a solution to the original problem, though non-portable: open the entire buffer of numbers as a stream in one fmemopen call, and then use fscanf to march through it.
https://pubs.opengroup.org/onlinepubs/9699919799/functions/f...
All your magic on format string doesn't change that poor behavior unless you manually add '\0' somewhere to null-terminate the input string early.
If you parse the year-month-day in 100 places of that input string, it will strlen 100 times. And now replace the "100 places" to a linear function of the input string length (strlen(input) / 11?), that is quadratic.
As someone who's been learning some C, I wonder if you have a recommendation for a guide on how to handle strings properly.
There are at least a few libs like sds that store the length information in various ways, which solves at least this kind of problem.
(Languages with more advanced string types and slicing mechanics are not impacted; C++, Rust, Go, etc. They do have their own issues of course.)
Then why does it support the "%n" specifier? Isn't it exactly for that use case?
It's almost like JSON and XML were invented so people didn't have to understand lex/yacc.
And behold... they don't.
There are two basic implementation strategies. The BSD (FreeBSD and OpenBSD and more than likely NetBSD too), Microsoft, GNU, and MUSL C libraries use one, and suffer from this; whereas the OpenWatcom, P.J. Plauger, Tru64 Unix, and my standard C libraries use another, and do not.
The 2002 report in the comp.lang.c Usenet newsgroup (listed in that discussion) is the earliest that I've found so far.
There have been several assertions that library support for strings in C++ is poor. However, note that the fix here was to switch from the Standard C library function sscanf() to a C++ library.
* https://github.com/fastfloat
One might advocate using other people's libraries, on the grounds that they will have got such things right. However, note that this is RapidYAML, and the bug got fixed several years after GTA Online was released. Other third-party parsing libraries have suffered from being layered on top of sscanf() in other ways.
* https://github.com/json-c/json-c/issues/173
* https://github.com/kgabis/parson/commit/96150ba1fd7f3398aa6a...
MUSL, with the same implementation strategy of nonce FILE objects, ends up performing multiple passes over those big input strings, once to look for NUL and a second time to pull each character out one by one. And it's doing those multiple passes multiple times, as it sets up and tears down the nonce FILE object each call to sscanf(), meaning that it's scanning the same parts of the string for NUL over and over again as it repeats the same buffer fill(s) from the input string a few characters further on. As I said before, it's not entirely free of this behaviour. With the other implementation strategy, in contrast, there's only a single pass over the whole input string, even with those multiple calls to sscanf().
It fills in a buf with a string_read callback, but if we look at what vfscanf does, it uses shgetc, which eventually calls into __uflow, which is where the read is performed:
https://git.musl-libc.org/cgit/musl/tree/src/stdio/vfscanf.c https://git.musl-libc.org/cgit/musl/tree/src/internal/shgetc... https://git.musl-libc.org/cgit/musl/tree/src/stdio/__uflow.c
This will call into the string_read callback with a buf size of 1. The string_read will internally do a memchr(buf, 0, len+256);, so, OK, it's memchr-ing at most 257 bytes, sorry. The read data then gets stored in a stdio-internal buffer (the rpos/rlen below), so this code path will actually be skipped on future reads.
If it then needs to unparse a character, it will use ungetc to put it back in a stdio-internal buffer. Reading this code isn't that hard, I don't know why you keep insisting it does something else.
Sure, it's scanning a bit more than maybe it needs to, but it's much closer to a non-scanning libc in the amount of data it's consuming.
It still doesn't guarantee that the behaviour is stable (because dependencies within the linked range may change before a reader gets to it) but it can help support your message (and hopefully your argument) both at time-of-authorship and in future.
(and yep, it's a tradeoff of your own time -- and response time -- against future readers' time; often tricky to handle. I think code browsers could do a lot more to make it easier)
Of course, the underlying vfscanf() implementation that reads from a FILE* simply proceeds one character at a time; the pseudo-quadratic behaviour arising from the strlen() call is simply an artifact of the simplistic "simulate a file using a string" wrapper in sscanf().
[1] https://minnie.tuhs.org/cgi-bin/utree.pl?file=4BSD/usr/src/l...
In general there are three cases:
- Developer forgot to strip/obfuscate symbols, in which case you’re golden
- Symbol you seek is in a shared library: patch the DLL. Often what we did was do investigations in a VM, where the n most common shared library calls were patched to send stack and symbol data to a local DB. Very similar play to RR, but only in the patched code.
- If the (missing) symbol you seek is pre linked (static), things get interesting. We built a pattern matcher that was incredible. Idea is, if you have the disassembly and a loop, like:
int fun(char* str) { call processChar on each uppercase character }
It pattern matched a name like rax_foreach_iflte_cx_call_obfuscatedSymbolForProcessChar
90+% of code either:
- enters the same dynamic library trampoline
- displays very very common programming patterns.
From what I know, we never released this outside early MS DN members because the 10% is often very useful code and developer tools weren’t given away for free yet.
But once the DLL is patched, you propagate names up graph. Any remaining symbols were hand annotated (and we run the propagator again).
Brian Parks went on to extend this system to work with C++ templates. C++ mangling is very predictable, so his program could scavenge for symbols (obfuscators only worked on the root symbol name) and build out the class at compile time for us to reference.
I can see why the devs would not trust that though. So it’s the inefficiency of the dedup’ which bugs me more than the dedup itself.
If you need to parse JSON or some high level format look for a lib, for instance I can vouch for cJSON which is very small and has a rather nice API (by C standards at least).
If I need to parse a simple format string like "id:name:description" something like strtok_r does the job to split the components and then you can parse individual components any way you like, potentially offering better error messages.
If I need to parse a complex format then chances are scanf wouldn't do the trick anyway, so it's either custom made parser state machine or lex/yacc.
Note for instance that in the GTA case the call was `sscanf(p, "%d", buf)` which is a glorified strtol with vastly worse performance.
I would seriously advise not using scanf and friends for anything serious. Too many footguns.
(1) Naive people find liberation in text-based protocols that don't require tools to understand.
(2) With experience one finds that variable-length data structures are much more treacherous than fixed-length data structures: you find simplicity instead in binary data formats that are dead easy to parse.
(3) Floating point numbers are astonishingly expensive to parse and unparse to ASCII compared to how many FLOPs you could do in that time. All that, for a function that isn't quite correct... That is, the number 1/5, which you write as "0.2" doesn't actually exist in IEEE floats. You don't notice at first because that number unparses as "0.2" -- but then note that "0.1 + 0.2 == 0.3" is false in most computer languages.
(4) Null-terminated strings are the definition of insanity, rivaled only by "pointers are integers, integers are pointers"
One pet project of mine lately is a "Data Assembler" that takes a graph of data structures (say test data for a family of logic chips, graphics data for a laser light show, or "maps" to upload to the engine control unit in your car) in Python and converts it to a byte array that is incorporated into an Arduino program.
The total size of the structure could be up to 32kbytes on a small device, or up to 250kbytes or so on a big device, but the program memory is huge compared to the RAM which might be as little as 2kb so the unpacking code is going to stream the data with a small working area.
The minimum it has to do is put together blocks of data that have references to one another that are only resolved once all of the blocks are otherwise complete.
I deliberated a lot about how to handle variable length data structures and decided to embed a two-byte length before the block for all blocks. I could have saved some bytes when the length of the structure was known, but then I'd have to figure out how to code lengths then (e.g. for an array of 2-byte ints do I count the ints or bytes) and figured I'd make up my mind once.
A fixed-length structure is only fixed until you need to add a field to it. I understand the appeal, but good solutions seem to be very rare here. ASN.1 was possibly the first attempt but turned into a total nightmare. Protobuf seems to be doing a reasonable job, but only as a wire format and not a persistence format.
Text always "works", even if inadequately. And it can be handled with your normal diff and editing tools.
I guess I could see that being an interesting property in software that wants to deal with incompatible forks, or software that for whatever reason wants other people to read their data but is unwilling to ship a schema with it, but just for persistence on backend servers it doesn't seem like a very interesting property or worth the extra space requirements.
Windows vs Linux end-of-line markers often screws-up things. UTF often screws-up things. Tab-vs-space sometimes screws-up things. Text is related to Postel's principle, which is often questioned because it leads to sloppy implementations on one side and complex implementations on the other side, and other bug-is-now-a-feature niceties.
I designed a couple of small internal protocols, and have been maintaining them for at least a decade. Maintenance time is nearly 0 because I never had to add a field to them. Even though the software using them evolved a lot and in unexpected ways.
Yes, they stood the test of time because they solved the problem we had, and those problems didn't evolve much. Not every protocol or format needs to be hyper-extendable and scalable.
> even if inadequately
So it works poorly but it works, so it can be shipped. That's why you buy a new smartphone every two year and don't see much improvement in response time or battery life.
As a huge fan of text-based protocols (but not of JSON or YAML), calling a whole class of debuggability, interoperability, and onboarding benefits "naive" is a bit much.
I would trust programmers at Intel to handle FORTRAN-style data structures, and maybe even make a compiler in C.
I wouldn't trust anybody from Intel to write a web server or even a web application with PHP in that wild west of applications programming with strings they are sure to mess it up and thus leave a vulnerability in the management engine on your CPU. It might superficially look like good code, but American Fuzzy Lop will find the branch that isn't obvious, and so will some kid.
That said, what I like about C is that it is a simple language where you can throw out the standard library and start anew: it makes me think of the old versions of ALGOL that didn't specify mechanisms for I/O, and now we know you can specify that behavior in libraries and leave the language pure.
Definitely! I wish the C and C++ communities would take this problem seriously.
> That said, what I like about C is that it is a simple language where you can throw out the standard library and start anew
This definitely is a nice property. But it's not just C that can do this. You can for example do this with Rust, and get an excellent package management system (that can indeed package these alternatives to the stdlib) to boot.
Wasn't ragel responsible for the cloudbleed vulnerability? Wherever possible we really shouldn't be using C at all.
1: https://blog.cloudflare.com/incident-report-on-memory-leak-c... (As for cloudflare, I don't think it's good idea to let single company to MITM the whole internet)
Our tech with regard to type and memory safety is more advanced than our car tech: we already have the tech to prevent these errors. And I believe we should be using it wherever possible. Obviously there will be exceptions (the more obvious being legacy codebases), but they should be exceptions.
They'd recycle my scripts when dealing with very large files but wouldn't learn it themselves.
I don't think text is necessarily the problem here. Other protocols might allow you to use less memory, but they aren't running into a memory limitation in this instance.
You can divide parsers into two groups: parsers designed to read data which follows a specific schema and parsers designed to read arbitrary data.
Schema-aware parsers have additional context which allows them to process the data much more efficiently. The stricter the schema, the more efficient the parser can be. Meanwhile, parsers for arbitrary data have no such context and must rely completely on syntax to parse the data. A poorly optimized schema-aware parser may still outperform a well optimized parser for arbitrary data.
You can still encode data as text and even use popular syntaxes like JSON or YAML with schema-aware parsers. I've personally helped design a schema-aware parser for XML data. But you should keep in mind that most third-party parsers for those syntaxes will be designed to parse arbitrary data. Parsers for arbitrary data can still be very fast - fast enough for many purposes. However, schema-aware parsers have a significant advantage when it comes to computational performance.
I can store fixed-length data using text. A strict schema can specify things like fixed lengths, value padding, orders for maps, exact whitespace rules, etc which can entirely remove the need for delimiters. A schema could even require length prefixes, just like you might use for a non-text format.
These concepts do not necessitate abandoning text. If you need the data to be readable/writable by a human with minimal tooling, you can still leverage text without sacrificing much parsing efficiency.
Syntaxes like XML, JSON, YAML, etc have become popular for their convenience and versatility, but they do not represent the whole of what is possible with text.
Personally I blame the C standard for not splitting strtod in to 1 parsing function (returning an integer and the power of ten), and 1 computation/conversion function. Without this you're basically hosed with respect to locales.
The floating-point numbers should be converted to text as hexadecimal, e.g., in C99 and later, using printf with the %a or %A conversion specifiers.
Even when a human uses the textual floating-point numbers, if they are used only for making copies or comparisons, the hexadecimal numbers are more convenient and they are free of conversion errors.
double v;
memcpy(&v, "\x00\x00\x00\x00\x00\xe4\x94\x40", sizeof(v)); // LECould you elaborate on this one? I've only used C in college and don't know why this is bad.
Or, compare string manipulation in C with any high level language like Ruby, Python, JS and the ease of use is night and day.
But saner ways did not.
Null-terminated strings try to save space by encoding the string length by convention. This can fail due to off-by-one errors, mistakenly allocating a fixed buffer, strings that have an unexpected null in them, and more.
Those can lead to all kinds of nasty problems like buffer overflows, which allow someone who can craft an input to write arbitrary stuff into memory.
God only knows how many security vulnerabilities and performance problems could be traced back to null terminated strings.
1. Such a string cannot contain arbitrary bytes. Since it is frequently useful to work with arrays of arbitrary bytes, we have to repeat every library function twice, once for strings that can contain nuls (the mem* series of functions), and once for strings terminated by nuls (the str* series of functions). The latter are much better supported, and are easier to work with, so we frequently see programs that trip over strings with nuls in them for no particularly good reason.
2. We often want to modify strings, not just read them. Modifications often need more space than was in the original string. A nul terminated string cannot say how much room beyond the current length is allowed to be written to. By default, most library functions assume that there is enough space, however this is frequently an unsafe assumption. This again causes us to duplicate every library function twice over in order to specify the size of the buffer (the strn* series of functions). Since, again, these functions are poorly supported and harder to use, we see many many instances of programs that have buffer overflow bugs, and therefore likely RCEs, for no good reason at all.
On (3): never perform equality tests on FP numbers in languages that use IEEE standard FP format.
On (4): I have worked both with ASCIIZ and counted strings, I think both have their pros and cons. But if you are writing stuff that has any contact with C/C++, it's often a bad idea to go against the stream. No, auto-adding a final 0 to your counted strings won't save the day every time.
There is nothing inherently wrong about a null-terminated string convention.
Basically when it comes to performance we're still in the place that we used to be with correctness: "either read the entire stack of source code to understand everything that's happening, or try and remember to read the docs which may or may not be correct and exhaustive and just hope for the best".
The performance characteristics of a piece of code are part of its contract with the outside world, but that contract is in no way formalized right now.
1: https://github.com/djhaskin987/degasolv/blob/develop/cli-src...
You parse by processing a long string/file via "get-a-token(), add-token-to-structure()" repeating until you have built your structure or made an error in the construction process (here a token a kind, a type of substring and get-a-token retrieves the substring type in some form).
The problems come because the naive programmer is only thinking about their structure building process, so they use some library function to get their tokens. Whether that's a regular expression interpreter or some other flexible facility by provided by the library, it's easy for it's definition of a token-string to include things "not in the string" as well as things in the string - a "quoted string" begins and ends with a quote, so to discover that X isn't a quoted string, you have to scan to the end. So if thing go wrong, each call to get-a-token may scan the while file to make sure this token isn't that. Just an example. The problem here seems to be tokenizing floats, which also has a "where it end" problem in c++. And the alternative is parsing floats yourself, which indeed no one wants to do since they are complex and defined by a complex standard.
Maybe I misremember, but I think long ago people made arguments that C strings were superior because of how simple they were. Such was the thinking 30 years ago.
Older computers had a lot less memory. Programs would write their own virtual memory systems to swap code and data to and from disk. C was born in this era. By the time large flat memory and many registers were available, the old C string was already the standard. It is very difficult to get away from that because everything from the OS on up assume standard C strings.
(Under the heading “is it rapid?” Measuring performance).
> The 450MB/s read rate for JSON puts ryml squarely in the same ballpark as RapidJSON and other fast json readers (data from here). Even parsing full YAML is at ~150MB/s, which is still in that performance ballpark, albeit at its lower end. This is something to be proud of, as the YAML specification is much more complex than JSON: 23449 vs 1969 words.
How did YAML get so complicated. I’d assume it would be much more similarity YAML’s quirks adding a few extra things to work on. But that spec is quite huge by comparison (of words alone).
YAML has both more syntactical options, and broader semantic coverage; particularly, it includes the concept of schemas defining YAML profiles.
Also, the YAML spec has much more informational content. Examples are a stark illustration: RFC 8259 has 5 example JSON documents (3 of which are simple scalar values). The preview chapter of the YAML 1.2 spec has 28 examples, only a handful of which are single scalars.
Meaning sscanf is actually designed to scan all your values in one pass?
https://blog.cloudflare.com/details-of-the-cloudflare-outage...
Lol just say thank you
Also you just have to google "sscanf slow"....