Why is Swift so slow (timeout) in compiling this code? (2022)
forums.swift.org
forums.swift.org
Actually, no. There is no portable "incbin" in C/C++, so one way to bake assets into the .rodata section is to convert a binary file into a huge "const unsigned char data[8192567] = { 0xFF, 0xD8, 0xFF, 0xE0, ... }" string and stash it into a .h or .c file. And yes, people do that often enough that gcc and clang actually had to optimize for this case specifically.
However, even though there's a clear need, as that post explains, the compilers varied between "Bad at this" and "Completely awful at this" which is part of what was so frustrating for JeanHyde over years of getting this through committee.
In principle #embed is just shoving the bytes in as comma separated values, but in practice by the "as if" rule in C and C++ the compiler won't do that because it's stupid.
I'm not sure who posted that comment, but if they are on the Swift team and blaming users, that should stop. Code like that is out there, and it's the job of tooling to handle it gracefully.
Actually it’s common advice to replace very large literals with `JSON.parse(“…”)` because it’s faster according to Chrome engineers [1]. At Notion we did so for our emoji unicode tables for a noticeable time-to-interactive improvement over the large literal.
My understanding of Swift’s teams thoughts, though a few years out of date, is it’s very very hard to get the Swift type inference engine to handle a huge blob, without type hinting, or with literals that could be multiple types.
Big wtf moment.
I don't miss Xcode too much. IIRC I built some custom script using arcanda dropped by Apple employees to output per-function build times. It was great! I knew exactly what functions were slow. But now I am old and lazy lol.
Rather than trying to force Swift to create fixed dimension arrays, the uncomfortable part about Swift for control freaks is that you should just make the array dynamic, and trust that the compiler is smart.
It's not blaming the user, but the compiler team has to balance functionality ("Why can't the type system give me a better error?") and speed. There's no way they can handle every pathological edge case gracefully.
So, an hour's worth of compilation time = 3600 seconds, divided by a million elements, that's 3.6 ms, or about 10 million instructions per element. Just what in the heck is this compiler doing with 10 million instructions on an integer literal? Like, it's never seen integer literals before? And they're all integers. Every element. It's not a "pathological edge case". The whole thread is absurd. People are speaking up and blaming programmers, the language, the type system, when this is so obviously the compiler's fault.
I can barely believe that a single thread, copy pasted quicksort can sort 1M elements in less than 50ms on my computer.
Fast compilers such as TCC should be frequently used as a reality check for compiler teams. Of course the constraints are not the same but still...
This is exactly the attitude that needs to be addressed. It's hostile to argue with users who are doing something completely legal and tell them they shouldn't do that, they should do it another way, that that behavior needs to be discouraged, etc.
We went through this all the time with V8, trying to tell JS developers they just shouldn't do that because V8 had couldn't run that code fast. Or worse, that it had a particular pathology that made certain code absurdly slow. It just doesn't fly. It's V8's job to not go off the rails for user inputs; it should provide good default performance all the time, not get stuck in deopt loops, use absurd amounts of memory, etc. Yeah, and that's hard work.
I would love for the Swift compiler to be dramatically faster, but I understand the challenge, with a powerful type inference engine that supports Generics. It's a resource scarcity problem. If the Swift team spends 10 hours to handle long strings of integer literals, that's 10 hours they haven't put toward features that would benefit a larger audience.
Sometimes it is the fault of the language design too. Maybe the language spec needs to be changed. Imagine all the wasted effort to optimize V8 when you could have put static typing into Javascript itself.
..and if you're in a context where you need that data in the binary, because you don't have a reliable place to store a file?
> That's now extra code to maintain in the compiler just to support a behavior that should be discouraged.
If you don't want to maintain the code to properly support a language feature, don't include that feature in the language.
Now, #embed and include_bytes! always give you bytes (in Rust these are definitely u8, an unsigned 8-bit integer, I don't know what is promised in C but in practice I expect you get the same) because that's what modern files are - whereas maybe you want, say, 32-bit big endian signed integers in this Swift example. But, Swift is a higher level language, it's OK if it has some more nuance here and maybe pays a small performance penalty for it. 10% slower wouldn't be objectionable if we can specify the type of data for example.
I thought the concern was about compile time performance? I'm not sure what Swift being higher-level has to do with that.
The application was a highly optimized encoding of Aho-Corasick state machines for megabytes of static dictionary strings. The code generation (not in C) was trivial, and along with .rodata came all the benefits of shared pages and lazy loading.
Across a number of compilers I only ran into one bug in 32-bit gcc, which was worked around easily by disabling the (unhelpful) optimization pass that was getting snarled.
You're right of course but experienced developers know to write those as strings instead of arrays, especially for very large content. Otherwise both the compiler and IDE become painfully slow. ie your example would be written "\xFF\xD8\xFF\xE0..."
And yes, it is dumb that we have to do tricks like this in 2023.
https://thephd.dev/finally-embed-in-c23
tl;dr It was such a big issue they actually managed to get it through the C Committee
I kinda hate this kind of argument. I mean, it's true, as .incbin is a GNU assembler directive (FWIW: binutils/llvm objcopy is a better mechanism still for this sort of thing in most contexts, as it doesn't involve source compilation of any kind).
But it leads ridiculous design decisions, like "I'm going to write 8MB of source code instead of doing the portability work on my own to turn that data into a linkable symbol".
There's a huge gulf in the space between "non-portable" and "impossible". In this case, the problem is trivially solved on every platform ever. Trivial problems should employ trivial solutions, even where they have to involve some per-platform engineering.
Why is that ridiculous? It strikes me as not necessarily the best, but the most obvious approach, the most portable, and possibly the quickest to implement.
Your toolchain might have a special way to import binary blobs, but a) you’ll have to dig through the docs to find it, b) you’ll probably need to solve the problem again when porting to a different platform, and c) who knows if it actually works, or if there are hidden gotchas?
Sure, if there’s a known tool or option that does the job, you should go ahead and use it. But in general, writing a little script to generate a bunch of boilerplate code is perfectly workable.
This is a corrollary: "I don't want to learn my tools, so I'll learn the language standard instead" is fundamentally exactly the problem I'm talking about.
Straight up: C linkage is a 1970's paradigm full of tools that had to run on a PDP/11, and it's vastly simpler than learning C++ or Rust. It's just not "modern" and no one taught it to you, so it looks weird and mysterious. That's the problem!
Why do we generate big C/C++ strings instead, and compile those? Because object files have lots of compiler/platform/architecture specific stuff in them. Outputting C (or C++) then compiling it is a much more convenient way to generate those object files, because it works on every OS, compiler and architecture. Even systems that you don't know about, or that don't exist yet.
I hear what you're saying and I'm torn about it. I mean, aren't binary blobs just a simpler version of the problem protobuf faces? C code is already the most portable way to make object files. Why wouldn't we use it?
The case at hand is an 8MB static array of integers. Obviously yes, of course, absolutely: you choose the correct/simple/obvious/trivialest implementation strategy for the problem. That's exactly what I'm saying!
In the case of protobuf (static generation of an otherwise arbitrarily complicated data structure with reasonably bounded size), code generation makes a ton of sense.
You lose your bet, because it doesn't, neither its inline assembler nor actual MASM shipped with Visual Studio support any "incbin"-like directives. I guess you can generate an .asm with a huge db/dd, I guess, if you don't like a large literal array in .c files, but that's it.
Right, meaning you have to implement N solutions instead of just one. It's a common enough and useful enough feature for the language to support it. I think it would be a different story if linkers were covered by the language specification.
> I remain shocked at how controversial this is.
I am a bit shocked that you think the right solution to making data statically available to the rest of your program is somehow outside the scope of the programming language.
[1] Which FWIW is much less portable as an issue of practical engineering than assembler or binutils tooling!
It's "I don't want to learn and debug _everybody else who may possible want to build this otherwise portable C program_'s tools."
When those tools change how they do this unportable thing every couple of years in subtle and incompatible ways, which require #ifdef's to handle the different ways those linked against symbols can be accessed, multiplied by dozens of different platforms, then yes, I'm going to compile an 8MB literal.
Also, you argument sounds exactly like those that the author of https://thephd.dev/finally-embed-in-c23 has been fighting against for 5 years straight.
I don’t know what to say except that I’ve worked with C linkers for a long time, since before I learned C++ and before Rust even existed, and I still don’t like ‘em.
people use scripts to do this kind of thing and never think about it again, and it's still more portable than writing a custom build step.
I used to agree with that, but honestly a simple `xxd` to get a C array avoids so many issues with `objcopy` that I'd rather just use that now. With `objcopy` even just getting the names of the produced symbols to be consistent is a pain, and you have to specify the output target and architecture which is just another thing you have to update for more platforms (and if someone's using a cross-compiler, they have to override that setting too).
In contrast if you just produce a C array then it compiles like normal C code and links like normal C code, problem solved and all the complexity is gone.
I understand this might not be a high priority, because it's a somewhat contrived usecase rare to matter in "the real world", but I strongly suspect this must be due to bug(s) in the compiler. I cannot think of a single reason it _has_ to be this way per the design of the language.
I recently started a macOS app project to learn Swift and SwiftUI. A week or two into the project, I managed to hit an honest-to-god compiler bug. A particular combination of syntaxes would reliably cause the compiler to crash (not yield a compile error; crash). I tried restarting XCode, restarting my computer, updating everything. Whenever I brought that syntax into the editor, XCode's Swift compiler integration would crash.
It wasn't anything exotic; I was making a straightforward app, and wasn't trying any Swift language funny-business (I didn't know the language well enough to even try)
It was bizarre to so easily find a bug in the headlining compiler made by the world's most valuable company; I can only think it's the result of this being a language that (effectively) only targets one company's systems, and the effect that has on the size and involvement of the community
"I found a trace of the compiler kind of surprising; the time is not all being spent in the type checker, as you observed:" - 5th reply by scanon
The Java approach is roughly: write types everywhere, keep them simple (and limited in scope), it's easy for the compiler to infer the types.
The Haskell/ML approach is roughly: know a lot about strong type systems, be purely functional, infer most of the types, produce bad error messages when you can't.
The Swift approach attempts to be the best of both worlds, lots of inference, strong capabilities in expressing complex types, but the trade-off is that it can't use the inference algorithms of either of the above examples. I believe Swift necessarily has poor time complexity in type checking – it can't be better with the design it has and requirements it puts on authors.
The trade-off is a somewhat rare sharp edge for Swift and a much easier onboarding, vs Rusts noticeably harder onboarding and more predictable sharp edges.
It managing an even more complex type system without any of the disadvantages seems to make it clearly great!
-Xfrontend -warn-long-function-bodies=<limit-ms>
-Xfrontend -warn-long-expression-type-checking=<limit-ms>
Set these to ~100ms and add a few explicit types whenever they log, and compile time will be significantly improved.Swift is a slow compiler in exceptional cases, maybe in the 99th percentile, but when you know where this is happening it's fairly trivial to avoid it and keep the compiler nimble and stay productive as an engineer.
I’ve worked in a large Swift project with those timeout warnings enabled, and didn’t find them too helpful. They showed up a lot and it was rarely obvious how to quickly fix them. Possibly it would have been useful if those warnings were enabled from the start and developers had always fixed them proactively. I’m skeptical that that could work in a big fast-moving team project, though.
I'm far from a Swift expert, but I usually found I could reduce the compilation time with ~10 mins of work. When I turned these on for a small codebase (60k lines, ~50% of an engineer split across 3 people), it only took a few hours to solve all the instances of this, and they were mostly obvious cases where a 100 line function could be split into 3 and solve the issue.
Also, this sort of works the other way too. Just because something failed at the 100ms threshold doesn't mean it would complete in 101ms, it might take 500ms, 10 minutes, or be undecidable, and it's important to fail CI in those sorts of cases to protect other developers from issues, so having a limit is useful even if it's relatively high for the default case.
I'm a huge fan of error'ing on warnings, but I really don't see it as appropriate for nondeterministic cases.
Personally I think it's the job of tools to make humans productive instead of the other way around.
I discovered very quickly that building a string by concatenating the input one word at a time in Java, my preferred language at that time, was a huge mistake, because of the way string literals work in Java. StringBuilder to the rescue!
At first I was astounded at how slow it was, but when I learned how the underlying language semantics worked, it made sense.
It's worth mentioning another, more common, cause. There can be computational cost explosion when Swift resolves the types for function calls (and operators). You notice it when you get the "too complex" error, but you may be tolerating slow compiles which are just slightly less than too complex. A few explicit type annotations in these situations can do wonders for speeding your development cycle.
Edit: redact I unfortunately can't tell you how to find these. Maybe a "give it a go with reduced complexity limit" compiler flag would help find them.
Edit: Add see reply below for answer.
That said, we should also get these to be fast enough that we have to introduce new flags that take µs instead of ms.
something like this
enum State { case A, … } switch (fromState, toState) { case (.A, .B): … }
in which you want to make sure you didn't forget a combination.
(I could not find the link anymore, but I remember reading a blog post about how Rust `match` statement are in fact solving an instance of SAT).
Instead, I went with Luc Maranget's classic approach and figured out a way to adapt it to a language with subtyping (with a ton of work from Johnni Winther to figure out all of the hard complex cases around generics):
https://github.com/dart-lang/language/blob/main/accepted/fut...
The performance (in the prototype!) was dramatically better. You can always make pattern matching go combinatorial, but I haven't seen any real-world switches get particularly slow with our approach yet, and we have some fairly large tests of matching on tuples of enums.
It breaks at around 50 cases in the switch.
I consider this a major hindrance for productivity.
Fast iteration/test cycle is the best way to stay in the flow while fixing small bugs on the go.
I'd say the limit for me is about five seconds, but I find it much better when I can keep it below one second.
I think that Jai is engineered to have fast compilation in debug mode, I hope that will be the case once released.
But the other part was the reaction amongst developers to this. Layers of abstractions introduced just to please the compiler gods. C++ was particular famous for this, all for the sake of partial compilation still working -- sometimes this had the benefits of decoupling, but often it was a few cake layers on top of that.
So you end up with convenient languages, but have to forsake some of that convenience to actually be productive. But you still had to learn the convenient unused part and know the difference, even if you almost never could afford to use it (if it's just about runtime performance, you often don't need that. But you don't want to cause the projects compilation or test phase to double, just because that functional-combinatoric template metaprogramming approach is so much neater)
Oh man, if only we would have Wirth-dows, not Windows ;)
I think they are making a PR mistake by not making it more open, but I would be surprised if it is never released.
https://github.com/ziglang/zig/issues/89#issuecomment-122118...
Although I'm not sure the "self hosted compiler" is the same as "llvm free backend".
It might have just been a few kb of text but I couldn't compile even with 16gb ram
1M elements array, included in the source code, then sorted with a random C quicksort implementation I found while googling (far from optimal)
Compilation time : 154ms on first run, 95ms on successive runs.
Run time : 693ms on first run, ~105ms on successive runs.
The CPU is a 13900K.
Compilation time is 700-800ms. Run time is about half as tcc, compiled with -O3 (45-55ms)
Running C code at 50% max speed is OK for me when running debug builds.
Clang being slower to compile is not an issue because I can use tcc. Unfortunately I can't use all those nice C++ features with tcc.
I tried to replace the quicksort with a radix sort, and the difference is much closer on this better optimized code (15ms vs 17ms).
Rust compiled in about a minute and the runtime was 2 secs
That seems really slow as well! Since reading the same thing in JSON only takes a fraction of a second.
I think it would be well worthwhile speeding up this kind of thing. Obviously a million-element static array isn’t common usage, but if you can find and fix the bottlenecks there, it will help make everyday usage faster too in lots of small ways.
Parsing a programming language is just always going to be slower than just reading a file into memory, which is why Rust had include_bytes! from the outset and it's silly that C didn't do likewise even in 1989, let alone C++ in 1998.
This particular exercise wants a machine-word size integer, so in Rust that's isize, but you could isize::from_le_bytes() or whatever with chunked conversions, which will happen at runtime but ought to go very much faster than parsing text.
If there’s a good and insurmountable reason why it must be a minute and no less, okay. But I’d be amazed if there aren’t some easy wins to be made that could speed it up by a decent amount.
That’s not worth doing just because of the silly million-element array case, of course; but if you can make that silly case faster, lots of more important use cases will get faster too.
You can submit a PR and after a robot puts it in a pile to be looked at, actual humans will ask you about your proposed change, they can ask other robots to check whether it works OK on the huge piles of real world Rust out there, and so on.
``` list = [1,2,3,4,5,6 ..., 7999, 8000] ```
i.e. a written out literal, that the swift compiler has trouble type-checking if everything in there is an int.
Today I get the felling that Swift is a science project used by millions and backed by a trillion dollar company.
I don’t know how you’d escape this. Other platforms have their own problems. As far as I can tell, they’re all science projects, because our understanding of language design and library design keeps changing.
Seconded -- Obj-C was really nice to work with at that time.
I do like a lot of the new stuff in Swift, but in many ways it’s a return to the bad old days in terms of the overall dev experience.
It depends on the angle you’re coming from IMO.
For me Swift has been a major improvement overall for a couple reasons: it has a ton of quality of life features that can only be had in Objective-C with a laundry list of CocoaPods, and its type system allows me to do fairly major refactors on a regular basis that I wouldn’t dream of trying with Obj-C. Not having to maintain header files is also a bigger deal than I thought it’d be…
If we go by the Objective C timeline, Swift will be pretty good in the year 2042.
let x: Int = 123
let y: Double = 123
But it looks like here the problem was not the type checker anyway.In the end what I ended up doing was to import it as JSON or binary plist or something at runtime.
Not great. But at least I could compile my code again.
Another surprise is that Apple’s protobuf generator is able to produce iodiomatic Swift, the same is not true for Kotlin :)
- If there is no concrete type that you can infer for the given literal, throw a type error.
- If there is no concrete type that you can infer for the given literal, fall back to a default, e.g. Int.
But this sort of situation is why backtracking during type inference can lead to pathological behavior.