368 karma · joined March 23, 2024
Whether or not you have to write the "case" keyword 10 times is an aesthetic choice.
I don't think this has anything to do with program optimization. On all non-theoretical ISA's I'm aware of, you don't need a JUMP instruction to go to the next instruction. We're debating names.
I'm a Ziguana so my answer to the programming philosophy questions would be that we need a language where the complexity emerges in the code, not in the language itself, and we generally want a shared language that can be read and used by anyone, anywhere. If everyone has their own subset of the language (like C++) then it's not really just one language in practice. If every project contains its own domain specific language, it may be harder for others to read because they have to learn custom languages. That's not to say you should never roll your own domain specific language, or that you should never write a program that generates textual source code, but the vast, vast majority of use cases shouldn't require that.
And, yes, be opinionated. I'm fine with some syntactic sugar that makes common or difficult things have shortcuts to make them easier, but again, if I learned a language, I should generally be able to go read someone's code in that language.
What do you consider "advancing the field as a whole"?
Are there people who see a "match" statement, smash both hands on the table, and shout, "WHAT THE ___ is a ------- MATCH STATEMENT?!!! THIS IS SO $%^&@*#%& CONFUSING!! I DON'T KNOW THAT WORD!! I ONLY KNOW SWITCH!!"
2020: I don't care how it performs
2030: I don't care why it performs
2040: I don't care what it performs
You can go read about the real differences on sites like Chips and Cheese, but those aren't pop-sciencey and fun! It's mostly boring engineering details like the size of reorder buffers and the TSMC process node and it takes more than 5 minutes to learn. You can't just pick it up one day like a children's story with a clear conclusion and moral of the story. Just stop. If I can acquire all of your CPU microarchitecture knowledge from a Linus Tech tips video, you shouldn't have an opinion on it.
If you look at the finished product and you prefer the M series, that's great. But that doesn't mean you understand why it's different from the Zen series.
I tend to disagree with most of these rulesets that are meaningless to "engineering". The idea that a function should only be 40 lines long is offensive to me. Personally, I would rather have one 400 line function than ten 40 line functions. I'm a Ziguana. I care about handling edge cases and I think my programming language should be a domain specific language to produce optimal assembly.
I would not constrain other people who feel differently. I read an article where some project transitioned from Rust to Zig, even though the people on the team were all Rustaceans. Obviously their Rust people hated this and left! To me, that's not a step in the right direction just because I prefer Zig to Rust! That's a disaster because you're taking away the way your team wants to build software.
I think hardly any of the things we disagree on actually have much to do with "Engineering". We mostly aren't proving our code correct, nor defining all the bounds in which it should work. I personally tend to think in those terms and certain self-contained pieces of my software have these limits documented, but I'm not using tools that do this automatically yet. I'd love to build such tools in the coming years though. But there's always the problem that people build tools that don't notice common use-cases that are correct, and then people have to stop doing correct things that the tool can't understand.
Daniel Spill, who worked with Parkes directly in England, founded several companies with Parkes selling Celluloid in England.
Spill and Hyatt spent the better part of a decade in court against each other over who invented it first and who has the right to the patents. The judge ultimately ruled that both of them can continue their businesses, and that Parkes invented it first.
He's "polishing a pig"? He's hiding "all issues" with Zig in internal mailing lists to "project a polished facade"? ALL?! You got all that from one sentence saying he wishes the author took a different approach?
Alright, fine. Here's my analysis of your character and lifelong patterns of behavior based on your first two sentences:
You just want to tear down everybody who is trying to do good work if they make any mistake at all. You look for any imperfection in others because criticizing people is the only approximation of joy in your existence. You are the guy that leaves Google reviews of local restaurants where you just critique the attractiveness of the women who work there. You see yourself as totally justified and blameless for your anti-social behavior no matter the circumstances, and you actually relish the idea of someone being hurt by you because that's all the impact you could hope for.
If that's not accurate to who you are, well, ¯\_(ツ)_/¯ that's just how it reads to me.
You said, "anybody [...] can write whatever they want". "They are under zero obligation ...". "Members of the Zig project are free to reach out ..."
Do you not realize that you have not at all addressed the point about what is the most productive way to criticize?
All you have done is go off about people's rights, freedoms, and lack of obligations. But nobody actually said "People shouldn't be able to post critical blog posts" or "People are obligated to participate by filing issues or contributing code to open source". So what was the point in saying this? Do you think people believe anything contrary to what you said?
He only said he wishes the author would have taken a different approach. So what? Why does everyone have to jump in and start psychologizing or essentializing Andrew based on one paragraph?
Why does one paragraph have to say so much about who he is as a person? Even if it did piss him off for a few hours, so what? He's not allowed to wish someone took a different approach?
I tend to think Andrew Kelley is a great guy, not just technically but as a person. And I think that because I've listened to him talk for dozens of hours. I can guarantee you that that one sentence he wrote is not the beginning of a character assassination campaign against the author of this blog.
He made Zig because he wanted to put something good into the world and improve the state of software. How about we include that in our analysis of Andrew's character? I'll leave it to the reader to consider whether the multi-year full time dedication to Zig should be weighed more heavily than a personal feeling he had for two minutes that he expressed respectfully without attacking anyone's character.
The bigger issue is instruction ordering and register allocation. On code where the compiler effectively has to lower serially-dependent small snippets independently, I think the compiler does a great job. However, when it comes to massive amounts of open code I'm shocked at how silly the decisions are that the compiler makes. I see super trivial optimizations available at a glance. Things like spilling x and y to memory, just so it can read them both in to do an AND, and spill it again. Constant re-use is unfortunately super easy to break: Often just changing the type in the IR makes it look different to the compiler. It also seems unable to merge partially poisoned (undefined) constants with other constants that are the same in all the defined portions. Even when you write the code in such a way where you use the same constant twice to get around the issue, it will give you two separate constants instead.
I hope we can fix these sorts of things in compilers. This is just my experience. Let me know if I left anything out.
"Not today, CIA" "Nice try, FBI" "Nice attempt at obscurity, Department of Homeland Security"
And regulating Apple is quite different from regulating someperson. If you made a Linux phone in your basement, nobody would tell you what kind of charger you should use. But companies that claim ownership of a substantial economy and can dictate the rights and culture and economic output for a large section of society do need to have more checks on their power than just, "Well if I'm so wrong, then why do I have so much money? Maybe you should make your own phone that won't work well with anyone else's and see if you can sell it."
I would like to see more effort dedicated to basic one liners that show up in real code like counting how many of a given character are in a string. E.g. `for (str) |e| count += e == '%'`. For this, LLVM spits out a loop that wants to do horizontal addition every iteration on x86-64 targets with vectors, at least. Let's focus on issues that can easily net a 2x performance gain before going after that 1-2% that people think pointer aliasing gets you.
24 grams. You won't even feel the weight of it.
I'm glad to know that's the design of Cranelift, since that's how I would think it should be done, although I haven't written a middle or backend for a compiler yet.
E.g. I wrote a `pextComptime` function, which will compile to just a `pext` instruction on machines that have a fast implementation, otherwise it will try to figure out if it can use a few clever tricks to emit just a couple of instructions, but if those aren't applicable it will fallback on a naïve technique.
https://github.com/Validark/Accelerated-Zig-Parser/blob/8782...
I'm not so sure. How many cycles would you expect this code to take?
mov dword [rsi], eax
add dword [rsi], 5
mov ebx, dword [rsi]
According to Agner Fog, these 3 instructions have a latency of 15 cycles on an AMD Zen 1 processor. On Zen 2, its latency is 2 cycles. This is because the CPU was given the ability to assign a register to `dword [rsi]`, overcoming the limit of 16 registers.This optimization is subject to problems, obviously pointer aliasing will enable the CPU to make the wrong assumption at times, and cause a situation not entirely unlike a branch mispredict.
There are constraints imposed by the micro-architecture for this feature. For you and I, a big one is it only works with general purpose registers. But is there a reason it couldn't or shouldn't be done for vectors? It seems like a micro-arch issue to me. Perhaps in a few years or in a lot of years, we'll have a CPU that can do this optimization for vectors.
That video is very interesting to me. I don't yet know enough about all those languages to speak intelligently on how they could be used to simplify my code, but I do agree it would be amazing if we could generalize my techniques so that every project could easily achieve what I have through a lot of hard work.
I'm not sure if the specific thing demo'ed in that video is helpful for my use-case. That workload, as well as other examples on the site, seem(s) to get a ton of mileage because their data is two-dimensional. Therefore there are a lot of different choices that could be made about the order in which things should be done. Being able to switch those orders quickly to test out how well vectorization works for different choices is therefore a major benefit to those examples.
The tokenizer, on the other hand, is one-dimensional. It's obvious what's the best way to traverse a source file: Just go forward.
Automatic scaling to bigger or smaller vectors could be easier though, but how I want that to be achieved atm is for LLVM to have:
1. u512 operations happening in AVX-512 vectors
2. ARM backend support for automatically preferring interleaved vectors over regularly-ordered vectors. Maybe even better would be a type at the language level.
3. A SWAR emitter for LLVM vectors for machines without vectors
4. Movemask support in the PowerPC, WASM, and MIPS backends. I don't want to write intrinsics for every little bitcast.
I'm also just skeptical in general that offloading to the compiler is a good idea with the current state of the technology. There are just too many routines it doesn't know about, it misses canonicalization too much, it doesn't seem to be architected correctly for having several register files, it doesn't seem to be architected correctly for supporting bit manipulation efficiently, and the fact that SWAR support hasn't arrived yet at this point in history doesn't inspire confidence either.
In cases like the video shows where you're just coercing the compiler to do basic auto-vectorization, I'm sure that works perfectly.
But in this work, my hunch is that the optimal solution could not be trivially written in a nice functional style. This code does not use the "easy solution" of being able to reach forwards and backwards for data. I.e. The easy solution expressed as some pattern matching over `{ file[p-1], file[p], file[p+1], ... }` (I believe) imposes a constraint that we have to have read in the next vector and produce the bitstrings for it in order to process the current vector. For simple languages this is fine but we have to do a lot of work for something like Zig and you just don't have enough registers to go around. That's why my solution uses pseudo-starts (putting a fake start position at the beginning of a chunk due to carried-over values) and a more manual carry-over system. The logic here is NOT position-independent, solely for performance reasons. Do I think the necessary transformation between position-agnostic logic and the logic I'm talking about could be converted automatically? Maybe. But my suspicion is that a parser generator makes more sense as the sort of tool that should solve this.