641 karma · joined February 15, 2013
It's not strictly true to say that it's "being conservative". What is more correct is to say that floating point operations have different semantics to integer operations, and an optimisation that retains the semantics of an expression over integers may not do so when applied to an expression over integers. Hence, it may be possible to apply one optimisation to an integer expression, but applying that to a floating-point expression may result in a different program meaning.
C/C++ compilers give you a way out of this with the `--ffast-math` flag, which essentially allows compilers to relax the constraints on floating-point optimisation passes.
For an example of how this works in GCC, take a look here: https://gcc.gnu.org/wiki/FloatingPointMath
If you're solely interested in single-core performance, then I would agree that they are a good stress test, but I think for a processor that is being sold on it's parallelism, they are not a great benchmark.
I wish this was the case. Unfortunately, companies that work on non-Chromium browsers need to employ dedicated web compatibility teams to either a) help website users fix non-standard (i.e. Chrome only) HTML/CSS/JS, or b) replicate Chromium-like behaviour for specific (very popular) websites so that they work "correctly".
There's also the websites that deliberately block certain browsers which is what tools like "chrome-mask"[1] are built to solve.
[1] https://addons.mozilla.org/en-GB/firefox/addon/chrome-mask/
I accept that if you're extremely used to writing this style of C code, it might be something that you're used to, and understand implicitly. As a C++ engineer that infrequently comes across this precise pattern, having the precedence made explicit makes it much easier to understand.
block
statement (assignment)
expression
operator (dereference)
operator (post-increment)
variable
expression
variable *lwr = x; lwr++;
This could be be represented with something like this (and this is a very vague approximation of an AST): block
statement (assignment)
expression
operator (dereference)
variable
expression
variable
statement
expression
operator (post-increment)
variable
The second form, that looks this in the source: *lwr++ = x
might look like this in the AST: block
statement (assignment)
expression
operator (post-increment)
operator (dereference)
variable
expression
variable
If these two ASTs are to take the same form in the compilation pipeline, there needs to be some kind of pass that transforms one into the other.As for why the second form is faster (or rather, why the compiler can generate faster code): There is likely an optimisation pass somewhere in llvm that recognises a pattern that this fits into, which allows it to generate branchless instructions. For instance, in the second form, there is a pattern of:
operator (post-increment)
operator (dereference)
that might be recognised by a pass. In the former form, the two operators are far apart in the tree, so a pass would have to "look further" to match them up. A single pass likely won't do this, either for (compiler) performance reasons, or for correctness reasons.Finding which pass that is can be non-trivial, as it's more than a matter of enabling individual passes until one works. It might be that an earlier pass does some code reshaping that allows the relevant pass to work. My suggestion would be to dump the llvm ir at the end, and find the rough pattern that you're looking for, then re-run the compilation with `-mllvm -print-after-all` to see what the IR looks like after each pass, and then manually "look back" until you can't see the pattern any more.
cmp ecx, 1
je .LBB0_3
vs cmp ecx, 2
jne .LBB0_2
LBB0_3 and LBBO_2 were the same in both outputs (up to alpha renaming).Oddly, both sources seemed to be quite sensitive to match switch and enum reordering, resulting in very different generated code. Possibly something to look into further.
Pre-Christian religions had many associations with yew trees (they live for a long time, give off mildly hallucinogenic gasses on hot days, discourage animals), and so built their holy sites around them. When Christianity came to Britain, churches were deliberately built on pagan holy sites to overrun the old religions, in the same way that early Christianity took over roman holy days (Saturnalia -> Christmas, Lemuria -> All Saint's Day). This led to churches being built next to sites with copious yew trees.
This is most obvious when network effects are present (e.g. local immunisation efforts vs country-wide immunisation), but it's surprisingly common in other government-related areas like welfare, childcare, social security etc.
Edit: Another comment has reminded me that affordable public transport is the perfect example of this: Incrementally building out a public transport system will almost always fail, as the initial lines (be they buses, light rail, etc) will typically not be successful enough to justify the cost of building the line. If, instead, a system is built out universally and simultaneously, the utility (and thus income) of each line increases due to the interconnected nature of the network.
On the other hand, declaring the options through composition means that the API for "plot" remains static, and adding/removing options can be done trivially without an API change.
Composition (rather than parameters) is also more flexible. Let's say you want to divide your plot into three sub-plots, two of which are 200x200, and another which is 200x400. How do you express this as a keyword parameter? In composition, you could do something like:
plot( ggsubplot(ggvsplit(ggsize(200,400), gghsplit(ggsize(200,200), ggsize(200,200)))) )
At what point do we declare that a company has "grown" and now must make money? OpenAI is a multi-billion dollar company right now, surely that's a point at which they should be profitable, instead of propped up by further investment and borrowing.
> We have very strong indicators that inference is not a money loser for these companies
All of the economic analysis that I've read strongly states the opposite. Running a GPU is a net loss /even for the data centre operators/. For them to break even, they currently charge OpenAI/Anthropic/Etc more than OpenAI/Anthropic/Etc make per-token.
Also, don't forget that WASM is designed to replace JavaScript, thus it must interoperate with it to smooth the transition. Rosetta and Prism also work to smooth the transition from x86 -> ARM, and much of the difficult work that they do actually involves translating between the calling conventions of the different architectures, and making them work across binaries compiled both for and not for ARM, not with the bytecode translation. WebAssembly is designed to not have that limitation: it's much more closely aligned to JS. That's why it wouldn't make sense to use a subset of x86 or similar, as it would simply produce more work trying to get it to interface with JavaScript.
For everyone here complaining about this, have you ever looked at how many ways there are to access your history on Firefox? At my last count, there were 4 different ways to do it, depending on which menu you picked first. Cutting down this kind of inconsistent, and repetitive flow is something that we should be applauding.
Also, from a quick look at your profile, you seem to have quite a lot of comments criticizing or commenting on CodePlay. Do you have some sort of relationship or animosity with them?
For example, I worked as part of the team that managed the software that allowed pilots to submit flight plans. Any upgrades had to go through multiple weeks of reviews and testing (I don't mean code review - I mean reviews through managers and processes), and was run on some rather ancient hardware. Moreover, thanks to pressure from the pilots union, the system had to be able to accept flight plans by fax, so had a lot of legacy cruft to support that too.
The problem isn't agility, corner cutting or moving too fast - it's moving too slow.