There are many cases of UB that would be cheap to check, but there are many more that are incredibly expensive to check.
183 karma · joined October 31, 2015
There are many cases of UB that would be cheap to check, but there are many more that are incredibly expensive to check.
There's nothing in LLVM itself that makes it use larger sizes, it just depends on the ABI and what's fastest when ABI doesn't matter.
C however states :
> When the processing of the abstract machine is interrupted by receipt of a signal, the values of objects that are neither lock-free atomic objects nor of type volatile sig_atomic_t are unspecified, [...] The representation of any object modified by the handler that is neither a lock-free atomic object nor of type volatile sig_atomic_t becomes indeterminate when the handler exits.
This wording is not present in C++, as it instead defines how signal handlers fit into the memory model.
This means that (with adjustments for C atomics):
int val = 0;
std::atomic<bool> flag{false};
void handler(int sig) {
if (!set) {
val = 1;
flag = true;
}
}
int main(void) {
signal(SIGINT, handler);
while (!flag) { /* Spin waiting for flag */ }
return val;
}
Is valid in C++, but not in C.In C++ there's actually a lot more freedom. You can access non-atomic non-volatile-std :: sig_atomic_t variables as long as you don't violate the data race rules.
However, this is a reasonable rare circumstance, and it's only getting easier to charge.
Also the power consumption makes this mostly useless for battery operation. Overall I'm still not impressed by the 2040. I'll stick with the NRF52840.
This is very obviously wrong. The economics of crypto mining are almost entirely dominated by power cost. Nvidia's decision here simply means they will use other hardware, not more of the LHR hardware.
Nvidia already can't make enough GPUs to satisfy demand. Why would they try to increase undesirable demand (which is going away soon anyway, at least for ETH).
ASLR has nothing to do with the type of memory layout discussed here. ASLR only impacts compiled code/data, and only entire shared objects/executables at a time.
A bad memory layout can have a huge impact on perf. A simple example is iteration order of a 2d array, where not doing sequential access can result in a ~5x slow down.
> Computer thermals when running the test can contribute 40%.
Only if you forgot to apply thermal paste.
Not quite true. For example quite a few orbits around the moon are unstable and you will end up crashing into the surface. This is due to essentially all celestial objects not being spherical.
Additionally, implementers are in completely agreement here that this works. There are zero standardization/implementation concerns with this method, and I would highly advise against scaring users away from it when necessary.
All code transformations have a compile time cost and runtime perf impact. We don't add transformations unless the runtime perf impact greatly outweighs the compile time cost.
These optimizations are added because they measurably improve the performance of real code. This comes up in every review for new or updated optimization passes. This claim that they aren't justified is actually rather insulting to the effort put in to improve perf without taking days to compile.
For this kind of behavior there is actually no limit on what the implementation is allowed to do, including assuming it never happens and optimizing accordingly. Implementations just need to document what they do.
I don't think these attempts to change the definition of UB are useful. As a compiler vendor it doesn't help to just say "it has a behavior", because that doesn't stop me from doing exactly what I do today. If people want some specific behavior, or to limit behaviors, then they need to actually say that in the spec.
To take left shift for example. Instead of saying it's implementation-defined behavior, they should say that it produces a implementation-defined non-trap value in the range of the resulting type that is consistent for the same inputs.
This would be pretty short to write in standardees, doesn't allow UB based optimizations, and allows the required implementation divergence.
X86-64 uses SSE registers for all floating point operations. I'm not sure that the author realized that they were looking at an -O0 binary. -O0 does not do vectorization (or anything else for that matter).