HNHacker News
TopNewBestAskShowJobs

BeeOnRope

2,629 karma · joined August 7, 2017

submissionscomments
BeeOnRope··on Intel's Redwood Cove: Baby Steps Are Still Steps
Yes I don't really get it. For non-loop branches if the compiler has knows enough to insert a branch hint it would also already arrange the generated code so that was the fall through cases, which is already much more efficient when predicted and aligns the default fall through assumption in the CPU.
BeeOnRope··on July 2024 Update on Instability Reports on Intel Core 13th/14th Gen Desktop CPUs
What's special about the E core's L2 cache such that it gets on-chip regulated voltage?
BeeOnRope··on Things I learned while writing an x86 emulator (2023)
Sounds a bit like the jcc erratum?
BeeOnRope··on Beating the Compiler
Modern compilers are not doing much searching in general. It's mostly apply some feed-forward heuristic to determine whether to apply a transformation or not.

I think a slower, search based compiler could have a lot of potential for the hottest parts you're willing to spend exorbitant time on a search.

BeeOnRope··on Things I learned while writing an x86 emulator (2023)
I'm not following: as long as you are introducing a new, incompatible instruction for leading zero counting, you'd definitely choose LZCNT over BSR as LZCNT has definitely won in retrospect over BSR as the primitive for this use case. BSR is just a historical anomaly which has a zero-input problem for no benefit.

What would be the point of offering a new variation BSR with different input semantics?

BeeOnRope··on Things I learned while writing an x86 emulator (2023)
Probably, if the uops come from the uop cache you get the fast speed since the prefix and any decoding stalls don't have any impact in that case (that mess is effectively erased in the uop cache), but if it needs to be decoded you get a stall due to the length changing prefix.

Whether a bit of code comes from the uop cache is highly dependent on alignment, surrounding instructions, the specific microarchitecture, microcode version and even more esototeric things like how many incoming jumps target the nearby region of code (and in which order they were observed by the cache).

BeeOnRope··on Show HN: Dut – a fast Linux disk usage calculator
Yes, files can be sparse but the actual disk usage information is also returned by these stat-family calls, so there is no special cost to handling sparse files.
BeeOnRope··on Show HN: Dut – a fast Linux disk usage calculator
This seems difficult since I'm not aware of any way to get approximate file sizes, at least with the usual FS-agnostic system calls: to get any size info you are pretty much calling something in the `stat` family and at that point you have the exact size.
BeeOnRope··on Detecting a PS2 Emulator: When 1*X does not equal X
The rules and mechanisms for SMC detection are essentially the same in both modes as far as I am aware.

Both Intel and AMD implement SMC detection that is a bit stronger than required by the specification as well.

BeeOnRope··on AMD Unveils Ryzen 9000 CPUs for Desktop, Zen 5
> of the high-performance programs that is possible in 512-bit AVX-512 due to the equality between register size and cache line size, so the consumer Intel CPUs will remain a worse target for the implementation of high-performance algorithms.

Can you elaborate here? I love full-width AVX-512 as much as the next SIMD nerd, but I rarely considered the alignment of the cache line and vector width one of the particularly useful features. If anything, it was a sign that AVX-512 was probably the end of the road for full-throughput full-width loads and stores at full AVX register width, since double-cache line memory operations are likely to be half-throughput at best and a doubling of the cache line width seems unlikely.

BeeOnRope··on Transforming a QLC SSD into an SLC SSD
It seems like that would be a likely common behavior for the FTL, but other options are possible (e.g., reading the old blocks) and it wasn't guaranteed by the spec, which is why they added this NVMe flag (so-called "DZAT") so that you can actually rely on it.
BeeOnRope··on Transforming a QLC SSD into an SLC SSD
Not guaranteed by default for NVMe drives. There's an NVMe feature bit for "Read Zero After TRIM" which if set for a drive guarantees this behavior but many drives of interest (2024) do not set this.
BeeOnRope··on Transforming a QLC SSD into an SLC SSD
Yes.
BeeOnRope··on An informal comparison of the three major implementations of std:string
libc++ string is smaller with a higher SSO capacity which in many scenarios can overwhelm the code generation. So it's hard to draw an absolute conclusions.
BeeOnRope··on So you think you want to write a deterministic hypervisor?
That's not how it works on x86 as far as I know. The atomic ops are simply performed by the ALU, against the L1 cache when the instruction is about to retire. Atomicity is guaranteed by not allowing the line to be stolen by another core between the operation, and memory order is ensured by draining the store buffer before executing the op and other speculation (eg loads can speculatively pass even atomic operations).
BeeOnRope··on The return of the frame pointers
> Doesn't explain why there's no -fshadow-stack-only option to pass in.

I thought you were asking about the design of the hardware: it's designed that way because compatibility means that the vast majority of people want something backwards compatible.

Very little is fully "compiled from scratch" when you consider for example that libc is in the set of things that must be recompiled.

BeeOnRope··on The return of the frame pointers
ABI compatibility.
BeeOnRope··on The return of the frame pointers
Which is the slow unwinding path? The one from libbfd?
BeeOnRope··on GhostRace: Exploiting and mitigating speculative race conditions
This kind of memory order speculatiom is basically required on x86 since the strong semantics would otherwise prevent many useful reorderings (especially L-L which is absolutely critical).

The basic way it works is that pretty much any reordered is allowed, speculatively, even if it violates the explicit or implicit barriers, but then the operations are tracked in a memory order buffer (MOB) until retirement which can detect whether the reordering was detectable and if so flush the pipeline (so-called "memory order nuke", visible in performance counters).

BeeOnRope··on Cloning a Laptop over NVMe TCP
In EC2, most of the "storage optimized" instances (which have the largest/fastest SSDs) generally have more advertised network throughput than SSD throughput, by a factor usually in the range of 1 to 2 (though it depends on exactly how you count it, e.g., how you normalize for the full-duplex nature of network speeds and same for SSD).
BeeOnRope··on Downfall Attacks
> But that doesn't prevent issues due to switching between customers on the same physical core, no?

Yes they are explicit that customers may be time-shared on a physical core ("burstable" instances don't really make sense without that). Most of these attacks aren't known to be possible in that scenario and in any case the mitigations are much easier since flushing sensitive state at group scheduling boundaries is much less costly than permanent dynamic changes to how concurrent SMT threads interact.

BeeOnRope··on Downfall Attacks
Granted but the vendors accept these are serious problems given that they are immediately patched and the mitigation all enabled by default even at significant performance cost (most chip generations are down double digit perf % based versus "zero mitigations" at this int).

So they don't need to build the "perfect processor" but why aren't they discovering any of these issues themselves?

BeeOnRope··on Downfall Attacks
Shouldn't the functions of "development" and "QA" both reside under the umbrella of the chip-maker though?

In fact, chip-makers famously invest an insane amount of money into "QA" (aka "validation") and many features or lack thereof are often put down to the cost of QA rather than the cost of development.

BeeOnRope··on Downfall Attacks
That's true, but it leads the odd assumption that the vendor managed to fix N side-channel attacks before release but 0 thereafter, while random individuals fixed M thereafter over a period of years with N >> M.

This seems to be much less likely than the conclusion that vendors are not in fact fixing many prior to the release and then stopping "cold turkey" after that. Especially since these attacks seem to cross chip versions, in many cases 6+ generations of chips: if vendors had substantial and increasing efforts on new chip versions they'd also be catching issues that applied to old released chips as well. We don't see that happening.

BeeOnRope··on Downfall Attacks
I dabble in this space (hardware reverse-engineering) and write software for a living and in my opinion the gaps are huge.

I should disclose have been paid by a chip-maker for a blog post that I wrote which "disclosed" an optimization which could be uses for a side channel attack (though I did not even suggest that aspect) and which was subsequently patched away via a microcode update. The whole process was very surprising to me in that there must have been several people inside the chip-maker who knew about the optimization I described in much deeper detail ... after all they conceived and implemented it.

So by what path does a blog post mentioning it get treated as the disclosure that results it it being removed when they knew about it all along?

> that's not how these attacks work it's by some oversight in memory handling not that different from software.

I think it is very different. Assembly is merely a somewhat less convenient form of the original semantics that embeds all the relevant semantics related to the attack surface since the original source has been "erased". Many analysis tools such as fuzzers operate directly on assembly with little loss in functionality.

These attacks are against completely unspecified aspects of the instruction execution and lean heavily on the actual hardware implementation (almost at the level of "how the transistors are laid out") such as what hidden buffers are used, when they are filled, how they are shared with sibling threads, etc.

In my experience there are very few people interested in these details outside of the vendors themselves and these folks and the ones creating the exploits would fit in a modestly sized lecture hall. The scope has increased a bit lately (see Tavis's fuzzer work) but it was originally a small group with little or no funding.

BeeOnRope··on Downfall Attacks
Nevermind, AWS explicitly documents that all instance types, including burstable, never co-locate different tenants on the same physical core at the same time:

https://docs.aws.amazon.com/whitepapers/latest/security-desi...

BeeOnRope··on Downfall Attacks
I'm not directly in charge of hiring QA, no!

I think this sort of excuses the initial blindness to Spectre style attacks in the first place, but once the basic pattern was clear it doesn't excuse the subsequent lack of discovering any of the subsequent issues.

It is as if someone found a bug by examining the assembly language of your process which was caused by unsafe inlined `strcpy` calls (though they could not see the source, so they had no idea strcpy was the problem), and then over a the subsequent 6 years other people slowly found more strcpy issues in your application using brute-force black-box engineering and meanwhile you never just grepped your (closed) source for strcpy uses and audited them or used many of the compiler or runtime mitigations against this issue.

BeeOnRope··on Downfall Attacks
> A fundamental problem is that the attack surface is so, so huge. Even if their security researchers are doing blue-sky research on both very small and very broad areas of processor functionality, they're going to miss a lot.

Sure. If they had patched a bunch of Spectre vulnerabilities and independent researchers had discovered a few more that would be one thing, but as far as I can tell they have patched _zero_ while independent researches have found many and it has been years since the initial attack. Many of these follow very similar patterns and "in what cases is protected data exposed via speculative execution" is something that an architect or engineer could definitely assess.

BeeOnRope··on Downfall Attacks
I think the comparison between CPU and software exploits holds at a very high level, but in the case of software the gap between internal and external researches seems lower. Much software is open-source, in which case the play field is almost level and even closed source software is available in assembly which exposes the entire attack surface in a reasonably consumable form.

Software reverse-engineering is a hugely popular, fairly accessible field with good tools. Hardware not so much.

> When billions use something I expect them to find more problems, flaws, and exploits in it than the creator/manufacturer did. The presence of this does nothing to indicate (or refute) any further conclusion about why.

To be very clear none of these errors have been found by billions of random users but by a few interested third parties: many of them working as students with microscopic funding levels and no apparent inside information.

I'm not actually suggesting that the nefarious explanation holds: I'm genuinely curious.

BeeOnRope··on Downfall Attacks
> Also you’re dealing with a company that has been running to stand still for a long time

I'm not just talking about Intel, but also Arm and AMD. As far as I know none of these has obviously been making proactive Spectre fixes.

← PreviousPage 3 of 30Next →