HNHacker News
TopNewBestAskShowJobs

electricshampo1

97 karma · joined November 3, 2020

submissionscomments
electricshampo1··on Faster Index I/O with NVMe SSDs
Depending on the IOPS rate for your app; SPDK can result in less CPU time spent in issuing IO/reaping completions compared to ex. io_uring.

See Ex. https://www.vldb.org/pvldb/vol16/p2090-haas.pdf What Modern NVMe Storage Can Do, And How To Exploit It: High-Performance I/O for High-Performance Storage Engines

for actual data on this.

OFC, If your block size is large enough and/or your design is batching enough etc. that you already don't spend much time in issuing IO/reaping completion then as you say, SPDK will not provide much of a gain.

electricshampo1··on H.R.1 Sec. 70302. Full expensing of domestic research, experimental expenditures
From the pdf of the bill, software is still classified as R&D on page 303 of this bill; the change is that domestic (US) software R&D is no longer forced into a 5yr amortization schedule.

Foreign software R&D is still forced to be amortized over a 15 year period; see pages 304,305,306 of the bill.

electricshampo1··on Post-Silicon Validation of Static Lockstep Mode
Nice to see SDC concerns being taken more seriously by hardware folks. Once software gets to sufficient quality (which we have achieved in many cases), these kinds of rando hw issues are the only remaining causes of "impossible" bugs that waste endless engineering time to debug.

I wonder how much of this relies on or is made easier by the clustered core architecture of E-Core Xeons. In comparison each physical core of P-Core Xeons is its own island basically.

electricshampo1··on Using high bandwidth PCIe 5 SSDs in databases
The gap between the efficiency displayed here and that which can be found in ex. postgres/mysql is insane.
electricshampo1··on Pushing AMD's Infinity Fabric to Its Limit
It is integer factors better overall total BW than ddr5 spr; I think they went for minimal investment + time to market for the spr w/ hbm product rather than heavy investment to hit full bw utilization. Which may have made sense for intel overall given business context etc
electricshampo1··on Better-performing “25519” elliptic-curve cryptography
Completely agree re: firedancer codebase. There is a level of thought and discipline wrt performance that I have never seen anywhere else.
electricshampo1··on Optimizing global message transit latency: a journey through TCP configuration
On prod servers I see a bunch of frontend stalls & code misses in the L2 for the kernel tcp stack; having each process statically embed its own network stack may make that worse (though using dynamic shared quic lib for ex. in userspace shared across multiple processes partially addresses that but with other tradeoffs).

Of course depending on usecase etc the benefit from first-order network behavior improvements is almost certainly more important than the second-order cache pollution effects of replicated/seperate network stacks.

electricshampo1··on Optimizing your programs for Arm platforms
Thanks for this link; did not realize that they did this.
electricshampo1··on Measuring CPU core-to-core latency
Answering only the latter question:

A Primer on Memory Consistency and Cache Coherence, Second Edition

https://www.morganclaypool.com/doi/10.2200/S00962ED2V01Y2019...

(free online book) would help

electricshampo1··on Things to know about databases
" Like many modern analytical engines [18, 20], Procella does not use the conventional BTree style secondary indexes, opting instead for light weight secondary structures such as zone maps, bitmaps, bloom filters, partition and sort keys [1]. The metadata server serves this information during query planning time. These secondary structures are collected partly from the file headers during file registration, by the registration server, and partly lazily at query evaluation time by the data server. Schemas, table to file mapping, stats, zone maps and other metadata are mostly stored in the metadata store (in Bigtable [11] and Spanner [12])."

https://storage.googleapis.com/pub-tools-public-publication-...

electricshampo1··on Graviton 3: First Impressions
The whole chip in general will be used in aggregate by independent vms/containers etc that do NOT read and write to the same memory. Some kernel datastructures within a given vm are still shared, ditto for within a single process, but good design minimizes that (per cpu/thread data structures, sharded locks, etc etc).
electricshampo1··on AlloyDB for PostgreSQL under the hood: Columnar engine
It might be helpful to read something like

DB2 with BLU Acceleration: So Much More than Just a Column Store

https://db.cs.pitt.edu/courses/cs3551/16-1/handouts/db2BLU.p...

or

Real-Time Analytical Processing with SQL Server

https://www.vldb.org/pvldb/vol8/p1740-Larson.pdf

to get a sense of the gap between postgres/open source databases vs something designed (at least partially) to better support analytical workloads.

electricshampo1··on Using Java's Project Loom to build more reliable distributed systems
Unlike goroutines, seems here you have control over the execution schedule for the virtual threads if you provide an executor. This is pretty great.

Think this will obsolete go over the next few decades.

electricshampo1··on Removing characters from strings faster with AVX-512
This is only on the client side; server still has and will have AVX512 for the foreseeable future.
electricshampo1··on Choosing the Right Integers
"However, we were talking about array indexes, for loops, and file offsets for a single file. These are 8 byte variables within a running program."

This depends on storage/page layout etc. See for ex. https://db.cs.pitt.edu/courses/cs3551/16-1/handouts/db2BLU.p...

electricshampo1··on S2n-QUIC (Rust implementation of QUIC)
Is Swift (https://dl.acm.org/doi/pdf/10.1145/3387514.3406591) expected/designed to be used in non-intra dc environments where primarily quic is expected to have an advantage relative to tcp?

I agree that it still would be nice to have as an option.

electricshampo1··on Are you sure you want to use MMAP in your database management system? [pdf]
Generally for perf critical use cases you dedicate the machine to the database. This simplifies many things (avoiding having to reason about sharing, etc etc).
electricshampo1··on Ramp up your distributed transactions
Seems like facebook's internal db TAO supports it according to https://www.vldb.org/pvldb/vol14/p3014-cheng.pdf
electricshampo1··on AWS Identity service handles 400M API calls every second
Seems 400M is aggregated qps worldwide. Wonder what avg qps looks like per iam server (and size of server).
electricshampo1··on Paxos vs. Raft: Have we reached consensus on distributed consensus?
This is just best effort on google's end right? Don't think anything is documented/guaranteed such that you would be able to, for ex. rely on it like spanner's use of true time.
electricshampo1··on It’s Not Always iCache
Very often people are looking at icache misses instead of something more precise when regarding perf effects due to code size/layout, etc. That more precise thing is frontend stalls: you only care about misses when they cause stalls; otherwise they are overlapped with actual work being done by the execution units.

You can measure frontend stalls on many recent intel chips by

IDQ_UOPS_NOT_DELIVERED.CORE

https://perfmon-events.intel.com/

Neoverse N1 from Arm has STALL_FRONTEND; see

https://developer.arm.com/documentation/PJDOC-466751330-5476...

electricshampo1··on Programming Language Memory Models
"Java and JavaScript have avoided introducing weak (acquire/release) synchronizing atomics, which seem tailored for x86."

This is not true for Java; see

http://gee.cs.oswego.edu/dl/html/j9mm.html

https://docs.oracle.com/en/java/javase/16/docs/api/java.base...

electricshampo1··on Intel and AMD Contemplate Different Replacements for x86 Interrupt Handling
This is essentially the approach taken by

https://www.dpdk.org/ (network) and https://spdk.io/ (storage)

Anything trying to squeeze perf doing IO intensive work should switch to this model (context permitting of course).

electricshampo1··on Cores that don’t count [pdf]
Thanks for the reference.
electricshampo1··on “Is Parallel Programming Hard, and, If So, What Can You Do About It?” v2 Is Out
For this load shedding in a single node, are you imagining something more than work stealing style approaches?

In a multi-node cooperative setting you need some way to transmit information that a given node is overloaded, some way to find nodes that have available capacity, and a low overhead way to shift the work over to them. If the work to be done depends solely on data that you have local to you, it seems silly to shift the data as well (depending on how big it is); this would only make sense when you have to combine data from a variety of nodes (which can be done on any other node, assuming that the current node is overloaded).

Probably not worth doing anything about short lived hotspots (<1s). I wonder what kind of granularity you have used in your systems (probably different for within node vs across nodes).

electricshampo1··on Speed of Rust vs. C
I don't know about latest; but try look at

Scheduling Parallel Programs by Work Stealing with Private Deques https://hal.inria.fr/file/index/docid/863028/filename/full.p...

electricshampo1··on On Summing Integers
(author here)

Using srand(0) as in your gist; changing int8_t to char ; using int for sum:

BM_hn 10155300 ns

BM_avx512 9218652 ns

-----

BM_avx512 9372319 ns

BM_hn 10428792 ns

Your code is getting ~18-19GB/s read bandwidth compared to roughly ~21GB/s for my code.

I wonder how much of that is due to interleaved summing vs not compared to AVX512 vs AVX2.

electricshampo1··on On Summing Integers
(Author here)

The main focus of the article is to show a worked example of how to analyze the behavior of simple programs. The circuitous path taken in the article is similar to that which you might face when analyzing your own programs. The article is only incidentally about AVX512 and summing numbers. Thus any specific technique used or even the precise runtimes measured should not be given too much weight. Forgive me, but the journey is more important than the destination.

The article is also meant to inspire others to learn and spend time deeply understanding (by looking at data that the hardware makes available) what actually happens to their code when it runs. It is too easy to lose sight of that in a professional setting, where business needs/requirements leave little time for deep analysis.

electricshampo1··on On Summing Integers
NOTE: The article has been updated and expanded since the initial post.
electricshampo1··on On Summing Integers
That is essentially the approach mentioned in the article at

" UPDATE: see https://www.realworldtech.com/forum/?threadid=200693&curpost... for a dramatic simplification. Not catching this is an oversight on my part. This post will be updated to include numbers with the mentioned strategy.

UPDATE: To my surprise and after much fiddling, I did not manage to write a version that was measurably faster (indeed they were at least a percent slower) than the hand written sum_avx512 shown below. There is almost certainly something that I am doing wrong but I can’t seem to figure out what it is. I will take this opportunity to leave this as an exercise for the reader :). "

Page 1 of 2Next →