HNHacker News
TopNewBestAskShowJobs

matt_d

22,061 karma · joined April 21, 2014

submissionscomments
matt_d··on Disambiguating Arm, Arm ARM, ARMv9, ARM9, ARM64, AArch64, A64, A78, ...
AArch64 SoC features:

https://marcin.juszkiewicz.com.pl/download/tables/arm-socs.h...

https://github.com/hrw/arm-socs-table

matt_d··on Myths and Legends in High-Performance Computing
- Myth 1: Quantum Computing Will Take Over HPC!

- Myth 2: Everything Will Be Deep Learning!

- Myth 3: Extreme Specialization as Seen in Smartphones Will Push Supercomputers Beyond Moore’s Law!

- Myth 4: Everything Will Run on Some Accelerator!

- Myth 5: Reconfigurable Hardware Will Give You 100X Speedup!

- Myth 6: We Will Soon Run at Zettascale!

- Myth 7: Next-Generation Systems Need More Memory per Core!

- Myth 8: Everything Will Be Disaggregated!

- Myth 9: Applications Continue to Improve, Even on Stagnating Hardware!

- Myth 10: Fortran Is Dead, Long Live the DSL!

- Myth 11: HPC Will Pivot to Low or Mixed Precision!

- Myth 12: All HPC Will Be Subsumed by the Clouds!

matt_d··on Still Entombed: The continuing twists and turns of a maze game [pdf]
HTML version: https://intarch.ac.uk/journal/issue59/3/full-text.html
matt_d··on Compiling Swift Generics [pdf]
FWIW, https://forums.swift.org/t/a-possible-vision-for-macros-in-s...
matt_d··on Are you sure you want to use MMAP in your database management system? [pdf]
One example is DBOS: A Database-oriented Operating System, https://dbos-project.github.io/ / https://github.com/DBOS-project (more details under "Publications").
matt_d··on How to learn compilers: LLVM Edition
More program analysis & LLVM resources (books, courses, and talks): https://gist.github.com/MattPD/00573ee14bf85ccac6bed3c0678dd...
matt_d··on NOELLE Offers Empowering LLVM Extensions
- https://github.com/scampanoni/noelle

- https://liberty.princeton.edu/Projects/NOELLE/

matt_d··on How does Clang 2.7 hold up in 2021?
I can definitely recommend https://book.easyperf.net/perf_book

The author's blog has been consistently great throughout the years, https://easyperf.net/notes/

See also microarchitectural performance analysis tools & readings, https://github.com/MattPD/cpplinks/blob/master/performance.t... and "Comments on timing short code sections on Intel processors", http://sites.utexas.edu/jdm4372/2018/07/23/comments-on-timin...

matt_d··on Advanced Compilers: Self-Guided Online Course
Compilers books: I'd start with "Engineering a Compiler" by Keith Cooper and Linda Torczon. http://craftinginterpreters.com/ is also a pretty great, programming-oriented intro, which may be good to work through alongside. For more on the analysis & compiler optimization side, "SSA-based Compiler Design" (http://ssabook.gforge.inria.fr/latest/; GitHub Mirror: https://github.com/pfalcon/ssabook) is a good follow-up.

Further readings: Book recommendations in https://github.com/MattPD/cpplinks/blob/master/compilers.md#... as well as program analysis resources (in particular lattice theory, type systems and programming languages theory, related notation): https://gist.github.com/MattPD/00573ee14bf85ccac6bed3c0678dd...

Courses: I can recommend the following: https://github.com/MattPD/cpplinks/blob/master/compilers.md#...

Particularly (in alphabetical order--I think these are all great, so including highlights of what I've liked about them):

- IU P423/P523: Compilers (Programming Language Implementation) - Jeremy Siek, with the course book "Essentials of Compilation: An Incremental Approach" (pretty interesting approach, with programming language features developed incrementally having a fully working compiler at each step, cf. http://scheme2006.cs.uchicago.edu/11-ghuloum.pdf; implementation language Racket),

- KAIST CS420: Compiler Design - Jeehoon Kang (good modern treatment of SSA representation itself, including the use of block arguments, https://mlir.llvm.org/docs/Rationale/Rationale/#block-argume..., as well as SSA-based analysis and optimization; Rust as an implementation language),

- UCSD CSE 131: Compiler Construction - Joseph Gibbs Politz, Ranjit Jhala (great lecturers, both Haskell and OCaml edition were interesting; fun extra: one of the Fall 2019 lectures (11/26) has an interesting discussion of the trade-offs between traditional OOP and FP compiler implementation),

- UCSD CSE 231: Advanced Compiler Design - Sorin Lerner (after UCSD CSE 131: for more on analysis & optimization--data flow analysis, lattice theory, SSA, optimization; fun extra: the final Winter 2018 lecture highlighted one of my favorite papers, https://pldi15.sigplan.org/details/pldi2015-papers/31/Provab...),

- UW CSE CSEP 501: Compilers - Hal Perkins (nice balanced introduction, including x86-64 assembly code generation, with the aforementioned "Engineering a Compiler" used as the course textbook).

matt_d··on Formulog: ML + Datalog + SMT
I can recommend the references listed in "Background: Notation" section of Program Analysis Resources: https://gist.github.com/MattPD/00573ee14bf85ccac6bed3c0678dd...
matt_d··on What Is the Minimal Set of Optimizations Needed for Zero-Cost Abstraction?
FWIW, this technique has been known in the C++ community as "hoisting", too (often applied to templates, particularly container class templates, but also smart pointers); cf. "Designing and Coding Reusable C++" by Carroll and Ellis, 1995 (http://cpptips.com/hoisting).

"Given this, what techniques can be used to reduce template instantiation time? One technique is called "hoisting." This is a generalization of the "wrappers for pointer containers" technique that showed up as a tip within the past week or so. The idea is quite simple: when writing a template class, consider each method in turn. If the method does not depend upon the template parameters, split the template class into a non-template base and a template derived class. Move (hoist) these parameter independent methods up into the base class. Note that with experience, you will begin to recognize opportunities for "hoisting" even in cases where methods initially appear to depend upon template parameters."

See also "thin template" (demonstrating the technique for containers, together with a Symbian OS adoption example, https://en.wikibooks.org/wiki/More_C%2B%2B_Idioms/Thin_Templ...) as well as "Minimizing Dependencies within Generic Classes for Faster and Smaller Programs", Dan Tsafrir, Robert W. Wisniewski, David F. Bacon, and Bjarne Stroustrup (ACM OOPSLA 2009). https://www.stroustrup.com/SCARY.pdf

"Reducing bloat by replacing inner classes with aliases can be further generalized to also apply to member methods of generic classes, which, like nested types, might uselessly depend on certain type parameters simply because they reside within a generic class’s scope. (Again, causing the compiler to uselessly generate many identical or nearly-identical instantiations of the same method.) To solve this problem we propose a “generalized hoisting” design paradigm, which decomposes a generic class into a hierarchy that eliminates unneeded dependencies."

matt_d··on Principles of Programming Languages (1997) [pdf]
General Program Analysis Resources may also be of interest: https://gist.github.com/MattPD/00573ee14bf85ccac6bed3c0678dd...
matt_d··on Hoare’s Rebuttal and Bubble Sort’s Comeback
See my other reply: https://news.ycombinator.com/item?id=23371399
matt_d··on Hoare’s Rebuttal and Bubble Sort’s Comeback
FWIW, there are some interesting tools for microarchitectural performance analysis which are getting closer (although it's a work in progress): https://github.com/MattPD/cpplinks/blob/master/performance.t...

One particularly interesting example in this context is OSACA (Open Source Architecture Code Analyzer), https://github.com/RRZE-HPC/osaca

The related publications explain the approach used to model and perform critical path analysis relevant for the modern superscalar out-of-order processors:

- Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels (2019): https://arxiv.org/abs/1910.00214

- Cross-Architecture Automatic Critical Path Detection For In-Core Performance Analysis (2020): https://hpc.fau.de/files/2020/02/Masterarbeit_JL-_final.pdf

There's also a broader line of research on performance modeling in this vein (some pretty detailed, including microarchitectural details like branch misprediction penalties, ROB capacity, etc.): https://gist.github.com/MattPD/85aad98ee8b135e675d49c571b67f...

More on modeling microarchitectural details (chronological order):

- Tejas S. Karkhanis and James E. Smith. "A First-Order Superscalar Processor Model." (ISCA 2004) - http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.79.....

- Stijn Eyerman, Lieven Eeckhout, Tejas Karkhanis, and James E. Smith. "A mechanistic performance model for superscalar out-of-order processors." ACM Trans. Comput. Syst. 27, 2 (2009) - http://www.elis.ugent.be/~leeckhou/papers/tocs09.pdf

- Maximilien B. Breughe, Stijn Eyerman, and Lieven Eeckhout. "Mechanistic analytical modeling of superscalar in-order processor performance." ACM Trans. Architec. Code Optim. 11, 4, Article 50 (2014) - https://users.elis.ugent.be/~leeckhou/papers/taco2015-breugh....

- "Modeling Superscalar Processor Memory-Level Parallelism", Sam Van den Steen and Lieven Eeckhout, IEEE Computer Architecture Letters (CAL), Vol 17, No 1 (2018) - https://users.elis.ugent.be/~leeckhou/papers/cal2018-MLP.pdf

matt_d··on Cpp-Taskflow: A General-Purpose Parallel and Heterogeneous Task System at Scale
Source code: https://github.com/cpp-taskflow/cpp-taskflow
matt_d··on GNU Binutils: The ELF Swiss Army Knife
Nice write-up!

FWIW, there's a bunch of other interesting tools for ELF, https://github.com/MattPD/cpplinks/blob/master/executables.m...

LIEF (Library to Instrument Executable Formats, https://github.com/lief-project/LIEF, an interesting project on its own) also has good references, https://lief.quarkslab.com/doc/latest/references.html

matt_d··on MLIR: A Compiler Infrastructure for the End of Moore's Law
See also: https://mlir.llvm.org/docs/LangRef/#blocks

"Context: The “block argument” representation eliminates a number of special cases from the IR compared to traditional “PHI nodes are operations” SSA IRs (like LLVM). For example, the parallel copy semantics of SSA is immediately apparent, and function arguments are no longer a special case: they become arguments to the entry block [ more rationale ]."

Block Arguments vs PHI nodes: https://mlir.llvm.org/docs/Rationale/#block-arguments-vs-phi...

matt_d··on Engineering Faster Sorters for Small Sets of Items
Perhaps it's worth considering succinct data structures (https://en.wikipedia.org/wiki/Succinct_data_structure), e.g.:

- BitMagic: http://bitmagic.io/, https://github.com/tlk00/BitMagic

- Succinct Data Structure Library 2.0: https://github.com/simongog/sdsl-lite, https://github.com/simongog/sdsl-lite/wiki/Literature

matt_d··on Is parallel programming hard, and, if so, what can you do about it?
Thanks!

FWIW, I've used the original post title ("Parallel Programming: December 2019 Update") to focus on the updates (and partially driven by the fear of comments made right after reading just the book title, "Is Parallel Programming Hard, And, If So, What Can You Do About It?", and focusing solely on answering the question while completely ignoring the content... :-]).

matt_d··on Clang Format Tanks Performance
Focusing on x86-64 (don't waste your time starting with the 32-bit variant when learning, let alone x87 FPU instructions, there's plenty of up-to-date materials for contemporary architectures nowadays):

- Chapter 10, Assembly Language (https://www3.nd.edu/~dthain/compilerbook/compilerbook.pdf#ch...) of http://compilerbook.org/

- intro_x86-64: Introduction to x86_64 assembly - https://gitlab.com/mcmfb/intro_x86-64

- Introduction to 64 Bit Assembly Language Programming for Linux and OS X - Ray Seyfarth - http://rayseyfarth.com/asm/

- Introduction to Computer Organization with x86-64 Assembly Language & GNU/Linux - Robert G. Plantz - http://bob.cs.sonoma.edu/IntroCompOrg-x64/book.html

- Modern X86 Assembly Language Programming - 2018; Daniel Kusswurm - Covers x86 64-bit, AVX, AVX2, and AVX-512 - https://github.com/Apress/modern-x86-assembly-language-progr...

- Understanding Assembly Language - a.k.a. Reverse Engineering for Beginners; https://yurichev.com/blog/UAL/ - https://beginners.re/

- x86-64 Assembly Language Programming with Ubuntu - Ed Jorgensen - http://www.egr.unlv.edu/~ed/x86.html

- Assembly Programming and Computer Architecture for Software Engineers (APCASE) - https://github.com/brianrhall/Assembly - Videos: https://www.youtube.com/channel/UCr0svQEez3UQvlj6-5EYS6w

Or, if you prefer talks:

- Just enough Assembly for Compiler Explorer - Anders Schau Knatten - NDC TechTown 2019 - https://www.youtube.com/watch?v=soeFwz0cOqU

- Modern x64 Assembly - https://www.youtube.com/playlist?list=PLKK11Ligqitg9MOX3-0tF...

- Bluff your way in x64 assembler - ACCU 2017; Roger Orr - https://www.youtube.com/watch?v=RI7VL-g6J7g

- Enough x86 Assembly to Be Dangerous - CppCon 2017; Charles Bailey - https://www.youtube.com/watch?v=IfUPkUAEwrk

More:

- Arm / AArch64: https://github.com/MattPD/cpplinks/blob/master/assembly.arm....

- RISC-V: https://github.com/MattPD/cpplinks/blob/master/assembly.risc...

- x86: https://github.com/MattPD/cpplinks/blob/master/assembly.x86....

matt_d··on Write Fuzzable Code
"Generating Software Tests" (https://www.fuzzingbook.org/) is pretty great (independent of your programming language) - arguably a must read for anyone interested in software testing.

John Regehr (the author of the blog post) has written more great posts:

- How to Fuzz an ADT Implementation - https://blog.regehr.org/archives/896

- Better Random Testing by Leaving Features Out - https://blog.regehr.org/archives/591

- Tricking a Whitebox Testcase Generator - https://blog.regehr.org/archives/672

- Fuzzers Need Taming - https://blog.regehr.org/archives/925

- Levels of Fuzzing - https://blog.regehr.org/archives/1039

- API Fuzzing vs. File Fuzzing: A Cautionary Tale - https://blog.regehr.org/archives/1269

- Reducers are Fuzzers - https://blog.regehr.org/archives/1284

In terms of software, DeepState (https://github.com/trailofbits/deepstate) may be a good place to start for C and C++. Relevant links:

- Fuzzing an API with DeepState: https://blog.trailofbits.com/2019/01/22/fuzzing-an-api-with-..., https://blog.trailofbits.com/2019/01/23/fuzzing-an-api-with-...

- NDSS 18 paper, "DeepState: Symbolic Unit Testing for C and C++": https://www.cefns.nau.edu/~adg326/bar18.pdf

In terms of choosing among fuzzing solutions, https://blog.trailofbits.com/2018/10/05/how-to-spot-good-fuz... is also worth a read -- as well as the article it refers to, http://www.pl-enthusiast.net/2018/08/23/evaluating-empirical.... For a broad survey, see "The Art, Science, and Engineering of Fuzzing": https://arxiv.org/abs/1812.00140, https://jiliac.com/pdf/fuzzing_survey19.pdf

More resources:

- Effective File Format Fuzzing – Thoughts, Techniques and Results (Black Hat Europe 2016): https://j00ru.vexillium.org/talks/blackhat-eu-effective-file...

- libFuzzer – a library for coverage-guided fuzz testing: http://tutorial.libFuzzer.info, http://llvm.org/docs/LibFuzzer.html, https://github.com/ouspg/libfuzzerfication

- Materials of "Modern fuzzing of C/C++ Projects" workshop: https://github.com/Dor1s/libfuzzer-workshop

- Introduction to using libFuzzer with llvm-toolset: https://developers.redhat.com/blog/2019/03/05/introduction-t...

- Fuzzing workflows - a fuzz job from start to finish: https://foxglovesecurity.com/2016/03/15/fuzzing-workflows-a-...

- Materials from "Fuzzing with AFL" workshop (SteelCon 2017, BSides London and Bristol 2019): https://github.com/ThalesIgnite/afl-training

- Making Your Library More Reliable with Fuzzing (C++Now 2018; Marshall Clow): https://www.youtube.com/watch?v=LlLJRHToyUk, https://github.com/boostcon/cppnow_presentations_2018/blob/m...

- C++ Weekly - Ep 85 - Fuzz Testing - https://www.youtube.com/watch?v=gO0KBoqkOoU

- The Art of Fuzzing – Slides and Demos: https://sec-consult.com/en/blog/2017/11/the-art-of-fuzzing-s...

matt_d··on Design and Evolution of C-Reduce (Part 1)
Updated URL: https://blog.regehr.org/archives/1678
matt_d··on A Look at the AMD Zen 2 Core
(Note: Also replied to another comment, but seems relevant here, too.)

There's a very approachable explanation in the following (around the 1:01:40 mark): https://youtu.be/8I_1TSs695I?t=1h1m40s -- part of Design of Digital Circuits - Lecture 18: Branch Prediction II (ETH Zürich, Spring 2019). Related readings: https://safari.ethz.ch/digitaltechnik/spring2019/doku.php?id....

These lectures are pretty great, by the way, highly recommended to anyone interested in computer architecture: http://people.inf.ethz.ch/omutlu/lecture-videos.html

Incidentally, this year's High-Performance Computer Architecture Test of Time Award has been given to "Dynamic Branch Prediction with Perceptrons" referenced in the lecture (from 2001, https://www.cs.utexas.edu/~lin/papers/hpca01.pdf): https://engineering.tamu.edu/news/2019/02/jimenez-receives-h....

matt_d··on A Look at the AMD Zen 2 Core
There's a very approachable explanation in the following (around the 1:01:40 mark): https://youtu.be/8I_1TSs695I?t=1h1m40s -- part of Design of Digital Circuits - Lecture 18: Branch Prediction II (ETH Zürich, Spring 2019). Related readings: https://safari.ethz.ch/digitaltechnik/spring2019/doku.php?id....

These lectures are pretty great, by the way, highly recommended to anyone interested in computer architecture: http://people.inf.ethz.ch/omutlu/lecture-videos.html

Incidentally, this year's High-Performance Computer Architecture Test of Time Award has been given to "Dynamic Branch Prediction with Perceptrons" referenced in the lecture (from 2001, https://www.cs.utexas.edu/~lin/papers/hpca01.pdf): https://engineering.tamu.edu/news/2019/02/jimenez-receives-h....

matt_d··on I Got a Knuth Check for 0x$3.00
The following (can be read in chronological order) give a pretty good idea:

- J.E. Smith and G.S. Sohi, "The Microarchitecture of Superscalar Processors," Proc. IEEE, vol. 83 (1995) - ftp://ftp.cs.wisc.edu/sohi/papers/1995/ieee-proc.superscalar.pdf, http://www.eng.ucy.ac.cy/theocharides/Courses/ECE656/supersc...

- Tejas S. Karkhanis and James E. Smith. "A First-Order Superscalar Processor Model." (ISCA 2004) - http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.79....

- Stijn Eyerman, Lieven Eeckhout, Tejas Karkhanis, and James E. Smith. "A mechanistic performance model for superscalar out-of-order processors." ACM Trans. Comput. Syst. 27, 2 (2009) - http://www.elis.ugent.be/~leeckhou/papers/tocs09.pdf

- Maximilien B. Breughe, Stijn Eyerman, and Lieven Eeckhout. "Mechanistic analytical modeling of superscalar in-order processor performance." ACM Trans. Architec. Code Optim. 11, 4, Article 50 (2014) - https://users.elis.ugent.be/~leeckhou/papers/taco2015-breugh...

- "Modeling Superscalar Processor Memory-Level Parallelism", Sam Van den Steen and Lieven Eeckhout, IEEE Computer Architecture Letters (CAL), Vol 17, No 1 (2018) - https://users.elis.ugent.be/~leeckhou/papers/cal2018-MLP.pdf

- A whirlwind introduction to dataflow graphs - https://fgiesen.wordpress.com/2018/03/05/a-whirlwind-introdu...

matt_d··on Show HN: IDE for Learning RISC-V
Here's a collection of RISC-V Instruction Set Architecture resources: https://github.com/MattPD/cpplinks/blob/master/assembly.risc...

Assembly tutorials: https://github.com/MattPD/cpplinks/blob/master/assembly.risc...

matt_d··on Design Continuums and the Path Toward Self-Designing Key-Value Stores That Know…
"The Research Challenge. The long-term challenge is whether we can easily or even automatically find the optimal storage design for a given problem. This has been recognized as an open problem since the early days of computer science. In his seminal 1978 paper, Robert Tarjan includes this problem in his list of the five major challenges for the future (which also included P Vs NP) [85]: “Is there a calculus of data structures by which one can choose the appropriate data representation and techniques for a given problem?”. We propose that a significant step toward a solution includes dealing with the following two challenges:

1) Can we know all possible data structure designs?

2) Can we compute the performance of any design?"

matt_d··on What makes processors fail – and how to prevent it
Slides (PDF): https://alastairreid.github.io/talks/what-makes-processors-f...
matt_d··on Fuzzing the .NET JIT Compiler
If you're interested, here's a collection of resources on compilers correctness (testing & fuzzing, validation, and verification): https://github.com/MattPD/cpplinks/blob/master/compilers.cor...
matt_d··on Numerical Tools for Non-Experts - Microsoft Research
- Demo: https://youtu.be/oYtnXEZC0jk?t=28m15s

- Herbgrind: http://herbgrind.ucsd.edu/

- Herbie: http://herbie.uwplse.org/

- FPBench: http://fpbench.org/

- Tools from the Floating-point Research World: http://fpbench.org/community.html

← PreviousPage 3 of 5Next →