HNHacker News
TopNewBestAskShowJobs

matt_d

22,064 karma · joined April 21, 2014

submissionscomments
matt_d··on Introduction to LLVM [video]
Slides and code examples: http://www.mshah.io/fosdem18.html
matt_d··on RustBelt: securing the foundations of the Rust programming language
Two recent talks about the project:

- Derek Dreyer, Workshop on Software Correctness and Reliability 2017 (48 min): https://www.youtube.com/watch?v=Y9vemQmVeLI

- Ralf Jung, POPL 2018 (25 min): https://plv.mpi-sws.org/rustbelt/popl18/

matt_d··on Auto-Tuning Compiler Transformations with Machine Learning – Dr. Biagio Cosenza
Slides (PDF): http://biagiocosenza.com/talk/LLVM-Berlin-Meetup-Nov2017.pdf

Description: https://www.meetup.com/LLVM-Social-Berlin/events/244936204/

"To deliver higher performance, today's computer architecture has evolved in complexity. Hardware design is taking an irreversible step toward parallel architectures, which burdens application programmers in porting and tuning their codes. It is desirable to write programs that execute efficiently on highly parallel computing systems, but peak performance is notoriously hard to reach, and the valuable cost of wasting these precious resources motivates application programmers to devote significant time to tuning their codes.

Program automatic tuning (autotuning) is an emerging approach that relies on automated search or machine learning to off-load the traditionally time-consuming manual tuning of applications, and while it can apply to very different scenarios, it has become particularly important for parallel architectures.

This talk will show how machine learning can be a powerful tool to design portable and efficient autotuners. Machine learning application to this context is challenging and requires to address very specific problems (encoding, modeling, and training data availability). I will show practical examples of such autotuners for vectorization, loop unrolling, heterogeneous task partitioning and stencil computations, and which have been applied on a variety of compiler infrastructures such as LLVM, GCC, Insieme, and Patus. In particular, I will show how modeling can be greatly enhanced by structural learning methods that adapt to the structure of the problem."

matt_d··on CppCon 2017: Chandler Carruth “Going Nowhere Faster”
Slides: http://chandlerc.github.io/talks/cppcon2017/going_nowhere_fa...
matt_d··on 2017 LLVM Developers' Meeting Videos
Slides (coming soon) & more information: http://llvm.org/devmtg/2017-10/

2017 LLVM Lightning Talks: https://www.youtube.com/playlist?list=PL_R5A0lGi1ABrnDbIkbiX...

matt_d··on LLVM 5.0.0 Release
"LLVM compiler recognizes opportunities to transform a branch into IR select instruction(s) - later it will be lowered into X86::CMOV instruction, assuming no other optimization eliminated the SelectInst. However, it is not always profitable to emit X86::CMOV instruction. For example, branch is preferable over an X86::CMOV instruction when:

- Branch is well predicted

- Condition operand is expensive, compared to True-value and the False-value operands"

https://reviews.llvm.org/D34769

"We have seen periodically performance problems with cmov where one operand comes from memory. On modern x86 processors with strong branch predictors and speculative execution, this tends to be much better done with a branch than cmov. We routinely see cmov stalling while the load is completed rather than continuing, and if there are subsequent branches, they cannot be speculated in turn."

https://reviews.llvm.org/D36858

matt_d··on Dave Patterson: Evaluation of the Tensor Processing Unit
Event description: https://eecs.berkeley.edu/research/colloquium/170315

Slides: https://drive.google.com/file/d/0B5vMMuEXssYnR1oxcXVaaTFTQWM...

Paper: https://arxiv.org/abs/1704.04760

matt_d··on 2017 EuroLLVM Developers' Meeting
- Slides: http://llvm.org/devmtg/2017-03/

- Detailed program: http://llvm.org/devmtg/2017-03/2017/02/20/accepted-sessions....

matt_d··on EMI Testing: Finding 1000+ Bugs in GCC and LLVM in 3 Years – Zhendong Su
Slides: http://www.srl.inf.ethz.ch/workshop2016/Su.pdf

Video: https://www.youtube.com/watch?v=x4JUUlO9XGY

Homepage: http://web.cs.ucdavis.edu/~su/emi-project/

matt_d··on GCC code generation for C++ Weekly Ep 43 example
A `constexpr` function storing result in a `constexpr` variable (you need both for the compile-time evaluation guarantee): http://en.cppreference.com/w/cpp/language/constexpr

An extra complication in this example is that `std::accumulate` currently lacks `constexpr` support: http://stackoverflow.com/questions/32395408/why-arent-stdalg...

There's been a proposal to relax this limitation: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2016/p020...

For now you could use a workaround in form of a range-based for loop: https://godbolt.org/g/aVEqlm or even https://godbolt.org/g/GDrFr5 (homogeneous types).

A fold expression is much cleaner, though (also supporting heterogeneous types): https://godbolt.org/g/Hrnkc1

matt_d··on Concepts: The Future of Generic Programming [pdf]
I believe the contracts (preconditions, postconditions, assertions) proposal may address this issue.

The latest proposal is P0380, "A Contract Design": https://wg21.link/P0380.

An earlier work, "Simple Contracts for C++" (https://wg21.link/P0287), gives some further background and motivation.

Notably, contracts are currently solely in the proposal stage (and not in the upcoming standard).

matt_d··on Concepts: The Future of Generic Programming [pdf]
> can it be used as a type name of a variable or function argument?

If I understand correctly, placeholders may work for this. This use is described in the "Programming with placeholders" section of the "Introducing concepts" article: https://accu.org/index.php/journals/2157

The remaining articles in Andrew Sutton's C++ concepts series are pretty good, too; so, just in case anyone is interested, here are the links:

- "Defining Concepts" - https://accu.org/index.php/journals/2198

- "Overloading with Concepts" - https://accu.org/index.php/journals/2316

matt_d··on Intel X86 Encoder Decoder
Try Bitsavers (an amazing work in the service of historic preservation): http://www.bitsavers.org/

Previously discussed here: https://news.ycombinator.com/item?id=10143295

Internet Archive collection (the Intel documents part): https://archive.org/details/bitsavers_intel

Library of Congress reference: https://www.loc.gov/item/lcwa00096459/

matt_d··on Godbolt: Enter C, get Assembly
STOKE, a stochastic superoptimizer, is pretty interesting in this context: http://stoke.stanford.edu/ & https://github.com/StanfordPL/stoke
matt_d··on Developer Preview – EC2 Instances with Programmable Hardware
Here's a collection of get-started resources: http://tinyurl.com/fpga-resources

You can start with the EDA Playground tutorial, practice with HDLBits, while going through a book alongside (e.g., Harris & Harris) for examples, exercises, and best practices.

Similarly to a sibling thread, I'd also go with a free and open source flow, IceStorm (for the cheaply available iCE40 FPGAs): http://www.clifford.at/icestorm/

You can follow-up from the aforementioned tutorial and continue testing the designs on an iCE40 board -- starting here: http://hackaday.com/2015/08/19/learning-verilog-on-a-25-fpga...

Here are some really great presentations about it (slides & videos) by the creator (which can also serve in part as a general introduction):

- http://www.clifford.at/papers/2015/icestorm-flow/

- http://www.clifford.at/papers/2015/yosys-icestorm-etc/

Have fun!

matt_d··on A 32nm 1000-Processor Array
It's a modern version of the megahertz myth: https://en.wikipedia.org/wiki/Megahertz_myth#Modern_adaptati...

I think this makes even less sense as we go to 1,000 processing elements -- even Amdahl's law is too optimistic for the workloads deployed in these scenarios (i.e., having an embarrassingly parallel computation workload doesn't help if we still need to access memory, including going through the shared interconnect(s) to do this, while also keeping cache coherence in mind--and most workloads have phases that have to be synchronized, assuming we eventually want to write the results of our computation somewhere): http://blogs.msdn.com/b/ddperf/archive/2009/04/29/parallel-s...

matt_d··on Safe VSP – 30 year old Commodore 64 bug demystified (2013)
Reminded me a bit about the modern DRAM issue, i.e., row hammer -- although the details are of course different -- https://en.wikipedia.org/wiki/Row_hammer
matt_d··on A cache miss is not a cache miss
This discussion reminded me of a formula for calculating the average cost of a cache miss in this pretty cool paper ("An Analysis of the Effects of Miss Clustering on the Cost of a Cache Miss").

In particular, see Equation 4: http://researcher.ibm.com/files/us-viji/miss-cluster.pdf

matt_d··on Intel Kills “Tick-Tock”
Thanks for the reply!

> These numbers refer to the length of the channel in the MOSFET or FinFET between the Source and Drain dopant regions. [...] The only thing marketing has done here is fight between each other on whether or not channel size is the number to focus on, as it benefits their corporation or not.

Right -- I believe this is where the variety (or marketing) creeps in.

Here's what I mean: The various "gate length" definitions (printed, physical, effective, etc.) have a direct (and significant) impact on the ambiguity of the meaning of "process nodes" (that's in addition to the half-pitch interpretation in the DRAM industry); cf. Table 1: ITRS 2013 Data for CMOS Technology “Nodes” in http://semiengineering.com/a-node-by-any-other-name/

For reference (for anyone else reading, I know you know this), "channel length" vs. "gate length" is another difference to take into account: http://vlsi-soc.blogspot.com/2015/12/channel-length-vs-gate-... (More formally -- "Sidebar: Gate Length (Lg) versus Channel Length (L) and Experimental Data versus Equations" in http://www-inst.eecs.berkeley.edu/~ee130/sp06/chp7full.pdf)

Given that, I think it's fair to say the result is that this is more of a marketing name than a directly (in terms of physical feature sizes) interpretable number; as in:

- "At the December meeting, for example, Chenming Hu, the coinventor of the FinFET, began by mapping out the near future. Soon, he said, we’ll start to see 14-nm and 16-nm chips emerge (the first, which are expected to come from Intel, are slated to go into production early next year). Then he added a caveat whose casual tone belied its startling implications: “Nobody knows anymore what 16 nm means or what 14 nm means.”"

- "The switch to FinFETs has made the situation even more complex. Bohr points out, for example, that Intel’s 22-nm chips, the current state of the art, have FinFET transistors with gates that are 35 nm long but fins that are just 8 nm wide."

http://spectrum.ieee.org/semiconductors/devices/the-status-o...

Related: http://spectrum.ieee.org/semiconductors/design/shrinking-pos...

matt_d··on Intel Kills “Tick-Tock”
> Weird statement in the article: "it (TSCMs 7nm tech) should be very similar in terms of transistor density to Intel's 10-nanometer technology". This makes no sense as it would be comparing apples and pears. Surely they are referring to TSCMs 10nm tech?!

Process metrics are somewhat of a marketing nature nowadays (and not as directly related to the physical properties of the fabrication process as it may seem--at least not without adjusting for the differences between vendor-specific definitions in use).

Compare TSMC's 10nm to Intel's 14nm:

"In the case of TSMC they follow the “Foundry” node progress whereas Intel follows more of an “IDM” node transition 40nm versus 45nm, 28nm versus 32nm and 20nm versus 22nm. At the 14nm node TSMC has also chosen to call their node 16nm where everyone else is calling it 14nm."

Source: https://www.semiwiki.com/forum/content/3884-who-will-lead-10...

"Although the nominal gap in process nodes between Intel and TSMC appears to be narrowing, TSMC is not likely to catch up in terms of actual Moore’s Law scaling any time soon. TSMC’s 16FF+ process delivers only 20nm scaling, so they are still a generation behind Intel’s 14nm in terms of actual die area. TSMC said that 10nm shrinks by 0.52x from 16nm, nearly identical to the 0.53x scaling that Intel achieved from 22nm to 14nm. So if they stay on schedule, in 2017 TSMC will be in production on a 10nm process that is equivalent to the 14nm technology that Intel began producing in 2Q15. At that rate, even though Intel has slipped 10nm to 2H17, they will remain at least a year ahead of TSMC."

Source: http://www.eetimes.com/document.asp?doc_id=1327725

matt_d··on AVX2 faster than native popcnt instruction on Haswell/Skylake
Some of the resources I've ran across (assuming x86 context):

* Brief intros in Agner Fog's manuals:

- Chapter 12, Using vector operations in http://www.agner.org/optimize/optimizing_cpp.pdf

- Chapter 13, Vector programming in http://www.agner.org/optimize/optimizing_assembly.pdf

* http://www.whatsacreel.net76.net/asmtutes.html / https://www.youtube.com/playlist?list=PL0C5C980A28FEE68D

* Kusswurm (2014) "Modern X86 Assembly Language Programming - 32-bit, 64-bit, SSE, and AVX"

* Hughes (2015) "Single-Instruction Multiple-Data Execution"

* The author's website :-) http://wm.ite.pl/articles/#assembler-x86-fpu-mmx-sse-sse2-ss...

References:

* https://chessprogramming.wikispaces.com/x86-64

* https://software.intel.com/sites/landingpage/IntrinsicsGuide...

* x86 intrinsics cheat sheet (link at the very bottom): http://db.in.tum.de/~finis/

matt_d··on Decompiler Design
"No More Gotos: Decompilation Using Pattern-Independent Control-Flow Structuring and Semantics-Preserving Transformations" deserves a look: http://www.internetsociety.org/doc/no-more-gotos-decompilati...

Pattern-independent structuring (from the above paper) is implemented by fcd, an LLVM-based native program optimizing decompiler: https://github.com/zneak/fcd/

Related blog posts have some details and examples:

- http://zneak.github.io/fcd/2016/02/24/seseloop.html

- http://zneak.github.io/fcd/2016/02/17/structuring.html

- http://zneak.github.io/fcd/2016/02/21/csaw-wyvern.html

matt_d··on QIRA is a timeless debugger
How about rr? // http://rr-project.org/
matt_d··on Writing high-performance servers in modern C++
All right, then: This one is complete (compiles, runs, iterates through the data, accessing only fields "x" and "y"):

  #include <fstream>
  #include <json.hpp>
  
  int main()
  {
  	std::ifstream input_file("data.json");
  	nlohmann::json data;
  	data << input_file;
  	for (auto obj : data)
  		std::cout << obj["x"] << ',' << obj["y"] << '\n';
  }
matt_d··on Writing high-performance servers in modern C++
Take a look at JSON for Modern C++: https://github.com/nlohmann/json

Perhaps this example (benchmark, actually) is close enough: https://github.com/nlohmann/json/blob/master/benchmarks/benc...

Note that the above also performs mini-benchmark with multiple iterations, dumping parsed data to another file, and finishes with a clean-up of the aforementioned dump target file.

The core (read, parse, write) can be simplified to:

  std::ifstream input_file("data.json");
  nlohmann::json data;
  data << input_file;
  std::cout << data;
matt_d··on Writing high-performance servers in modern C++
I'd recommend to take a look at Seastar: http://www.seastar-project.org/

Briefly: "Seastar is event-driven and supports writing non-blocking, asynchronous server code in a straightforward manner that facilitates debugging and reasoning about performance." -- http://www.scylladb.com/2015/02/20/seastar/

Concurrency model: "Seastar futures/promises/continuations (f-p-c) are a subset of reactive programming.

Seastar performance derives from the sharded, cooperative, non-blocking, micro-task scheduled design, and f-p-c are a friendlier way of feeding tasks to the scheduler.

Seastar’s f-p-c do have some optimizations relative to other f-p-c designs. They trade off thread safety, which is unneeded due to the sharded design, for scheduling efficiency, and have a very low memory footprint." -- http://www.seastar-project.org/faq/

If you'd like to find out more, here are a few links with more information (including examples and tutorials):

- https://github.com/scylladb/seastar/wiki

- https://github.com/scylladb/seastar/blob/master/doc/tutorial...

- http://www.seastar-project.org/futures-promises/

- http://blog.cloudius-systems.com/2015/04/29/seastar-tutorial...

- http://www.slideshare.net/TzachLivyatan/seastar-sayeret-lamb...

In terms of practical applications, ScyllaDB (fully compatible--and significantly faster--drop-in replacement of Apache Cassandra) relies on Seastar: http://www.scylladb.com/

ScyllaDB has been discussed here before: https://news.ycombinator.com/item?id=10262719

matt_d··on Non-volatile Storage: CPUs no longer more performant than I/O devices
On a side note, it's interesting to me that emerging memory technologies currently seem to be mainly focused on addressing the "from-DRAM-to-disk" part of the memory hierarchy.

That is, as you mentioned, not directly competing with DRAM, and consistently on the same side of the 1 microsecond dividing line between memory and storage; as in:

http://www.rambusblog.com/2015/10/15/mid-when-memory-and-sto... (note SCM placed between DRAM and SSD)

http://semiengineering.com/the-memory-and-storage-hierarchy/

As far as the other side of the line is concerned, I think I've only seen proposals for hybrid-cache architectures (HCA) -- other than http://link.springer.com/chapter/10.1007%2F978-1-4419-9551-3... -- with a hybrid approach (e.g., combining SRAM/eDRAM/STT-RAM/PCRAM) probably making sense due to latency/endurance/bandwidth trade-offs.

If anything, there seems to be more development on the DRAM interface itself -- with multiple candidates for the (or a) DDR4's successor, so far involving Wide I/O (Samsung), Hybrid Memory Cube (Intel, Micron), High Bandwidth Memory (SK Hynix, AMD, Nvidia): http://www.extremetech.com/computing/197720-beyond-ddr4-unde...

(Latency and bandwidth improvements seem promising, http://semiengineering.com/which-memory-type-should-you-use/)

One interesting development I've seen involves reducing SRAM's footprint, by moving from 6T (6-transistors) cell to a 1T (one-transistor) one: http://www.eetimes.com/document.asp?doc_id=1328453

It's fairly recent development, though, and it remains to be seen how is it going to fare.

Other than the above, there doesn't really seem to be much progress around competing with/improving SRAM. However, this may become increasingly important, since some of the technological process scaling issues apply to SRAM, too.

matt_d··on Non-volatile Storage: CPUs no longer more performant than I/O devices
Interesting!

Incidentally (since this may be somewhat related), I'm wondering, what are your thoughts on the Persistent Memory Manager approach, as in the following:

Justin Meza, Yixin Luo, Samira Khan, Jishen Zhao, Yuan Xie, and Onur Mutlu: "A Case for Efficient Hardware/Software Cooperative Management of Storage and Memory." Workshop on Energy-Efficient Design, 2013.

Context: "emerging high-performance NVM technologies enable a renewed focus on the unification of storage and memory: a hardware-accelerated single-level store, or persistent memory, which exposes a large, persistent virtual address space supported by hardware-accelerated management of heterogeneous storage and memory devices. The implications of such an interface for system efficiency are immense: A persistent memory can provide a unified load/store-like interface to access all data in a system without the overhead of software-managed metadata storage and retrieval and with hardware-assisted data persistence guarantees."

The stated goals/benefits include eliminating operating system calls for file operations, eliminating file system operations, and efficient data mapping.

Paper: http://justinmeza.com/bin/meza_weed13.pdf

Presentation: https://users.ece.cmu.edu/~omutlu/pub/mutlu_weed13_talk.pdf

matt_d··on Test results for Broadwell and Skylake
Thank you for the reply!

Interesting about parallel misses handling, thanks!

One worry is that this tends to compound other effects -- say, non-prefetch-friendly access combined with TLB misses resulting in increasingly expensive slowdowns (as in the continuous-vs.-random array access example in the paper).

matt_d··on Test results for Broadwell and Skylake
I realize that this is another topic, although in a somewhat similar context, so I thought I may just ask: Have there been any advances in reducing the page walk latency?

I'm thinking of the virtual address translation costs having impact on the run times of common algorithms, e.g., as demonstrated in the following work by Jurkiewicz & Mehlhorn: http://arxiv.org/abs/1212.0703

The recent research I'm aware of is, e.g., Generalized Large-page Utilization Enhancements (GLUE) mechanism, proposed in "Large Pages and Lightweight Memory Management in Virtualized Environments" (from this year's Micro): slides: https://dl.dropboxusercontent.com/u/36554102/BPC-1.pdf ; paper: http://paul.rutgers.edu/~binhpham/phamMICRO15.pdf

Admittedly, it focuses specifically on one aspect (the Double Address Translation on Virtual Machines issue in the Jurkiewicz & Mehlhorn context).

What I'm wondering about is: Has there been any progress on that on the "practical implementation" side, in the recent/coming Intel (or other, for that matter) CPUs?

← PreviousPage 4 of 5Next →