22,064 karma · joined April 21, 2014
- Derek Dreyer, Workshop on Software Correctness and Reliability 2017 (48 min): https://www.youtube.com/watch?v=Y9vemQmVeLI
- Ralf Jung, POPL 2018 (25 min): https://plv.mpi-sws.org/rustbelt/popl18/
Description: https://www.meetup.com/LLVM-Social-Berlin/events/244936204/
"To deliver higher performance, today's computer architecture has evolved in complexity. Hardware design is taking an irreversible step toward parallel architectures, which burdens application programmers in porting and tuning their codes. It is desirable to write programs that execute efficiently on highly parallel computing systems, but peak performance is notoriously hard to reach, and the valuable cost of wasting these precious resources motivates application programmers to devote significant time to tuning their codes.
Program automatic tuning (autotuning) is an emerging approach that relies on automated search or machine learning to off-load the traditionally time-consuming manual tuning of applications, and while it can apply to very different scenarios, it has become particularly important for parallel architectures.
This talk will show how machine learning can be a powerful tool to design portable and efficient autotuners. Machine learning application to this context is challenging and requires to address very specific problems (encoding, modeling, and training data availability). I will show practical examples of such autotuners for vectorization, loop unrolling, heterogeneous task partitioning and stencil computations, and which have been applied on a variety of compiler infrastructures such as LLVM, GCC, Insieme, and Patus. In particular, I will show how modeling can be greatly enhanced by structural learning methods that adapt to the structure of the problem."
2017 LLVM Lightning Talks: https://www.youtube.com/playlist?list=PL_R5A0lGi1ABrnDbIkbiX...
- Branch is well predicted
- Condition operand is expensive, compared to True-value and the False-value operands"
https://reviews.llvm.org/D34769
"We have seen periodically performance problems with cmov where one operand comes from memory. On modern x86 processors with strong branch predictors and speculative execution, this tends to be much better done with a branch than cmov. We routinely see cmov stalling while the load is completed rather than continuing, and if there are subsequent branches, they cannot be speculated in turn."
- Detailed program: http://llvm.org/devmtg/2017-03/2017/02/20/accepted-sessions....
An extra complication in this example is that `std::accumulate` currently lacks `constexpr` support: http://stackoverflow.com/questions/32395408/why-arent-stdalg...
There's been a proposal to relax this limitation: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2016/p020...
For now you could use a workaround in form of a range-based for loop: https://godbolt.org/g/aVEqlm or even https://godbolt.org/g/GDrFr5 (homogeneous types).
A fold expression is much cleaner, though (also supporting heterogeneous types): https://godbolt.org/g/Hrnkc1
The latest proposal is P0380, "A Contract Design": https://wg21.link/P0380.
An earlier work, "Simple Contracts for C++" (https://wg21.link/P0287), gives some further background and motivation.
Notably, contracts are currently solely in the proposal stage (and not in the upcoming standard).
If I understand correctly, placeholders may work for this. This use is described in the "Programming with placeholders" section of the "Introducing concepts" article: https://accu.org/index.php/journals/2157
The remaining articles in Andrew Sutton's C++ concepts series are pretty good, too; so, just in case anyone is interested, here are the links:
- "Defining Concepts" - https://accu.org/index.php/journals/2198
- "Overloading with Concepts" - https://accu.org/index.php/journals/2316
Previously discussed here: https://news.ycombinator.com/item?id=10143295
Internet Archive collection (the Intel documents part): https://archive.org/details/bitsavers_intel
Library of Congress reference: https://www.loc.gov/item/lcwa00096459/
You can start with the EDA Playground tutorial, practice with HDLBits, while going through a book alongside (e.g., Harris & Harris) for examples, exercises, and best practices.
Similarly to a sibling thread, I'd also go with a free and open source flow, IceStorm (for the cheaply available iCE40 FPGAs): http://www.clifford.at/icestorm/
You can follow-up from the aforementioned tutorial and continue testing the designs on an iCE40 board -- starting here: http://hackaday.com/2015/08/19/learning-verilog-on-a-25-fpga...
Here are some really great presentations about it (slides & videos) by the creator (which can also serve in part as a general introduction):
- http://www.clifford.at/papers/2015/icestorm-flow/
- http://www.clifford.at/papers/2015/yosys-icestorm-etc/
Have fun!
I think this makes even less sense as we go to 1,000 processing elements -- even Amdahl's law is too optimistic for the workloads deployed in these scenarios (i.e., having an embarrassingly parallel computation workload doesn't help if we still need to access memory, including going through the shared interconnect(s) to do this, while also keeping cache coherence in mind--and most workloads have phases that have to be synchronized, assuming we eventually want to write the results of our computation somewhere): http://blogs.msdn.com/b/ddperf/archive/2009/04/29/parallel-s...
In particular, see Equation 4: http://researcher.ibm.com/files/us-viji/miss-cluster.pdf
> These numbers refer to the length of the channel in the MOSFET or FinFET between the Source and Drain dopant regions. [...] The only thing marketing has done here is fight between each other on whether or not channel size is the number to focus on, as it benefits their corporation or not.
Right -- I believe this is where the variety (or marketing) creeps in.
Here's what I mean: The various "gate length" definitions (printed, physical, effective, etc.) have a direct (and significant) impact on the ambiguity of the meaning of "process nodes" (that's in addition to the half-pitch interpretation in the DRAM industry); cf. Table 1: ITRS 2013 Data for CMOS Technology “Nodes” in http://semiengineering.com/a-node-by-any-other-name/
For reference (for anyone else reading, I know you know this), "channel length" vs. "gate length" is another difference to take into account: http://vlsi-soc.blogspot.com/2015/12/channel-length-vs-gate-... (More formally -- "Sidebar: Gate Length (Lg) versus Channel Length (L) and Experimental Data versus Equations" in http://www-inst.eecs.berkeley.edu/~ee130/sp06/chp7full.pdf)
Given that, I think it's fair to say the result is that this is more of a marketing name than a directly (in terms of physical feature sizes) interpretable number; as in:
- "At the December meeting, for example, Chenming Hu, the coinventor of the FinFET, began by mapping out the near future. Soon, he said, we’ll start to see 14-nm and 16-nm chips emerge (the first, which are expected to come from Intel, are slated to go into production early next year). Then he added a caveat whose casual tone belied its startling implications: “Nobody knows anymore what 16 nm means or what 14 nm means.”"
- "The switch to FinFETs has made the situation even more complex. Bohr points out, for example, that Intel’s 22-nm chips, the current state of the art, have FinFET transistors with gates that are 35 nm long but fins that are just 8 nm wide."
http://spectrum.ieee.org/semiconductors/devices/the-status-o...
Related: http://spectrum.ieee.org/semiconductors/design/shrinking-pos...
Process metrics are somewhat of a marketing nature nowadays (and not as directly related to the physical properties of the fabrication process as it may seem--at least not without adjusting for the differences between vendor-specific definitions in use).
Compare TSMC's 10nm to Intel's 14nm:
"In the case of TSMC they follow the “Foundry” node progress whereas Intel follows more of an “IDM” node transition 40nm versus 45nm, 28nm versus 32nm and 20nm versus 22nm. At the 14nm node TSMC has also chosen to call their node 16nm where everyone else is calling it 14nm."
Source: https://www.semiwiki.com/forum/content/3884-who-will-lead-10...
"Although the nominal gap in process nodes between Intel and TSMC appears to be narrowing, TSMC is not likely to catch up in terms of actual Moore’s Law scaling any time soon. TSMC’s 16FF+ process delivers only 20nm scaling, so they are still a generation behind Intel’s 14nm in terms of actual die area. TSMC said that 10nm shrinks by 0.52x from 16nm, nearly identical to the 0.53x scaling that Intel achieved from 22nm to 14nm. So if they stay on schedule, in 2017 TSMC will be in production on a 10nm process that is equivalent to the 14nm technology that Intel began producing in 2Q15. At that rate, even though Intel has slipped 10nm to 2H17, they will remain at least a year ahead of TSMC."
* Brief intros in Agner Fog's manuals:
- Chapter 12, Using vector operations in http://www.agner.org/optimize/optimizing_cpp.pdf
- Chapter 13, Vector programming in http://www.agner.org/optimize/optimizing_assembly.pdf
* http://www.whatsacreel.net76.net/asmtutes.html / https://www.youtube.com/playlist?list=PL0C5C980A28FEE68D
* Kusswurm (2014) "Modern X86 Assembly Language Programming - 32-bit, 64-bit, SSE, and AVX"
* Hughes (2015) "Single-Instruction Multiple-Data Execution"
* The author's website :-) http://wm.ite.pl/articles/#assembler-x86-fpu-mmx-sse-sse2-ss...
References:
* https://chessprogramming.wikispaces.com/x86-64
* https://software.intel.com/sites/landingpage/IntrinsicsGuide...
* x86 intrinsics cheat sheet (link at the very bottom): http://db.in.tum.de/~finis/
Pattern-independent structuring (from the above paper) is implemented by fcd, an LLVM-based native program optimizing decompiler: https://github.com/zneak/fcd/
Related blog posts have some details and examples:
- http://zneak.github.io/fcd/2016/02/24/seseloop.html
#include <fstream>
#include <json.hpp>
int main()
{
std::ifstream input_file("data.json");
nlohmann::json data;
data << input_file;
for (auto obj : data)
std::cout << obj["x"] << ',' << obj["y"] << '\n';
}Perhaps this example (benchmark, actually) is close enough: https://github.com/nlohmann/json/blob/master/benchmarks/benc...
Note that the above also performs mini-benchmark with multiple iterations, dumping parsed data to another file, and finishes with a clean-up of the aforementioned dump target file.
The core (read, parse, write) can be simplified to:
std::ifstream input_file("data.json");
nlohmann::json data;
data << input_file;
std::cout << data;Briefly: "Seastar is event-driven and supports writing non-blocking, asynchronous server code in a straightforward manner that facilitates debugging and reasoning about performance." -- http://www.scylladb.com/2015/02/20/seastar/
Concurrency model: "Seastar futures/promises/continuations (f-p-c) are a subset of reactive programming.
Seastar performance derives from the sharded, cooperative, non-blocking, micro-task scheduled design, and f-p-c are a friendlier way of feeding tasks to the scheduler.
Seastar’s f-p-c do have some optimizations relative to other f-p-c designs. They trade off thread safety, which is unneeded due to the sharded design, for scheduling efficiency, and have a very low memory footprint." -- http://www.seastar-project.org/faq/
If you'd like to find out more, here are a few links with more information (including examples and tutorials):
- https://github.com/scylladb/seastar/wiki
- https://github.com/scylladb/seastar/blob/master/doc/tutorial...
- http://www.seastar-project.org/futures-promises/
- http://blog.cloudius-systems.com/2015/04/29/seastar-tutorial...
- http://www.slideshare.net/TzachLivyatan/seastar-sayeret-lamb...
In terms of practical applications, ScyllaDB (fully compatible--and significantly faster--drop-in replacement of Apache Cassandra) relies on Seastar: http://www.scylladb.com/
ScyllaDB has been discussed here before: https://news.ycombinator.com/item?id=10262719
That is, as you mentioned, not directly competing with DRAM, and consistently on the same side of the 1 microsecond dividing line between memory and storage; as in:
http://www.rambusblog.com/2015/10/15/mid-when-memory-and-sto... (note SCM placed between DRAM and SSD)
http://semiengineering.com/the-memory-and-storage-hierarchy/
As far as the other side of the line is concerned, I think I've only seen proposals for hybrid-cache architectures (HCA) -- other than http://link.springer.com/chapter/10.1007%2F978-1-4419-9551-3... -- with a hybrid approach (e.g., combining SRAM/eDRAM/STT-RAM/PCRAM) probably making sense due to latency/endurance/bandwidth trade-offs.
If anything, there seems to be more development on the DRAM interface itself -- with multiple candidates for the (or a) DDR4's successor, so far involving Wide I/O (Samsung), Hybrid Memory Cube (Intel, Micron), High Bandwidth Memory (SK Hynix, AMD, Nvidia): http://www.extremetech.com/computing/197720-beyond-ddr4-unde...
(Latency and bandwidth improvements seem promising, http://semiengineering.com/which-memory-type-should-you-use/)
One interesting development I've seen involves reducing SRAM's footprint, by moving from 6T (6-transistors) cell to a 1T (one-transistor) one: http://www.eetimes.com/document.asp?doc_id=1328453
It's fairly recent development, though, and it remains to be seen how is it going to fare.
Other than the above, there doesn't really seem to be much progress around competing with/improving SRAM. However, this may become increasingly important, since some of the technological process scaling issues apply to SRAM, too.
Incidentally (since this may be somewhat related), I'm wondering, what are your thoughts on the Persistent Memory Manager approach, as in the following:
Justin Meza, Yixin Luo, Samira Khan, Jishen Zhao, Yuan Xie, and Onur Mutlu: "A Case for Efficient Hardware/Software Cooperative Management of Storage and Memory." Workshop on Energy-Efficient Design, 2013.
Context: "emerging high-performance NVM technologies enable a renewed focus on the unification of storage and memory: a hardware-accelerated single-level store, or persistent memory, which exposes a large, persistent virtual address space supported by hardware-accelerated management of heterogeneous storage and memory devices. The implications of such an interface for system efficiency are immense: A persistent memory can provide a unified load/store-like interface to access all data in a system without the overhead of software-managed metadata storage and retrieval and with hardware-assisted data persistence guarantees."
The stated goals/benefits include eliminating operating system calls for file operations, eliminating file system operations, and efficient data mapping.
Paper: http://justinmeza.com/bin/meza_weed13.pdf
Presentation: https://users.ece.cmu.edu/~omutlu/pub/mutlu_weed13_talk.pdf
Interesting about parallel misses handling, thanks!
One worry is that this tends to compound other effects -- say, non-prefetch-friendly access combined with TLB misses resulting in increasingly expensive slowdowns (as in the continuous-vs.-random array access example in the paper).
I'm thinking of the virtual address translation costs having impact on the run times of common algorithms, e.g., as demonstrated in the following work by Jurkiewicz & Mehlhorn: http://arxiv.org/abs/1212.0703
The recent research I'm aware of is, e.g., Generalized Large-page Utilization Enhancements (GLUE) mechanism, proposed in "Large Pages and Lightweight Memory Management in Virtualized Environments" (from this year's Micro): slides: https://dl.dropboxusercontent.com/u/36554102/BPC-1.pdf ; paper: http://paul.rutgers.edu/~binhpham/phamMICRO15.pdf
Admittedly, it focuses specifically on one aspect (the Double Address Translation on Virtual Machines issue in the Jurkiewicz & Mehlhorn context).
What I'm wondering about is: Has there been any progress on that on the "practical implementation" side, in the recent/coming Intel (or other, for that matter) CPUs?