MLIR Primer: A Compiler Infrastructure for the End of Moore’s Law
ai.google
ai.google
But this will not be a general library right ? You must have only included certain subset of functions of TF or PyTorch or whatever. Autodiffing is also included in certain proprietary libraries like ones from NAG. But I doubt its possible to achieve 50x speedup without compromising on functionality.
It mean you have efficient "crud" operators? This interest me because I'm building a relational language and toy with the idea of tensors as the "table" structure, but get rid of that because updates...
In short, you don't want to spend weeks on hand-optimized kernels, every time you want to implement a new type of layer for a large-scale network, or on an experimental target hardware.
The highest performance approach (on CPU) seems to be JIT assembly (as used in Intel OpenVINO). On Intel, that beats the living Jesus out of everything else, including my work. On ARM it's a free for all, and there's no clear leader, especially in quantized inference. On GPU whatever it is NVIDIA is doing is the right thing to do.
I'm skeptical that XLA/MLIR/TVM-like approaches can come close to, let alone exceed, the performance of hand-tuned kernels, for the same reason why hand-tuned assembly beats the shit of what compiler generates most of the time. I've yet to see it happen in practice. And you just need a few of those kernels, strategically placed where most of the computation happens, as per Pareto principle. For something like TF Google has the resources to get that done. It just chooses not to, to sell you more GPU-hours.
Google certainly has _technical_ capacity to do what PyTorch does, but not the _organizational_ capacity to scrap TF and start over. So they're trying to half-ass it with TF 2.0. It still sucks though.
Perhaps Google could copy PyTorch. But Tensorflow has a lot of overhead, and for both political/technical reasons, there's no easy path to go from Tensorflow to Pytorch.
You could just as easily ask: Why is Google sticking with {Hangouts/(Allo,Duo)/Angular} instead of doing/copying {Zoom/WhatsApp/React}? It's not like Google lacks the technical ability.
True, but it's important to note that if all the transistors in a modern CPU switched on at once, it would quickly overheat. This is the "power wall" -- we can squish more transistors in one die than we can actually turn on at one time due to electrical and thermal constraints.
> far, far more transistors in RAM than in the CPUs, overall
Also true, and this is an active area of research. Many people have tried various approaches to performing computations using DDR and other memory technologies. In the past, people were trying to trying to use DDR to run automata. These days there seems to be a lot of focus on processor-in-memory technologies; it turns out memristors can actually be used for computation, effectively turning the entire memory array into a hugely wide SIMD RISC processor. Here is some recent work presented on this subject:
Real Processing-in-Memory with Memristive Memory Processing Unit (mMPU) - https://ieeexplore.ieee.org/document/8825114
PPAC: A Versatile In-Memory Accelerator for Matrix-Vector-Product-Like Operations - https://ieeexplore.ieee.org/document/8825013
Parallel Stateful Logic in RRAM: Theoretical Analysis and Arithmetic Design - https://ieeexplore.ieee.org/document/8825150
> there's still a lot of room to grow, speed wise.
Yes and no. Single-threaded performance is close to tapping out. Production processes can only shrink so far before physics starts getting in the way. Pipelines and speculation can only get so deep (and broaden surface area for security vulnerabilities). Performance growth for massively parallel workloads is continuing along at a healthy clip, and will probably continue to do so for quite some time. Of course the trouble is that end-user desktop software is generally not massively parallel.
You're changing the effective throughput weather you do it by upping the clockrate or deepening the pipeline. Using half the data per cycle at twice the clock rate will cause the same memory pressure.
> which you can use with like explicit preloading into a tiny cache
That will kill it. As soon as you put it on the compiler designers or programmers to do something special to realize performance benefits, you're going to loose to architectures that don't.
Sure, compiler writers and programmers will optimize for your architecture... if it's popular and widely used. So you have a chicken and egg problem where you need to get adoption in the first place by running existing workloads faster.
> We can build 20 GHz CPUs now with passive cooling,
Citation? Like for real, that's cool and I'd like to read about it!
All the work that Google did around NBU (Next Billion Users), accessibility and Android are useless using this initiative.
They should have not chosen Swift. Kotlin Native would have worked equally well with day zero targetting for Android
https://resources.jetbrains.com/storage/products/kotlinconf2...
Naturally with Lather on board, it is Swift all the way, even though its support is WIP on Linux and nonexistent on Windows.
As for Kotlin/Native it is very imature, with incompatible memory semantics between other variants, and to be honest I don't see Kotlin ever taking off outside Android.