HNHacker News
TopNewBestAskShowJobs

Marat_Dukhan

625 karma · joined November 3, 2012

XNNPACK TLM @ Google, previously QNNPACK lead @ Facebook, Ph.D. student @ Georgia Tech, and the author of NNPACK library.

My Ph.D. projects:

www.github.com/Maratyszcza/NNPACK

www.github.com/Maratyszcza/PeachPy

www.github.com/Maratyszcza/Opcodes

Cool demos:

www.github.com/Maratyszcza/blis-bench

www.github.com/Maratyszcza/laff-demos

submissionscomments
Marat_Dukhan··on tolower() in bulk at speed
Linux-capable RISC-V cores often have 64-bit architecture and no SIMD/vector processing capabilities.
Marat_Dukhan··on Ampere Announces 5nm Arm Server CPU
IMO the author alludes to some enterprise software running on Wintel that has per-core licensing costs.
Marat_Dukhan··on Ten Lessons from Three Generations Shaped Google’s TPUv4i [pdf]
He's a Software Engineer on the TPU team. Are you confusing him for Thomas Kurian, GCloud SVP?

Note: I work for Google, but speak for myself.

Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
If by acceleration you mean offloading inference to a different IP block (GPU/DSP/NPU), then yes. XNNPACK is the inference engine for CPU.

CPU is the default backend in TensorFlow Lite, and CPU inference always works and produce correct result. GPU/DSP/NPU inference can be faster, particularly for large models on high-end SoCs, but generally you need to make sure that the model is supported on the IP block, the result is correct and performance is better than the CPU baseline. And that quickly gets very complicated:

1. NN API, and TFLite GPU/DSP backends support a limited subset of all TensorFlow Lite operators, and if a model is only partially offloaded to GPU/DSP/NPU, part of it will still run on CPU, and commonly synchronization overhead kills all potential speedups of the specialized hardware. The situation is even worse in CoreML, as CoreML doesn't provide an API to even learn which operators failed to offload to GPU/NPU.

2. Bugs in GPU shader compilers and NN API drivers do happen, and unless your model is a standard MobileNet, you're likely to hit them at least on some mobile phones. Then you'd need an infrastructure to detect this situation and disable offloading the model to this IP block on particular phones.

3. Low-end SoCs usually completely lack DSP and NPU, and their GPU is often slower than CPU even in nominal peak performance. This happens because CPU cores in low-end SoCs are typically just downclocked versions of the CPU cores in high-end SoCs, but low-end GPUs have 8-16 times fewer GPU cores than their high-end counterparts.

Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
In order to benefit from optimizations in *this blog post* the model needs to be quantized to 8-bit integers. However, XNNPACK supports floating-point inference as well (including with FP16 weights), see https://blog.tensorflow.org/2020/07/accelerating-tensorflow-...
Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
It performs fixed-point arithmetic on 8-bit integers. You can mimick lower than 8-bit precision by using output_min/output_max parameters in XNNPACK operators, but keep in mind that: 1. This functionality is experimental and not exposed in TFLite. You'd need to call XNNPACK APIs directly from C/C++ code. 2. Computations would still be done on 8-bit numbers.
Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
Yes, these optimizations work with existing tflite models, so long as the quantized operators they use are supported in XNNPACK.
Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
TensorFlow doesn't support quantized inference (it supports only mimicking quantization in floating-point for quantization-aware training), so it can't immediately benefit from these optimizations.
Marat_Dukhan··on Faster Quantized Neural Network Inference with XNNPack
Author here, happy to take your questions.
Marat_Dukhan··on Apple Starts Work on Its Own Cellular Modem
Good. I was surprised that Apple Silicon Macs don't have a built-in cellular modem. This reminds me how iPhone launched without a 3G modem, and I hope Apple will similarly fix the lack of cellular connectivity in the next generation of MacBooks.
Marat_Dukhan··on Supercharging TensorFlow.js with SIMD and multi-threading
Even WebGL2 doesn't expose compute shaders, so any NN computations work by abusing the graphics pipeline, with many inefficiencies involved. Shader dispatch is more expensive, no access to local memory, no control over dispatch blocks. Hopefully the upcoming WebGPU specification will close these efficiency gaps.
Marat_Dukhan··on UK officially in recession for first time in 11 years
What a time to be alive!
Marat_Dukhan··on Open-sourcing FBGEMM for state-of-the-art server-side inference
Performance on the plot is higher than FP32 peak, but there's no error - because FBGEMM does not compute in FP32, it computes in 8-bit fixed point. On a Broadwell CPU, you can do 16 FP32 multiply-adds (2x 8-wide FMA instructions via VFMAxxxPS instructions), but 32 8-bit multiply adds (1x 32-wide multiplication with accumulation of adjacent results via VPMADDUSBW instruction).
Marat_Dukhan··on Open-sourcing FBGEMM for state-of-the-art server-side inference
FBGEMM is faster than theoretical peak FP32 (single-precision floating-point) performance, therefore its faster than SGEMM/DGEMM in any BLAS library
Marat_Dukhan··on Qnnpack: PyTorch-integrated open source library for mobile deep learning
QNNPACK directly competes with the CPU backend of TensorFlow Lite and the gemmlowp library. The Caffe2 backend of PyTorch 1.0 integrates QNNPACK, and directly competes with TensorFlow Lite. QNNPACK targets only mobile CPUs, but Caffe2 integrates other backends for non-CPU targets, e.g. Apple's MPSCNN library for iPhone GPUs, Qualcomm's Snapdragon NPE for Qualcomm GPUs and DSPs, ARM ComputeLibrary for Android GPUs. Not sure what you mean by TensorFlow Cores: NVIDIA has TensorCores and TensorRT, and Google has Tensor Processing Units (TPU), but neither of these technologies are for mobile.
Marat_Dukhan··on Shipping a Neural Network on iOS with CoreML, PyTorch, and React Native
You can use the same toolchain to convert PyTorch model to Caffe2 through ONNX. Caffe2 supports both Android and iOS. There is even a tutorial: http://pytorch.org/tutorials/advanced/super_resolution_with_...
Marat_Dukhan··on Feasibility of low-level GPU access on the Web
It is possible to perform some computations using OpenGL ES 3.0 / WebGL 2.0, but many types of operations (e.g. anything that involves random-access writes) are impossible, and many others (anything that normally requires shared memory) are very inefficient. Programming GPU through WebGL 2.0 is akin to programming desktop GPUs pre-CUDA: it is too intricate to take off.

Compute shader extension for WebGL 2.0 would be cool, but it would require to port a large part of OpenGL ES 3.1: OpenGL ES 3.0 / WebGL 2.0 doesn't include even random access buffers (SSBOs)

Marat_Dukhan··on Feasibility of low-level GPU access on the Web
WebGL 2 is based on OpenGL ES 3.0, it doesn't give you compute. Compute shaders were added in OpenGL ES 3.1
Marat_Dukhan··on Show HN: detecting cache latency inside a Web browser
Hmm...I just tried on iPhone 7/iOS 11.2.2, and it still works, albeit takes very long to start.
Marat_Dukhan··on Show HN: detecting cache latency inside a Web browser
Asm.js is not necessary, simple JavaScript interpreter is enough. This demo used to work before most browsers implemented optimizers for Asm.js
Marat_Dukhan··on Show HN: detecting cache latency inside a Web browser
No, it is a static web page, and all code runs only locally in your browser.
Marat_Dukhan··on Show HN: detecting cache latency inside a Web browser
Author here. I made this demo and a related matrix-matrix multiplication demo [1] back in 2015 for Robert van de Geijn's Linear Algebra: Foundations to Frontiers MOOC class [2]. In the light of Spectre attack and recent browsers' changes to reduce precision of timers, I remembered of this project, and decided to check if it still works now, 3 years later. Surprisingly, it still works well!

The source code is available on GitHub [3].

[1] https://maratyszcza.github.io/laff-demos/dgemm.html

[2] https://www.edx.org/course/linear-algebra-foundations-fronti...

[3] https://github.com/Maratyszcza/laff-demos

Marat_Dukhan··on Ask HN: What did you work on in 2017?
CPU INFOrmation library: a cross-platform library to discover supported instruction sets, microarchitecture, and cache parameters of the CPU. Started as a "oh, I can do it over the weekend" project at first, took close to a year to get to production quality.

https://github.com/Maratyszcza/cpuinfo

Marat_Dukhan··on Making Pillow-SIMD, optimizing image processing in Python
There is transfer time. You source image is in cacheable CPU memory. Integrated GPUs normally work with uncacheable memory allocated in a special region of system memory. Some GPUs can access cacheable memory too, but it is much slower (because it has to maintain coherency with CPU caches), and requires that you allocate such cacheable memory using special OpenCL driver calls, not your normal malloc. So, in practice, you would do a copy to GPU-optimized buffer (in shared with CPU, but uncacheable memory).
Marat_Dukhan··on Tensorflow sucks
It should, if you find that converted model doesn't work, it is a bug, please file it with PyTorch and/or Caffe2.

Implementation details of course differ between frameworks, but luckily neural networks are very robust to noise. In my experience, changes like using Winograd/Fourier for convolutions, or even running the whole thing in FP16 do not result in noticeable artifacts, and these are among the biggest differences you could have between frameworks.

Marat_Dukhan··on Tensorflow sucks
Reference specification and validator are hosted in https://github.com/onnx/onnx
Marat_Dukhan··on Tensorflow sucks
Have you looked at ONNX? It is a neural network exchange format that, in particular, lets you deploy PyTorch models in production via Caffe2. Here is a tutorial: http://pytorch.org/docs/master/onnx.html

Disclaimer: I work on Caffe2 team (not on ONNX, though)

Marat_Dukhan··on The .feedback scam
Interestingly, amazon.feedback redirects to amazon.com, so Amazon did pay
Marat_Dukhan··on Facebook told advertisers it can identify teens feeling 'insecure', 'worthless'
Russian social network Odnoklassniki ("classmates" in English) tried subscription-based model for two years. It nearly killed them.
Marat_Dukhan··on How Many X86-64 Instructions Are There Anyway?
I made a Python package which documents x86(-64) instructions in a ready-to-use way: https://github.com/Maratyszcza/Opcodes (also `opcodes` on PyPI). With this package, its easy to collect ISA stats, e.g.

  import opcodes.x86_64
  isa = opcodes.x86_64.read_instruction_set()
  print(sum(len(instruction.forms) for instruction in isa))
  >>> 6020
As an another example, here is number of instruction forms (e.g. mnemonic name + operand types) over time on Intel CPUs: http://imgur.com/a/AVPcq
Page 1 of 4Next →