HNHacker News
TopNewBestAskShowJobs

crowwork

204 karma · joined September 30, 2016

submissionscomments
crowwork··on Open ABI and FFI for Machine Learning Systems
The goal of the project is to bring open ABI and FFI for machine learning systems.

- Stable, minimal C ABI designed for kernels, DSLs, and runtime extensibility. - Zero-copy interop across PyTorch, JAX, and CuPy using DLPack protocol. - Compact value and call convention covering common data types for ultra low-overhead ML applications. - Multi-language support out of the box: Python, C++, and Rust (with a path towards more languages).

crowwork··on LLM Microserving: a new RISC-style approach to design LLM serving API
Scale LLM serving with programmable cross-engine serving patterns, all in a few lines of Python
crowwork··on [dead]
XGrammar is an open-source library for efficient, flexible, and portable structured generation. Bring 2x-10x speedup in grammar grammar-guided(JSON and CFG) LLM serving.
crowwork··on In-browser LLM inference engine with WebGPU and OpenAI API
Comes with ability to do full structured generation with json schema

also a in-browser demo https://chat.webllm.ai/

crowwork··on Universal LLM Deployment Engine with ML Compilation
runs on qwen2 on iphone with 26 tok/sec and a OpenAI style swift API
crowwork··on Gemma locally on iOS, Android, web browsers, and GPUs with a single framework
2b model running at 20tok/sec on iphone, nice potential for future applications
crowwork··on Running LLM (phi-2) locally on latest Google Chrome Android
Runs Phi-2 on Samsung S23 with pretty decent speed on Google Chrome browser.

LLM on browser on a phone

crowwork··on Making AMD GPUs competitive for LLM inference
You can also try out the vulkan backend, which we know should work for windows, although speed might be slower than rocm
crowwork··on Making AMD GPUs competitive for LLM inference
Yes, it works out of box and the blog contains a prebuilt python package that you can try out
crowwork··on Making AMD GPUs competitive for LLM inference
There is also vulkan support which should be more universal(also included in the post), for example, the post also shows running LLM on a steamdeck APU.
crowwork··on Run Llama2-70B in Web Browser with WebGPU Acceleration
Checkout the latest docs https://mlc.ai/mlc-llm/docs/ MLC started with demos and it evolved lately, with API integrations, documentations into an inference solution that everyone can reuse for universal deployments
crowwork··on What Is ML Compilation
It certainly also involves generating code(e.g. WebGPU, vulkan) that are more akin to traditionally compiler, and more like graph and memory optimization. So indeed more than packaging.

Please checkout the course if you are interested

crowwork··on MLC-LLM: GPT/Llama on consumer-class GPUs and phones
There is a conda app that can be installed on macos
crowwork··on MLC-LLM: GPT/Llama on consumer-class GPUs and phones
You can try out the demo and benchmark yourself
crowwork··on A brief history of LLaMA models
https://mlc.ai/mlc-llm/
crowwork··on MLC LLM: Universal LLM Deployment with GPU Acceleration
Supported platforms include:

- Metal GPUs on iPhone and Intel/ARM MacBooks

- AMD and NVIDIA GPUs via Vulkan on Windows and Linux

- NVIDIA GPUs via CUDA on Windows and Linux

- WebGPU on browsers (through companion project WebLLM).

crowwork··on Web LLM – WebGPU Powered Inference of Large Language Models
tvm runtime is pretty decent(~700k-2M level depending on dependency included), you can checkout tvm community and bring up the question there, i think there might be some common interest. There are impl of runtime for vulkan, metal that can be used as reference.
crowwork··on Web LLM – WebGPU Powered Inference of Large Language Models
I think instead what would be needed is a wgpu native runtime support for TVM. Like the implementations in tvm vulkan, then it will be naturally link to any runtime that provides webgpu.h

Then yah the llm_chat.js would be high-level logic that targets the tvm runtime, and can be implemented in any language that tvm runtime support(that includes, js, java, c++ rust etc).

Support webgpu native is an interesting direction. Feel free to open a thread in tvm discuss forum and perhaps there would be fun things to collaborate in OSS

crowwork··on Web LLM – WebGPU Powered Inference of Large Language Models
The WGSL are generated and compiled through TVM and embedded into the wasm.

I think what you mean is wgpu native support. At the moment the web gpu runtime dispatches to the js webgpu environment. Once TVM runtime comes with wgpu native support (like the current ones in vulkan or metal), then it is possible to leverage any wgpu native runtime like what Zig provide.

Additionally, currently tvm natively support targets like vulkan, metal directly which allows targeting these other platforms

crowwork··on Web Stable Diffusion
Webgpu will ship this year, so it will be more widely available pretty soon
crowwork··on Transformers.js
checkout https://mlc.ai/web-stable-diffusion, which is builds on top of Apache TVM and brings in models from PyTorch2.0, ONNX and other means into the ML compilation flow
crowwork··on Running Stable Diffusion fully in browser with WebGPU
FP16 is already in the spec and hopefully will ship this year we believe
crowwork··on Running Stable Diffusion fully in browser with WebGPU
On the current setup(M2 and metal), it should run as fast as the local native environment.
crowwork··on Running Stable Diffusion fully in browser with WebGPU
The current demo uses CLIP model from openai, which is likely what you are looking for
crowwork··on Running Stable Diffusion fully in browser with WebGPU
on apple M2max, it takes 20sec, and with the fp16 support, there will be opportunities for improvement likely soon this eyar
crowwork··on Running Stable Diffusion fully in browser with WebGPU
This project brings stable diffusion models to web browsers. Everything runs inside the browser with no server support. Please check out our GitHub repo to see how we did it. There is also a demo which you can try out.
crowwork··on Integrating the TVM Deep Learning Compiler into PyTorch
As TVM continuously demonstrates improvements to the efficiency of deep learning execution, it has become clear that PyTorch stands to benefit from directly leveraging the compiler stack. A major tenet of PyTorch is providing seamless and robust integrations that don’t get in the user’s way. To that end, PyTorch now has an official TVM-based backend, torch_tvm.
crowwork··on Automating Optimization of Quantized Deep Learning Models on CUDA
With learning-based program optimizer, we can competitive performance on benchmark models and significant boost on emerging models against TensorRT(int8).
crowwork··on TVM deep-learning compiler framework transitions to Apache
“TVM is right for the Apache Software Foundation, and the Apache Software Foundation is right for TVM: One thing the ASF excels at is enabling collaboration across organizations, and encouraging collaboration even among competitors. With contributions from such a wide range of organizations, TVM clearly fits that profile. I am honored to help the project thrive in the ASF,” said Markus Weimer, the ASF member who championed the incubation of TVM at the ASF.
crowwork··on Talks from the first TVM and deep learning compilers conference
TVM is an open-source deep learning compiler stack for CPUs, GPUs, and specialized accelerators. It aims to close the gap between the productivity-focused deep learning frameworks, and the performance- or efficiency-oriented hardware backends.

The conference contained 20+ talks covering deep learning compilation, specialized accelerators, IoT, ML for systems, privacy/security and more

Page 1 of 2Next →