HNHacker News
TopNewBestAskShowJobs

colorant

150 karma · joined October 31, 2018

submissionscomments
colorant··on FlashMoE: DeepSeek-R1 671B and Qwen3MoE 235B with 1~2 Intel B580 GPU in IPEX-LLM
Hardware Requirements:

- 380GB CPU memory for DeepSeek V3/R1 671B INT4 model

- 128GB CPU memory for Qwen3MoE 235B INT4 model

- 1-2 ARC A770 or B580

- 500GB Disk space

colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
With ~1000 input, the TTFT is ~10 seconds
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
The ipex-llm implementation extends llama.cpp and includes additonal CPU-GPU hybrid optimizations for sparse MoE
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
Currently >8 token/s; there is a demo in this post: https://www.linkedin.com/posts/jasondai_run-671b-deepseek-r1...
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
Yes, but the context length will be limited due to VRAM constraint
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
Prompt length mainly impacts prefill latency (FTFF), not the decoding speed (TPOT)
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
This is based on llama.cpp
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
>8TPS at this moment on a 2-socket 5th Xeon (EMR)
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
See this section https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
Also see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex...
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
Yes, you are right. Unfortunately HN somehow truncated my original URL link.
colorant··on DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...

Requirements (>8 token/s):

380GB CPU Memory

1-8 ARC A770

500GB Disk

colorant··on IPEX-LLM Portable Zip for Ollama on Intel GPU
llamafile cannot use Intel GPU (including integrated GPU on your PC)
colorant··on Llama.cpp supports Vulkan. why doesn't Ollama?
There is ipex-llm support for Ollama on Intel GPU (https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...)
colorant··on BigDL-LLM: running LLM on your laptop using INT4
BigDL-LLM is a library for running LLM (language language model) on your local laptop using INT4 with very low latency on CPU. (It is built on top of the excellent work of llama.cpp, gptq, bitsandbytes, etc., and supports any Hugging Face Transformers model).
colorant··on Show HN: Analytics Zoo – Distributed TensorFlow and Keras on Apache Spark
- Data wrangling and analysis using PySpark

- Deep learning model development using TensorFlow or Keras

- Distributed training/inference on Spark and BigDL

- All within a single unified pipeline and in a user-transparent fashion!