67% Cost Savings with PD Disaggregation Using Ray and vLLM on AMD MI325Xanyscale.com·4 pts·robertnishihara·0
Major upgrades to Ray Serve: 88% lower latency and 11.1x higher throughputanyscale.com·2 pts·robertnishihara·1
vLLM large scale serving: DeepSeek 2.2k tok/s/h200 with wide-epblog.vllm.ai·147 pts·robertnishihara·54
An Open Source Stack for AI Compute: Kubernetes and Ray and PyTorch and VLLManyscale.com·1 pts·robertnishihara·0
AsyncFlow: An Asynchronous Streaming RL Framework for LLM Post-Trainingarxiv.org·4 pts·robertnishihara·0
Large-Scale Deployment of Ray in Tencent's Weixin AI Infrastructureanyscale.com·2 pts·robertnishihara·0
An Open Source Stack for AI Compute: Kubernetes and Ray and PyTorch and VLLManyscale.com·1 pts·robertnishihara·0
Building an LLM Router for High-Quality and Cost-Effective Responsesanyscale.com·1 pts·robertnishihara·0
RAG at Scale: 10x Cheaper Embedding Computations with Anyscale and Pineconeanyscale.com·1 pts·robertnishihara·0
Comparing LLM Performance: Introducing the Open Source Leaderboard for LLM APIsanyscale.com·2 pts·robertnishihara·0