HNHacker News
TopNewBestAskShowJobs

medicis123

7 karma · joined July 11, 2018

submissionscomments

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

4 pts·medicis123·
0

Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM

2 pts·medicis123·
0

A New Approach to GPU Sharing: Deterministic, SLA-Based GPU Kernel Scheduling

1 pts·medicis123·
0

Show HN: Disaggregating GPU compute from CPU in ML job execution to scale GPUs

woolyai.com·1 pts·medicis123·
0

Show HN: Run PyTorch on CPU boxes, offload kernels to remote GPUs

1 pts·medicis123·
0

Running Nvidia CUDA PyTorch container project/pipelines on AMD with no changes

1 pts·medicis123·
0

GPU-accelerated code on CPU-only environments -Remote GPU Kernel Execution

youtube.com·1 pts·medicis123·
1

Sharing base model in GPU VRAM across multiple inference stack process [video]

youtube.com·7 pts·medicis123·
1

Sharing actual GPU core and VRAM utilization metrics for query on 10 LLM models

woolyai.com·1 pts·medicis123·
1

Show HN: WoolyAI-CUDA Abstraction Layer to Decouple Kernel Shader Exec on GPU

woolyai.com·4 pts·medicis123·
0

Locally delivered and centrally managed macOS envs for privileged access setup

veertu.com·1 pts·medicis123·
0

Shopify scaling iOS CI with Anka

engineering.shopify.com·1 pts·medicis123·
0