HNHacker News
TopNewBestAskShowJobs

medicis123

7 karma · joined July 11, 2018

submissionscomments
medicis123··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
We did something similar - Streaming experts. Maintaining an expert cache, optimizing it to simulate running a multi-model agentic workflow on a 2-DGC Spark Cluster. The models we ran were: DeepSeek V4 Flash, Gemma 4 26B A4B, and Nemotron 3 Nano Omni 30B NVFP4. The results were very encouraging in terms of performance and model switching. Check it out here - https://woolyai.com/ai-compute-software/dgx-spark-inference-...
medicis123··on GPU-accelerated code on CPU-only environments -Remote GPU Kernel Execution
We built this feature in our GPU Hypervisor, which cleanly separates your user-space ML environment from the GPU runtime, so you can code locally, run remotely, and execute kernels remotely on GPU hosts. Would love to hear how this impacts your ML platforms. Pls comment or DM me.

https://youtu.be/f62s2ORe9H8

medicis123··on Sharing base model in GPU VRAM across multiple inference stack process [video]
We have just published a short demo of the WoolyAI GPU Hypervisor, showcasing VRAM memory sharing/deduplication. Load a single base model once, then run multiple isolated LoRA stacks or VLLM stacks on the same GPU.

Why this matters

Higher capacity: Share the base model in VRAM; add more adapters or vertical inference stacks per GPU without increasing memory usage.

Isolation & control: Each stack is its own process with independent batching and SLA-aware scheduling.

While vLLM supports multiple adapters on a single vLLM process, many teams need predictable per-adapter SLAs—this is where running independent stacks with a shared base model in VRAM can enable doing it all on the same GPU.

The demo uses LoRA inference using Pytorch, but the same applies when using vLLM. If you’re scaling LoRA inference across business units or model variants and need predictable latency without overprovisioning GPUs, I’d love your feedback. Comment or DM to chat.

medicis123··on Show HN: Ultra private AI meeting assistant
Looks interesting. Will give it a try. I have been following AI startups developing various agents(Agentic phenomenon). Are you using OPenAi and any other LLM service APIs or you fine-tuned an open source model?
medicis123··on Sharing actual GPU core and VRAM utilization metrics for query on 10 LLM models
We ran it on WoolyAI Acceleration Service https://docs.woolyai.com/getting-started/running-your-first-...

There are some interesting insights just looking at these numbers.

Environment Details Wooly Client: Linux non-GPU container running PyTorch scripts for all ten models Models were downloaded using Hugging Face Transformers library from vendor-specific repositories. Each model was executed 20 times using the same script to collect average Wooly Credits for both CPU and VRAM usage. Models Tested Llama-3.2-1B Llama-3.2-1B-Instruct Llama-3.2-3B Llama-3.2-3B-Instruct Mistral-7B-Instruct Falcon3-7B-Instruct Llama-3.1-8B-Instruct Llama-3.1-8B Dolly-v2-12B Llama-2-13B-Chat-HF Pytorch Script

from transformers import AutoTokenizer, AutoModelForCausalLM import torch torch.manual_seed(100000) # Model name or path model_name = "meta-llama/Meta-Llama-3.1-8B-Instruct" # Load tokenizer and model tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name, device_map="cuda") # Input text input_text = "What is the capital of United States of America" # Tokenize input text inputs = tokenizer(input_text, return_tensors="pt").to(model.device) # Decode and print output for z in range (1, 10): outputs = model.generate(*inputs, max_length=100) generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True) print(generated_text) GPU Core and Memory utilization Metrics

Llama-3.2-1B Core Wooly Credits Used - 46000 VRAM Wooly Credits Used - 31072 Llama-3.2-1B-Instruct Core Wooly Credits Used - 94868 VRAM Wooly Credits Used - 60964 Llama-3.2-3B Core Wooly Credits Used - 195936 VRAM Wooly Credits Used - 84715 Llama-3.2-3B-Instruct Core Wooly Credits Used - 502448 VRAM Wooly Credits Used - 258125 Mistral-7b-instr Core Wooly Credits Used - 525689 VRAM Wooly Credits Used - 397181 Falcon3-7B-Instruc Core Wooly Credits Used - 136094 VRAM Wooly Credits Used - 26528 Llama-3.1-8B Core Wooly Credits Used - 283458 VRAM Wooly Credits Used - 167515 Llama-3.1-8B-Instruct Core Wooly Credits Used - 574872 VRAM Wooly Credits Used - 403934 Dolly-v2-12b Core Wooly Credits Used - 767108 VRAM Wooly Credits Used - 342877 Llama-2-13b-chat-hf Core Wooly Credits Used - 313809 VRAM Wooly Credits Used -120067

medicis123··on RemoteMac.io – Dedicated Mac mini
Wanted to share that we developed a GitLab CI Runner/executor for Anka Build and have made it public (It was done for one of our user). It basically enables you to run your iOS/macOS GitLab CI pipelines/jobs on Anka build macOS cloud. You can configure Anka Build macOS cloud on-prem(on macs) or on hosted macs. Happy to share more details.https://github.com/veertuinc/gitlab-runner