Nice one! Let's say I'm serving local models via vllm (because ollama comes with huge performance hits), how would I implement that in gomodel?
docker run --rm -p 8080:8080 \
-e VLLM_BASE_URL=http://host.docker.internal:18000/v1 \
-e VLLM_BASEMENT_BASE_URL=http://host.docker.internal:18000/v1 \
enterpilot/gomodel:latest docker run --rm -p 8080:8080 \
-e OPENAI_API_KEY="some-vllm-key-if-needed" \
-e OPENAI_BASE_URL="http://host.docker.internal:11434/v1" \
...
enterpilot/gomodel
I'll add a more convenient way to configure it in the coming days.