Bringing Up DeepSeek-V4-Flash on AMD MI300X
fergusfinn.com
fergusfinn.com
I'm super curious what you would use them for.
We're kind of known for our low prices - our prices (our main usage is for our high throughput API - the async tier) is significantly below average openrouter prices - but cached prices is coming soon which will lower them even more :)
(CEO Hot Aisle)
It's not, which is why it would be nice if they did the actual work (on your hardware).
I would 100% pay $16/hr to run a self-hosted instance, but I won't spend thousands of dollars to (maybe) get it working (my time + the hardware).
https://docs.modular.com/max/models/
I agree with you though, serving up inference is secret sauce for a lot of teams and not everyone publishes how to do it because of the costs involved in doing so. They need an ROI.