HNHacker News
TopNewBestAskShowJobs

jrandolf

90 karma · joined November 7, 2024

submissionscomments
jrandolf··on Owner of grok.bot asks xAI for $1M
This is hilarious. Are you the owner?
jrandolf··on Show HN: sqlc-gen-sqlx, a sqlc plugin for generating sqlx Rust code
This plugin uses sqlx underneath which handles prepared statement caching. Regarding migration, we just used a coding agent to migrate our database infrastructure to it. It takes <20 minutes and remember this really only helps with static queries. We do support sqlc's dynamic queries though.

Tip: you can use case statements and etc. to create static queries even when you have conditionals.

Also, read https://news.ycombinator.com/newsguidelines.html#generated

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
We are aware of this. There was a bug that overcounted and now it's been fixed. If you'd like for us to delete your account, please contact support@sllm.cloud.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Fixed.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
First, thanks for signing up early. It means a lot.

The $10/mo price needed 465 people to fill a cohort before we could turn on a single GPU. People signed up and churned while waiting, so we looked at the reservation pattern and determined 80 slots was optimal. This reflects in the new price and throughput.

We're considering a 1-week option so people can test it out before committing to a full month. Would that help?

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
We collect emails to notify you when the cohort fills or any important information such as cancellation. No one's selling your email.

Also, please read https://news.ycombinator.com/newsguidelines.html. HN is a community for thoughtful discussion.

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
The audience here is developers buying API access. They want to see the model, the price, and the throughput, not a hero image and three paragraphs about our mission. Marketing copy between a developer and that information is friction.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
See https://news.ycombinator.com/item?id=47670843
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Yes.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
You're right that we're less flexible than OpenRouter or Chutes. We don't let you hop between models per-request. If you want that, use those. If you want predictable cost and guaranteed throughput on one model, that's us.

On TEE: yeah, it's stronger, but it also adds cost and latency. We run dedicated hardware with no prompt logging and an isolated proxy. For most people who just don't want their data in someone's training set, that's enough. If your threat model is more serious than that, we're not the right choice.

On models: we are focusing on Qwen for now. We add based on demand. Would you actually use MiMo-V2-Pro or Trinity if we had them?

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
15-25 was a rate based on oversubscription. Now it's 60 like others :).
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Thanks to everyone who shared feedback. We’re implementing it now.

Here’s what’s changed:

- We’ve removed the other LLMs for now and are focusing entirely on Qwen 3.5. We’ll bring back additional smaller models later, but most usage was already concentrated on Qwen 3.5.

- Pricing is now around $50. You get roughly 2× the throughput (61 tok/s vs. 31 tok/s, verified in testing), and it’s still unlimited. For context, that’s about 158M tokens per month. Comparable providers like Novita charge around $3.2 per million tokens, so this comes out to roughly 10% of typical token costs.

- Context size is now capped at 32K tokens. For the vast majority of use cases, this is more than sufficient.

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
You get an API key
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
The problem is different. OpenRouter is a router to LLMs. It doesn't solve GPU underutilization.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
20 tok/s is an average. It can be more, it can be less. If you are running off-peak I'm sure you'd get some crazy number.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Going on ChatGPT.com and using their AI for 24 hours doesn't mean you are actually using their LLM for 24 hours. It's only live for as long as the output is being generated. You reading, waiting for tool calls, etc. don't count toward concurrency. Factor in time-zones, lunch times, etc...it's more likely that we'd have an underutilization problem.

For filling up the cohorts, I agree and we're launching for a week to gather feedback.

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
There is vast.ai!
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Multiplexing on a GPU cloud.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
I'm feeling it Mr. Crabs.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Not if you are the only one. We have rate limits to prevent this in case, idk, you share your key with 1000 people lol.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
No cohorts have been filled yet. We're still early. We are seeing reservations pick up quickly, but I'd be able to give you a more concrete estimate of fill velocity after about a week.

That said, we're planning to add a 7-day window: if a cohort doesn't fill within 7 days of your reservation, it cancels automatically and your card is released. We don't want anyone's payment method sitting in limbo indefinitely.

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
24/7 LLM for $10/month.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Yes
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
We implement rate-limiting and queuing to ensure fairness, but if there are a massive amount of people with huge and long queries, then there will be waits. The question is whether people will do this and more often than not users will be idle.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
Thanks lol. I actually like Shadcn's style. It's sad that people view it as AI now.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
vLLM handles GPU scheduling, not sllm. The model weights stay resident in VRAM permanently so there's no loading/unloading per request. vLLM uses continuous batching, so incoming requests are dynamically added to the running batch every decode step and the GPU is always working on multiple requests simultaneously. There is no "load to VRAM and run" per request; it's more like joining an already-running batch.

TTFT is under 2 seconds average. Worst case is 10-30s.

jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
OpenRouter is a little different. We are trying to experiment with maximizing a single GPU cluster.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
1. It's an average. 2. We have sophisticated rate limiter.
jrandolf··on Show HN: sllm – Split a GPU node with other developers, unlimited tokens
That was an error on our part lol. We'll update with the price.