90 karma · joined November 7, 2024
Tip: you can use case statements and etc. to create static queries even when you have conditionals.
Also, read https://news.ycombinator.com/newsguidelines.html#generated
The $10/mo price needed 465 people to fill a cohort before we could turn on a single GPU. People signed up and churned while waiting, so we looked at the reservation pattern and determined 80 slots was optimal. This reflects in the new price and throughput.
We're considering a 1-week option so people can test it out before committing to a full month. Would that help?
Also, please read https://news.ycombinator.com/newsguidelines.html. HN is a community for thoughtful discussion.
On TEE: yeah, it's stronger, but it also adds cost and latency. We run dedicated hardware with no prompt logging and an isolated proxy. For most people who just don't want their data in someone's training set, that's enough. If your threat model is more serious than that, we're not the right choice.
On models: we are focusing on Qwen for now. We add based on demand. Would you actually use MiMo-V2-Pro or Trinity if we had them?
Here’s what’s changed:
- We’ve removed the other LLMs for now and are focusing entirely on Qwen 3.5. We’ll bring back additional smaller models later, but most usage was already concentrated on Qwen 3.5.
- Pricing is now around $50. You get roughly 2× the throughput (61 tok/s vs. 31 tok/s, verified in testing), and it’s still unlimited. For context, that’s about 158M tokens per month. Comparable providers like Novita charge around $3.2 per million tokens, so this comes out to roughly 10% of typical token costs.
- Context size is now capped at 32K tokens. For the vast majority of use cases, this is more than sufficient.
For filling up the cohorts, I agree and we're launching for a week to gather feedback.
That said, we're planning to add a 7-day window: if a cohort doesn't fill within 7 days of your reservation, it cancels automatically and your card is released. We don't want anyone's payment method sitting in limbo indefinitely.
TTFT is under 2 seconds average. Worst case is 10-30s.