HNHacker News
TopNewBestAskShowJobs

Hiteshjain118

5 karma · joined November 28, 2020

submissionscomments
Hiteshjain118··on Show HN: Fast inference for deep seek flash v4.1 469 tok/s for coding
We offer $5 free credits. But here's a $20 code for sign up from hackernews: HACKERNEWS20
Hiteshjain118··on Show HN: Fast inference for deep seek flash v4.1 469 tok/s for coding
We built a custom inference engine to deliver fast and cheap inference. We optimized this engine for coding and research workload and achieving an average speed of 469 tok/s and 4.5x lower costs than Openrouter. Software development at that speed feels different. Get an API key and try in your Opencode!
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
More power to you!
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
I think we would know the costs when they go public. Before that may have infer costs through various signals.
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
Thanks for sharing your workload. Impressive! And this one person(you) steering all this token usage across the month? I want to double click on -- if the prices were to go high, you wouldn't be spending these many tokens. Would you just delay all those projects?
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
The link in your HN is taking me to your list of Show HN posts. I wasn't able to get your github.
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
Yup, I have been testing Qwen and Kimi lately. Seeing comparable accuracy of Qwen-Instruct (not thinking) to closed source models. Here's a blog we published on that https://www.coralbricks.ai/blog/alphacumen-finance-benchmark...
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
I'm curious what your token/$ usage looks like, if you'd be willing to share :)
Hiteshjain118··on Show HN: Claude Code's $200 plan is a 17× subsidy on the raw API
My guess is you can push to $5k/month token usage with just a single Max subscription.

During my 30 days analysis window, I consumed $3371 of token and didn't hit rate limits even once.

I plan to keep pushing my token usage higher until I hit rate limits at least 5-10 times in a month.