HNHacker News
TopNewBestAskShowJobs

RussianCow

4,165 karma · joined May 26, 2013

Just a guy who loves tech but hates the industry. I regularly alternate between giddy excitement and wanting to light my computer on fire.

My current venture is Semi-Decent, a software consultancy that helps businesses grow through intensely pragmatic software. https://www.semi-decent.com/

I also run a side business making bespoke, sustainable craft cocktails for events: https://www.theminimalmixologist.com/

Please get in touch! I'd love to chat.

sasha+hn@chedygov.com

submissionscomments
RussianCow··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
> You don't need SOTA. You need a model that can accomplish your task.

I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.

> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.

Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.

RussianCow··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
> it sounds like a fumble by OAI.

The vast majority of their revenue comes from large businesses buying for their teams, which are almost certainly not going to juggle lower tiers of different subscriptions to save a few bucks.

RussianCow··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
People keep saying this but it's just patently not true, or at least not apples-to-apples. You can't seriously compare Qwen 3.8 27B to Fable or Astra. Even if local models get better, so will the frontier, and you'll always be at a disadvantage.

Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.

There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.

RussianCow··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I'm a software engineer and I'm in a similar boat. The models have definitely gotten better, but all the latest frontier models have been within the same order of magnitude of usefulness for many months now. Opus 4.5 was a huge boon in productivity, and I certainly write less code by hand than I did when it came out, but I can't say that my workflow has changed drastically for many months. Every model still takes some amount of babysitting to ensure it's doing the right thing, and they all make silly mistakes sometimes.
RussianCow··on Sonnet 5.5
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.

With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.

RussianCow··on Toyota is taking the Corolla electric
You're right, I should have said "in the US". But it is the second biggest car market in the world, so it's not like we're irrelevant in the discussion.
RussianCow··on Unreal Agent
Like what?
RussianCow··on Toyota is taking the Corolla electric
Seems like a chicken-and-egg problem. Part of the reason consumers don't want EVs is because there are limited options that are any good.
RussianCow··on Unreal Agent
I don't mean to be super negative, but I read through this all the way and still have no idea what it does or how it works. I think the problem statement needs to be outlined much better, and concrete examples of turns with/without your solution would be helpful to illustrate the difference.
RussianCow··on Mercury 2.5 LLM hits 770 tokens per second
Depends on how many times you need to iterate to get the result you want. If you need to run the fast model 5 times to get the results you need, compared to 1-2 times for a smarter but slower model, you've just eroded any advantage that the speed gave you.

So, as always: it depends on the use case.

RussianCow··on Mercury 2.5 LLM hits 770 tokens per second
But there are lots of use cases where a relatively "dumb" model is good enough.
RussianCow··on Mercury 2.5 LLM hits 770 tokens per second
Unfortunately, the lack of an input cache discount makes it prohibitively expensive for most use cases that aren't one-shot prompts.
RussianCow··on Mercury 2.5 LLM hits 770 tokens per second
The point is the speed.
RussianCow··on How did AMD Ryzen get 50% faster in two years?
> It is, but isn’t it all warranty in your scenario?

There are numerous recent stories of companies refusing to make good on their warranties and instead offering customers a refund for the original price paid. One example: https://www.tomshardware.com/pc-components/hdds/toshiba-refu...

RussianCow··on MiMo v2.6
Because the end goal is to ban non-US AI companies from being able to do business in the US.
RussianCow··on Grok 4.7
For caching, only if you don't specify your preferred providers and let OpenRouter route each request itself. I have stuff like this in my OpenCode config for each model I use and I regularly get ~90-95% cache hit rates.

    "order": ["relace", "coreweave", "novita", "baseten", "together"],
    "allow_fallbacks": false
It still won't be quite as high as you'd get by just using DeepSeek because occasionally a request will fail and you'll get routed to a backup provider with nothing cached, but it's close enough not to matter in most instances.

But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).

RussianCow··on Astra for Law
Honestly, I know several devs who do chores around the house or even play video games while AI does the bulk of the heavy lifting. If nobody cares or even realizes, does it matter? (To be clear, they all work remotely.)
RussianCow··on Nvidia dismisses "circular financing", says every $1 it invests brings back $100
Which doesn't matter if they're still not profitable.
RussianCow··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?
RussianCow··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
The issue is that all input (including context) counts towards that limit. So 10 requests with 50k of context will blow through the limit, even if little to no output was generated, which is incredibly easy to do with agentic workloads.
RussianCow··on Claude Fable 5.1 and Claude Mythos 5.1
This is pretty terrible advice when there are dozens of AI inference providers out there serving great models with significantly more cost effectiveness than you'd get from buying your own hardware.
RussianCow··on Can I opt out of my input or output data being used for training?
I don't trust any company with their word on anything. Luckily, privacy policies are legally binding.
RussianCow··on Hy4 preview
Presumably the number that OpenRouter shows is averaged across all requests.
RussianCow··on Hy4 preview
Most providers do what's called "prefix caching", where each turn in a session is cached such that sending new messages with the exact same "prefix" (set of previous messages) gives you the cache read price on that input instead of the full price. As long as you're not changing your system prompt, available tools, etc mid-session, you automatically benefit from this.
RussianCow··on Hy4 preview
They're not "docking points", they're calculating it in the most straightforward way. If I start a session and the majority of requests are sent to Provider A, and my last request gets routed to Provider B, I have a 0% cache hit rate with Provider B. I'm very curious how else you expect this to be calculated? Do you think they're completely omitting requests that switch providers mid-session?

FWIW, I get significantly higher than listed cache hit rates when I pin my session to a specific provider, which is further evidence of the above.

RussianCow··on Claude Session URL appended to commit messages and PR descriptions by default
But presumably everyone in your company/team is using Jira, so it's not an "ad" because it's a product already used internally. Claude is appending these links to all commits by default, whether or not others on the team use Claude. Those are very different things.
RussianCow··on Hy4 preview
How else would you expect them to calculate it?
RussianCow··on Hy4 preview
This is very much NOT my experience in practice, even though it's how I would expect it to work. OpenRouter will happily bounce you between several providers (none of which have downtime) even within the same session. Requesting specific providers is the only way I've been able to hit a cache rate above 90%.
RussianCow··on Boot a Virtual iPhone via Apple's Virtualization.framework
Can you give some examples of where it matters? I'm genuinely curious.
RussianCow··on Boot a Virtual iPhone via Apple's Virtualization.framework
That explains the difference, but what's the purpose?
Page 1 of 34Next →