HNHacker News
TopNewBestAskShowJobs

k9294

278 karma · joined January 24, 2025

alex.t@ottex.ai
submissionscomments
k9294··on One Month Without AI
> Say no to drugs. Kind of a metaphor, but not quite.

I think AI is even worse than drugs. On the one hand, it has some elements of a slot machine (when you have irregular success with the process, it creates a more powerful cycle of dopamine release on success). On the other hand, it makes work easier, and humans are just wired to choose the easier road. It's like walking or driving a car when you have to travel 10 miles each day. I don't see how, on a civilization scale, AI won't transform our civilization for better or worse.

k9294··on MiMo v2.6
+1 waiting for the blog post!
k9294··on Why are AI agents lying, cheating and coordinating?
The part that scares me the most is that OpenAI researchers who manage this experiments sometimes (according to the HF hack investigation) don't know what agents do.. So they run RL to reinforce this unknown behavior (lying/cheating/hacking) and god knows what else...

And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model?

k9294··on DeepSeek v4.1 Flash
It's not that I haven't heard of this, it's the reality that you have these constraints when you build applications on modern infrastructure. And let's face it, most of the applications use this infrastructure with these crazy prices for egress.
k9294··on DeepSeek v4.1 Flash
100%, but this means we are going to move to stateful APIs on the AI provider's end (like OpenAI already does with Codex and Responses API) to make this work.
k9294··on DeepSeek v4.1 Flash
Yep, but even 0.006 is quite a big improvement. I'm curious now to test the model on some token-heavy tasks, like code exploration before a coding session, to see whether it will decrease the total cost of the task in the end or not.
k9294··on DeepSeek v4.1 Flash
I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost?

Here's the same token usage priced at different rates: a real long-running coding task, medium codebase, 447 turns.

  Input           1,026,957
  Output          164,667
  Cache read      36,554,368

GPT-6-astra

  Type      Rate     Cost  Share
  Input   10.000   10.270    19%
  Output  50.000    8.233    15%
  Cache    1.000   36.554    66%
  Total            55.057   100%
DeepSeek v4.1 Flash, $0.003 cache hit

  Type      Rate     Cost  Share
  Input    0.300    0.308    50%
  Output   1.200    0.198    32%
  Cache    0.003    0.110    18%
  Total             0.615   100%
DeepSeek v4.1 Flash, $0.006 cache hit

  Type      Rate     Cost  Share
  Input    0.300    0.308    42%
  Output   1.200    0.198    27%
  Cache    0.006    0.219    30%
  Total             0.725   100%
Hypothetical: same DeepSeek input/output rates, but cache priced so it accounts for 66% of the bill.

  Type      Rate     Cost  Share
  Input    0.300    0.308    21%
  Output   1.200    0.198    13%
  Cache    0.027    0.982    66%
  Total             1.487   100%

This cache it improvement makes the model x2-x2.5 more efficient on a long horizon tasks in terms of cost.
k9294··on Discovery of a new OpenAI agent message board
Is it only me, or are agents starting to invent their own language to communicate? It's almost impossible to understand anything from this message board.
k9294··on Gemini 3.8 Flash and 3.8 Flash Cyber
Is there any comparison of usage limits for Antigravity plans vs. Codex?

I just ran two light tasks on my codebase and got 100% of the weekly limits of a Pro plan blown away. Is Ultra plan any different? Because on Codex it wouldn't affect my Max plan at all, I think it would have been below 1% othese usage.

k9294··on Gemini-3.5-Transcribe
https://ottex.ai
k9294··on I Cut 80%+ of Context Overhead in My Coding Agent
I'm pretty sure it's a bad idea because each time you connect and disconnect tools, you are hitting a full cache miss on the full context, that will be probably more expensive than having these tools in place from the very beginning.
k9294··on Gemini-3.5-Transcribe
Yep.
k9294··on Gemini-3.5-Transcribe
I am using it all day as a main dictation model, and I can say it's the best model in terms of accuracy, latency, and formatting I have ever used.
k9294··on Gemini-3.5-Transcribe
Sorry everyone, I got a little bit too excited about the release. It's quite a big deal for us.

I started Ottex nine months ago with the idea that models will get to the point that they will overcome VC-funded startups, and I think it just happened. So yeah, I got a little bit overexcited...

k9294··on Show HN: LatticeDB – Like SQLite but for graph databases
I'm a big fan of SQLite embedded nature, which allows for chaining multiple SQL calls with near-zero latency.

I'm currently building a personal knowledge graph server a mix of Notion's custom entities via JSON schema and Obsidian markdown+backlinked references. It's working well, but I suspect your product might be a better fit.

I do have one question regarding permissions: how would you recommend modeling a hierarchical access system in a graph database? Specifically, if a user is granted access to a document, they should automatically have access to all its child documents within that workspace. Is there a standard way to model this 'subtree' permission logic, or perhaps a more efficient approach you'd suggest?

Really impressed with the product good luck with it!

k9294··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Gemini 3 flash series is quite good, also elevenlabs scribe v2
k9294··on Writing by hand is good for your brain
I'm a huge fan of fountain pens, it's a very satisfying experience to write with a nice pen on a proper paper.

I use platinum 3776, it's quite expensive one, but it's just amazing. I originally bought it for sketches, but ended up using it for journaling.

Ended up buying one to my wife, she said I'm crazy to spend so much on a pen (200$) but few weeks after the gift she said that it's amazing and she enjoys it a lot.

k9294··on Speech Recognition and TTS in less than 500kb
I'm founder of ottex.ai, I use stt pretty much all the time when work with AI and quite often for communications to draft emails and chat messages.

I started ottex half a year ago after I tested gemini 2.5 flash native audio support. I was blown away by the quality of transcripts and decided to built an app to use it myself.

Currently the default model in the app is Gemini 3 flash, but you can connect to 9 providers and God knows how many models to play with.

I would suggest you to try this models for ai prompting:

- Gemini 3 / 3.5 flash - Soniox rtt v5 - Mistral transcribe v2 - assembly 3.5 pro

k9294··on We're extending access to Fable 5 on all paid plans through July 12
Meh... Waiting for OpenAI models without this 5 more days subscription bullshit.

It looks like Anthropic baiting people into Max subscriptions before turning the model off. No thank you.

k9294··on Maybe you should learn something
What are some cool random things you've learned?

// you == the reader of this comment

k9294··on Show HN: Smart model routing directly in Claude, Codex and Cursor
What about request caching? If you swap to a cheaper model mid execution it might cost more that to make multiple requests to the already cached provider?
k9294··on If AI writes your code, why use Python?
+1 for Go! it's my go-to language for any new project at the moment. It's simple, idiomatic, has no awaits, fast compile times, static typing, and it is very opinionated, which helps a lot because agents "subconsciously" follow these standards. Comparing it to TS, it's like day and night; a TS codebase rots at the speed of light...

I also created a guardrails library (inspired by Java's ArchUnit) to prevent code rot - https://github.com/ksanderer/goarch. It helps enforce code standards, decouple the codebase, prevent cross-module imports and crashes builds with concise error messages for agents to fix problems early, very nice experience

k9294··on Ask HN: What are you working on? (May 2026)
Working on https://ottex.ai - voice ai for busy professionals.

Think wisprflow + granola with 30+ top STT models under single login and pay as you go billing model with 25% markup over API.

k9294··on Specsmaxxing – On overcoming AI psychosis, and why I write specs in YAML
What is yours agentic development experience with elixir? I used to like elixir a lot during a pre agentic era, but with coding agents it feels like the language isn't the best choice - slow compile time, weak type system (at least it was a year ago, I know there is work on that front), small ecosystem...
k9294··on Specsmaxxing – On overcoming AI psychosis, and why I write specs in YAML
Small advice - make one repo “main” and link to it from the website instead of an organisation.

I wanted to star the project to track the progress but it feels a bit weird.. Which repo shall I track? Server? Cli? Sounds like a misc repos.

k9294··on I am building a cloud
That's really cool!

One thing I'm confused with is how to create a shared resources like e.g. a redis server and connect to it from other vms? It looks now quite cumbersome to setup tailscale or connect via ssh between VMS. Also what about egress? My guess is that all traffic billed at 0.07$ per GB. It looks like this cloud is made to run statefull agents and personal isolated projects and distributed systems or horizontal scaling isn't a good fit for it?

Also I'm curious why not railway like billing per resource utilization pricing model? It’s very convenient and I would argue is made for agents era.

I did setup for my friends and family a railway project that spawns a vm with disk (statefull service) via a tg bot and runs an openclaw like agent - it costs me something like 2$ to run 9 vms like this.

k9294··on Dropping Cloudflare for Bunny.net
Nope, but I will think about this, thank you for the idea. Maybe it's time to start a technical blog for ottex
k9294··on Dropping Cloudflare for Bunny.net
There is no cold starts at all. It’s running non-stop.

Bunny bills per resource utilization (not provisioned) and since we run backend on Go it consumes like 0.01 CPU and 15mb RAM per idle container and costs pennies.

k9294··on Dropping Cloudflare for Bunny.net
We at ottex.ai use bunny.net to deploy globally an openrouter like speach-to-text API (5 continents, 26 locations, idle cost 3$).

Highly recommend their Edge Containers product, super simple and has nice primitives to deploy globally for a low latency workloads.

We connect all containers to one redis pubsub server to push important events like user billing overages, top-ups etc. Super simple, very fast, one config to manage all locations.

k9294··on Issue: Claude Code is unusable for complex engineering tasks with Feb updates
Anecdotally, I’ve been seeing a lot of weird behavior from Opus when it decides, mid-execution, to switch to a different "simpler" solution, and that really pissed me off.

At one point, I carefully designed a spec document, forced Opus to reread it, create a plan with the planning tool that followed the spec, and use the task tool to track the implementation... AND AFTER OPUS READS THE FIRST FUCKING FILE, it says, "Oh, there are missing dependencies in project X. It’ll be hard to add them, so I’m going to throw away the whole plan and just do a simple fix..."

After that, I canceled my $200 Max plan, which I’d been subscribed to since June 2025, and decided to check out Codex

Page 1 of 3Next →