HNHacker News
TopNewBestAskShowJobs

numlocked

3,680 karma · joined August 17, 2009

Chris Clark.

COO, OpenRouter (https://openrouter.ai)

Previously: Co-founder & CTO Grove (www.grove.co, NYSE:$GROV)

Creator of SQL Explorer (www.sqlexplorer.io)

[ my public key: https://keybase.io/cc; my proof: https://keybase.io/cc/sigs/mvtne8Fa_G2XaRwEkDFpAifq6DYrB5PY5rpj9-RHZ4A ]

submissionscomments
numlocked··on Prompt-caching – auto-injects Anthropic cache breakpoints (90% token savings)
As per its own FAQ this plugin is out of date and doesn’t actually do anything incremental re:caching:

> "Hasn't Anthropic's new auto-caching feature solved this?"

> Largely, yes — Anthropic's automatic caching (passing "cache_control": {"type": "ephemeral"} at the top level) handles breakpoint placement automatically now. This plugin predates that feature and originally filled that gap.

numlocked··on OpenAI API Logs: Unpatched data exfiltration
At the risk of totally misunderstanding this...it seems to be exfiltration by the app developer, who already has access to all of these data sources and the data that the customer is inputting into the AI KYC app (in this example)...right? I don't believe this exposes any end-user information to a third party. The AI app developer is already 'trusted' and could get access to this information regardless of the exfiltration. Maybe someone can explain this to me more clearly.
numlocked··on Response Healing: Reduce JSON defects by 80%+
That is a way of doing that, but it's quite expensive computationally. There are some companies that can make it feasible [0], but it's often not a perfect process and different inference providers implement it different ways.

[0] https://dottxt.ai/

numlocked··on Mistral OCR 3
(I work at OpenRouter) If you send a PDF to our API we will:

1. Use native PDF parsing if the model supports it

2. Use this Mistral OCR model (we updated to this version yesterday)

3. UNLESS you override the "engine" param to use an alternate. We support a JS-based (non-LLM) parser as well [0]

So yes, in practice a lot of OCR jobs go to Mistral, but not all of them.

Would love to hear requests for other parsers if folks have them!

[0] https://openrouter.ai/docs/guides/overview/multimodal/pdfs#p...

numlocked··on Classical statues were not painted horribly
I just learned that the site/magazine publishing this, Works in Progress, is owned by Stripe! I have no idea why, but the content is great so...thanks Stripe!
numlocked··on Classical statues were not painted horribly
Good read! The idea that these marvels of artistry were painted like my 10th birthday at the local paint-your-own-pottery store always seemed incongruous, at best.

> Why, then, are the reconstructions so ugly?

> ...may be that they are hampered by conservation doctrines that forbid including any feature in a reconstruction for which there is no direct archaeological evidence. Since underlayers are generally the only element of which traces survive, such doctrines lead to all-underlayer reconstructions, with the overlayers that were obviously originally present excluded for lack of evidence.

That seems plausible -- and somewhat reasonable! To the credit of academics, they seems aware of this (according to the article):

> ‘reconstructions can be difficult to explain to the public – that these are not exact copies, that we can never know exactly how they looked’.

numlocked··on Developers can now submit apps to ChatGPT
We do this at openrouter and many apps use exactly that pattern!
numlocked··on Getting a Gemini API key is an exercise in frustration
(I work at OpenRouter) Certainly for individual developers / hobby projects that's the primary value prop; super easy access to all of the models.

But there's a lot more functionality that becomes relevant when building in production. We do automatic fallbacks, route between providers based on data policies, syndicate your data to agent observability tools / your logging platform of choice, user-level and api-key-level budget management and model allow/block lists, programmatic API key management, etc, etc. More good stuff shipping all the time!

numlocked··on Getting a Gemini API key is an exercise in frustration
(I work at OpenRouter) We add about 15ms of latency once the cache is warm (e.g. on subsequent requests) -- and if there are reliability problems, please let us know! OpenRouter should be more reliable as we will load balance and fall back between different Gemini endpoints.
numlocked··on GLM 4.5 with Claude Code
(OpenRouter COO here) We are starting to test this and verify the deployments. More to come on that front -- but long story short is that we don't have good evidence that providers are doing weird stuff that materially affects model accuracy. If you have data points to the contrary, we would love them.

We are heavily incentivized to prioritize/make transparent high-quality inference and have no incentive to offer quantized/poorly-performing alternatives. We certainly hear plenty of anecdotal reports like this, but when we dig in we generally don't see it.

An exception is when a model is first released -- for example this terrific work by artificial analysis: https://x.com/ArtificialAnlys/status/1955102409044398415

It does take providers time to learn how to run the models in a high quality way; my expectation is that the difference in quality will be (or already is) minimal over time. The large variance in that case was because GPT OSS had only been out for a couple of weeks.

For well-established models, our (admittedly limited) testing has not revealed much variance between providers in terms of quality. There is some but it's not like we see a couple of providers 'cheating' by secretly quantizing and clearly serving less intelligence versions of the model. We're going to get more systematic about it though and perhaps will uncover some surprises.

numlocked··on How big are our embeddings now and why?
I hear you, but the article is talking specifically about "embeddings as a product" -- not the embeddings that are within an LLM architecture. It starts:

> As a quick review, embeddings are compressed numerical representations of a variety of features (text, images, audio) that we can use for machine learning tasks like search, recommendations, RAG, and classification.

Current standalone embedding models are not intrinsically connected to SotA LLM architectures (e.g. the Qwen reference) -- right? The article seems to mix the two ideas together.

numlocked··on How big are our embeddings now and why?
I don’t quite understand. The article says things like:

“With the constant upward pressure on embedding sizes not limited by having to train models in-house, it’s not clear where we’ll slow down: Qwen-3, along with many others is already at 4096”

But aren’t embedding models separate from the LLMs? The size of attention heads in LLMs etc isn’t inherently connected to how a lab might train and release an embedding model. I don’t really understand why growth in LLM size fundamentally puts upward pressure on embedding size as they are not intrinsically connected.

numlocked··on OpenRouter is down
Hi folks -- I'm Chris from OpenRouter. This one hurts. We're back, but our database was down for about 45 minutes, which caused user and credit lookups to fail, and took down the API. We are investigating why, and of course going to look into improving durability so this failure mode can't happen again. We will share a post-mortem on the site when we have finished our investigation. I'm sorry to our users who count on us.
numlocked··on Show HN: Price Per Token – LLM API Pricing Data
(I work at OpenRouter)

We have a simple model comparison tool that is not-at-all-obvious to find on the website, but hopefully can help somewhat. E.g.

https://openrouter.ai/compare/qwen/qwen3-coder/moonshotai/ki...

numlocked··on Show HN: Price Per Token – LLM API Pricing Data
(I work at OpenRouter)

We have solved this problem by working with the providers to implement a prices and models API that we scrape, which is how we keep our marketplace up to date. It's been a journey; a year ago it was all happening through conversations in shared Slack channels!

The pricing landscape has become more complex as providers have introduced e.g. different prices for tokens depending on prompt length, caching, etc.

I do believe the right lens on this is actually the price per token by endpoint, not by model; there are fast/slow versions, thinking/non-thinking, etc. that can sometimes also vary by price.

The point of this comment is not to self promote, but we have put a huge amount of work into figuring all of this out, and have it all publicly available on OpenRouter (admittedly not in such a compact, pricing-focused format though!)

numlocked··on OpenAI dropped the price of o3 by 80%
Hi - I'm the COO of OpenRouter. In practice we don't expire the credits, but have to reserve the right to, or else we have a uncapped liability literally forever. Can't operate that way :) Everyone who issues credits on a platform has to have some way of expiring them. It's not a profit center for us, or part of our P&L; just a protection we have to have.
numlocked··on Claude 4
(Caveat on this comment: I'm the COO of OpenRouter. I'm not here to plug my employer; just ran across this and think this suggestion is helpful)

Feel free to give OpenRouter a try; part of the value prop is that you purchase credits and they are fungible across whatever models & providers you want. We just got Sonnet 4 live. We have a chatroom on the website, that simply uses the API under the covers (and deducts credits). Don't have passkeys yet, but a good handful of auth methods that hopefully work.

numlocked··on 4o Image Generation
Thank you for the kind words and email received!
numlocked··on 4o Image Generation
I’m a bit late here - but I’m the COO of OpenRouter and would love to help out with some additional credits and share the project. It’s very cool and more people could be able to check it out. Send me a note. My email is cc at OpenRouter.ai
numlocked··on Qwen2.5-VL-32B: Smarter and Lighter
We don’t particularly want our customers’ data :)
numlocked··on Qwen2.5-VL-32B: Smarter and Lighter
COO of OpenRouter here. We are simply stating the WE can’t vouch for the behavior of the upstream provider’s retention and training policy. We don’t save your prompt data, regardless of the model you use, unless you explicitly opt-in to logging (in exchange for a 1% inference discount).
numlocked··on The model is the product
Is this the Will Brown talk you are referencing? https://www.youtube.com/watch?v=JIsgyk0Paic
numlocked··on Boris Spassky: 1937–2025
Ohh, that is lovely. I had not seen it. Black does seemingly nothing wrong and is just in a world of hurt by the 12th move.
numlocked··on I tasted Honda’s spicy rodent-repelling tape and I will do it again (2021)
I have no insight but boy oh boy is this funny and well written. Like prime Dave Barry [0].

[0] https://www.davebarry.com/columns/how-to-make-board.php

numlocked··on Show HN: A Marble Madness-inspired WebGL game we built for Netlify
Minor -- the date on the Riot Games video "stop point" is likely wrong. It says 2024, but based on the timeline it should be 2023.
numlocked··on Thinking about recipe formats more than anyone should
After reading another comment here about a recipe being an “upside down tree” I now understand what at least this format is trying to accomplish.

It has some really nice properties, but trades off a key feature of the gantt format: your hands can only be doing one thing at a time. With the gantt format it’s very clear what you are supposed to be doing at any time and it preserves the order of operations. It doesn’t express how things are combined, however, which the tree format accomplishes.

My motivation for the gantt format was to prevent getting “meanwhiled” by a recipe. You are chugging along, and think you are in good shape, and come across that dastardly word in a recipe: Meanwhile. Turns out you should have beaten the eggs to a stiff whip 15 minutes ago.

numlocked··on Thinking about recipe formats more than anyone should
Hmm interesting. But I have to confess I have no idea, intuitively, how to read that format. I’m sure it works once you understand it, but if you need an instruction manual for the format then maybe you’ve lost the plot a bit.
numlocked··on Thinking about recipe formats more than anyone should
Many years ago I experimented with making recipes into Gantt charts. For more complex recipes this proved incredibly useful. I spent some time trying to automate turning some of the recipe formats into Gantts, but it was pretty cumbersome. I'll bet a good LLM would make this achievable now.

For an example, here's a gantt chart for Beef Bourguignon:

https://ibb.co/c3TVTnX

Note that when I print it on a (physical) recipe card, I have the 'prose' instructions underneath.

I still think this is a pretty good idea, and I still use the cards for this recipe, and Beef Wellington.

numlocked··on Ask HN: Why is Pave legal?
I believe wage-fixing would require companies to agree to…fix wages. Having knowledge of average compensation is not inherently problematic. Firms are still able to decide to pay more or less than the prevailing wage.
numlocked··on Show HN: SQL Explorer – Open-source reporting tool that Just Works
Amazing! Thanks so much!
← PreviousPage 2 of 14Next →