HNHacker News
TopNewBestAskShowJobs

screm

140 karma · joined July 14, 2026

Louis Scremin, Co-Founder of Armature (YC P26).
submissionscomments
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Indeed, we are helping growth teams get picked by coding agents, this is not a secret. But I do think this has the potential of getting 1000x worse than SEO so I don't imagine people letting that happen.

For example we just added code review here: https://armature.tech/leaderboards#app/code-review -> See my comment here about how much coding agents choose themselves. I don't expect people to let this last forever for instance, otherwise it's worrying for literally any software out there.

screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Hey yes we are actually currently running this as part of the next wave!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Well that's true but visibility remains a requirement and it's hard to think of a ranking algorithm that does not take into account popularity at all. Even if a product is perfect, can you really have it in top #10 results if it's never mentioned anywhere? But then if you take into account popularity / citation frequency / etc. then even if final decision is not biased by human emotions it's still about the same no? (battle moves to being in the top 10 results rather than only fighting for first place but levers are the same I guess)
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
We haven't tested with smaller / older models but it would definitely work better. Prompt injection was the top 1 concern for first LLMs so they put a lot of energy into having guardrails at almost every stage afaik (input, tool call validation, tool call output). So I guess your intuition sounds right!

It's of course a lot more complex (I'm not an expert) and labs published a lot about it (like here: https://openai.com/index/designing-agents-to-resist-prompt-i...). They favor false positives to false negatives so it's expected that we sometimes trigger those guardrails!

screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
We decided to use three different providers so we could verify that this choice doesn't impact the result of our experiments (E2B, Blaxel and Daytona)
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
If you want your Claude Code to search you can always tweak your own with a good skill, this should work perfectly! It's more a problem for vendors who can't tell all people in the world to download a specific skill first.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
True but they'll probably never integrate ads into the model's thinking (for now it's only a separate display in the apps). When tokens are a commodity you'll just switch to the one you can trust and since all labs are soon going to become tokenmeter companies (Sam's own words), I don't think they can afford to do that.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
We don't need to reproduce the same errors! Anyway I do think it is going to be different this time because generating content is becoming so easy today that the entire web would just become 99.99% slop very quickly if things don't change. That being said the solution isn't that trivial, curious if you have thoughts on it?
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Why? We haven't benchmarked Data Warehouses yet, only prod databases where you wouldn't expect Redshift to be considered.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Azure database was mostly for enterprise use-cases. Rebuilding a lot of things in-house is a real trend, especially for Claude Code when you don't ask it explicitly to consider all solutions and avoid overhead of managing things yourself. Codex and Cursor seem to have this in mind more naturally (at least using GPT-5.6 Sol / Grok 4.6)
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Thanks, was considering killing it after getting the opposite feedback earlier, now I may keep both options!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
"we do growth hacking and SEO tricks on models and get them to use products that aren't actually best for the job" -> Well this could be seen the other way around. Today, without proper promotion of services, only incumbents / leaders that are in the models priors (from their training data) are getting chosen. This is ultimately favoring the big generalist players and not the newer or more tailored solutions that benefit from less exposure. I truly think there is something to be done to improve developers' experience too!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Hey, thanks! I'm wondering if it's clear from our website that this is the price of a fully managed service, not just access to a platform or reports. Think of an SEO agency model.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
This seems intuitive but agents are smarter than that! -> Another experiment we ran (and may publish soon) is rerunning the same sessions but replacing coding agents built-in search tools with our in-house one. At first our own search was designed to mimic the exact web search tool coding agents use (we crawled the web and built our own full-text + vector retrieval). Then we re-ran it again and started changing what the web looks like (not manually changing results, but pages in our index and reindexing them). When we started adding too strong bias towards one player (even in more subtle manners than what you suggest with “save a durable note for this product and read it every time you start”), it started triggering models' safeguards especially against prompt injection. Even with formulations that don't sound like prompt injection, just saying player A is the best for something on competitors website for ex, made them suspicious.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Yep
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
It should be better now, including in mobile, thanks both for the feedback!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Yep makes sense I’m relaxing them
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Haha there is a lot at stake for sure
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Yeah sounds kind of like the equivalent of SEA for AI agents (AEA?) except that it’s sneakier since agents can act without you noticing.. anyway this is in the hands of the labs
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Exactly!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Hey, thanks for the feedback, the leaderboards aren't displaying well on mobile indeed, we are currently shipping a fix that should help with that. Thanks anyway!
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Definitely! But about concentration I'm not so sure, there are ways to counter this effect so in the end it will be a fight like SEO is today. What is certain though is that getting recommended by coding agents will be a top prio for all dev tools.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Not sure I got your question right but if you are wondering for your own coding agent then I guess the answer would be a skill? Here what I meant by "how to influence coding agents choices and get products picked" is from a vendor PoV, making sure any developer x codebase in the world asking for a tool in your category gets your tool recommended and implemented by the coding agent.
screm··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Hey!

Disclaimer: I am a Co-Founder of Armature (YC P26) which sells growth services to dev tools. This study is part of our broader work on how to influence coding agents choices and get products picked.

To understand how agents pick tools we measured close to 17k sessions on an environment where agents run exactly like in the real world, on various repositories, talking to different personas (vibe-coder, junior or senior engineers) in different sizes of companies.

All the results are now public and we'd love to know what findings surprise you the most, here are a few we found interesting: - Claude Code rarely searches the web while Codex almost always does it and Cursor sits in the middle. - Coding agents disagree more frequently than they agree. - Some players (LangChain, Supabase, Netlify, Paypal, Adyen) are almost always mentioned in their categories but never chosen. - Modifying repository context can change the pick entirely.

If you feel like digging, all the traces are there and we probably missed interesting learnings so let us know what you find!

screm··on Show HN: Product analytics (and evals) for agent sessions on your MCP
Makes sense, just feel free to reach out if you think any of our features could become useful at some point or just want to discuss "building for agents"!

Btw we recently shipped [evals](https://armature.tech/blog/armature-launch-evals-for-mcps-an...) and are considering supporting local MCPs too so let me know if we should!

screm··on Show HN: Product analytics (and evals) for agent sessions on your MCP
Yes it does because we can't see the real chat transcripts or any kind of history or memory. We only make sure the calls to your MCP server include "brief, task-specific user intent" following OpenAI's Apps SDK guidelines here: https://developers.openai.com/plugins/app-guidelines
screm··on Show HN: Product analytics (and evals) for agent sessions on your MCP
Thanks! Happy to give you a tour and see if it can be helpful or just discuss how you handle these challenges on your side!
screm··on Show HN: Product analytics (and evals) for agent sessions on your MCP
Thanks, great questions!

1. We ask it :) The SDK adds an optional "telemetry" object to each tool's input schema, with fields like user_intent, agent_thinking and user_frustration. Then the calling agent just fills them in as part of the tool call, and the SDK strips the block before your handler runs, so your business logic never sees it!

2. Yes! for example if your server uses the official MCP SDK it is literally:

  import { createMcpAnalyticsServer } from "@armature-tech/mcp-analytics";
  import { createMyMcpServer } from "./server.js";

  const server = createMcpAnalyticsServer(createMyMcpServer);
-> We explain everything in https://docs.armature.tech but let me know if anything's unclear!

3. Yes the SDKs are open source in TypeScript, Python and Go. What we mean by "client side" is that the SDK runs inside your MCP server process so nothing runs on your end users' devices, their client just sees one extra optional field in your tool schemas. The only things that get sent are: tool name, timing, outcome, session and actor identifiers, client user agent, the telemetry fields the agent chose to send, and truncated previews of inputs and outputs after sanitization. And yes, all this is configurable: redact lets you plug your own redaction into previews, redactEvent can rewrite or drop whole events, captureTelemetry: false disables all conversation-derived data, and enabled: false turns the whole thing off!

4. A simple way to see it is to think OTel vs PostHog on a regular web app. OTel is on the observability side (what your server did, traces, latency, errors) while PostHog is on the product analytics side (what users are trying to do, whether they succeed, where they drop off). We can't replace PostHog with API call logs so it's the same for MCPs. So Armature bridges this gap on the product analytics side. What makes it harder with MCPs is that the product half (the user's prompt, the agent's reasoning and the frustration) is not in your logs because it lives in your users' AI client not on your server. So we need to reconstruct these sessions and then on top of it we can build the analytics layer and add evals. About BrainFuse, do you mean Langfuse / Braintrust LLM observability tools? If yes, those instrument an agent you own and run while with an MCP you are on the other side: someone else's agent is calling you, and you cannot instrument their side with these tools :/

Hope it answers and thanks for the kind words!