HNHacker News
TopNewBestAskShowJobs

pranshuchittora

95 karma · joined August 13, 2018

submissionscomments
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Yes a lot. So with the acceleration of AI in the software engineering. Features are being shipped faster but causes regressions. The only way to verify is either you write tests with AI and spend hours reviewing them or you do manual QA. agent-qa aims to solve the later. First your product should work for the end user, later you can write clean test etc.
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Try the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Try the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Try the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Try the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
I did. Checkout the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Try the OSS alternative - https://github.com/vostride/agent-qa
pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Some digging FAST_MODEL = "google/gemini-3-flash" (fast mode primary) DEEP_MODEL = "openai/gpt-5.4" (deep mode primary) VISION_CLICK_MODEL= "openai/gpt-5.4" (the visual grounder)

fast: gemini-3-flash, falls back to gpt-5.4, 15-min run timeout, max 2 visual calls/step. deep: gpt-5.4, 15-min timeout, max 3 visual calls/step.

Why such a hard timeout, and why not latest models?

pranshuchittora··on Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps
Hey, I just gave it a try and ran a quick test on booking.com. It took ~3 mins for a basic test. Do you cache the test steps so that future runs are faster and they don't call LLMs for the subsequent runs?

Also your current pricing is $300 for 1K tests which means $0.3 for each test. We tried out playwright mcp and it easily consumes 1M+ tokens for a test with ~20 steps (including image input). So with this pricing are you guys default alive?

Also is there a benchmark which you ran to prove the efficacy of your testing agent? because in the current stage it is a trust me bro kinda thing.

pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
The word dictionary is curated with guardrails. Also the dictionary contains words which are 1 BPE token long. Under 5-6 characters
pranshuchittora··on Show HN: Open-Source Agentic QA Harness with Memory
Thanks for those kind words. The landing page's demo required lots of sculpting. I would say that agent-qa is not only frontend focused. As you can run hooks in sandboxed env to test apis, so a better way to put it is with agent-qa you can test the product end-to-end not only UI.

But the issue with API testing / backend is that coding harnesses are really good at it. A product manager who writes user stories should be able to write tests for the product, and usually PMs don't care about the APIs.

Do give agent-qa a try, and consider giving it a star on GH https://github.com/vostride/agent-qa

Thanks!

pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
It is being used in production at https://vostride.com/agent-qa The issues was agent-qa have many different kinds of files tests, memory etc and there's too much FK references which LLMs need to resolve. Using id-agent worked like a charm
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
V1StGXR8_Z5jdHi6B-myT 21 Characters, 14 Tokens Really really inefficient
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Yes, we have the validation methods to verify the output. https://github.com/vostride/id-agent/#validateid

A random "-" separated words will fail the validation check.

pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Looks like it ;)
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
LLMs are good at predicting words, since each word in the id is ~1 BPE token. But uuids are random hex characters, this is where LLMs struggle to output the right ids.

You can use the .from method https://github.com/vostride/id-agent/#idagentfrominput-opts

To convert uuid or any text to id-agent based id. Then do the LLM inference and then convert it back to UUID.

pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
No worries, Checkout https://vostride.com/agent-qa to see how we are using this in production.
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Kinda similar, but this is token efficient. Each word is ~1 BPE token
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Would love to, can you please create an issue on the GH repo.
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Using "_" separator increases the token usage.
pranshuchittora··on Show HN: Id-agent – Token efficient UUID alternative for AI agents
Yes, that a valid point. That's why we have a verification method which can be part of the harness to make sure the ids are not hallucinated.
pranshuchittora··on Open-Source Agentic QA Harness with Memory
This is what teams are doing today. But LLMs have a tendency to greedily write tests, which leads to hacky tricks to make the test succeed.

agent-qa is a harness where playwright works as an execution kernel and LLM works as a observer, planner and verifier.

pranshuchittora··on Open-Source Agentic QA Harness with Memory
Do give it a try https://vostride.com/docs/agent-qa/quickstart
pranshuchittora··on Open-Source Agentic QA Harness with Memory
Yes, https://vostride.com/docs/agent-qa/configuration/global-conf...
pranshuchittora··on Click (2016)
Peak unemployment ;)
pranshuchittora··on Open-Source Agentic QA Harness with Memory
Hey, I am the creator of agent-qa.

Coding agents have accelerated software development, allowing folks to ship features at lightning speed, but whether the feature works in production without breaking existing behavior is still questionable.

Conventionally, either a software engineer or a QA engineer converts user stories / feature PRDs into composable end-to-end tests, allowing teams to catch regressions.

But with AI writing code, tests become the bottleneck. Though you can ask the coding agent to write tests, and it does write tests with reasonable correctness, AI greedily chases passing tests and sometimes bends the rules. Also, having access to the code allows it to write tests with shortcuts that might not mimic real user behavior.

With agent-qa, you can write tests in plain English (natural language). It is built upon battle-tested testing frameworks (Playwright for web and Appium for mobile). Playwright and Appium work as a kernel executing the planned actions, while AI runs in the harness doing observation -> planning -> executing planned actions (via kernel) -> self-healing (in case a planned action fails) -> verification.

The agent also evolves with every test run. It generates learning & product memories from each run, improving itself over time.

This is in an early stage, and I’m looking forward to your feedback.

Thanks!

Live Demo - https://vostride.com/demo/agent-qa GitHub - https://github.com/vostride/agent-qa (Consider giving it star) Good Day!

pranshuchittora··on Show HN: Agent-QA: Open-source AI end-to-end testing for web and mobile apps
Hey, I am the creator of agent-qa.

Coding agents have accelerated software development, allowing folks to ship features at lightning speed, but whether the feature works in production without breaking existing behavior is still questionable.

Conventionally, either a software engineer or a QA engineer converts user stories / feature PRDs into composable end-to-end tests, allowing teams to catch regressions.

But with AI writing code, tests become the bottleneck. Though you can ask the coding agent to write tests, and it does write tests with reasonable correctness, AI greedily chases passing tests and sometimes bends the rules. Also, having access to the code allows it to write tests with shortcuts that might not mimic real user behavior.

With agent-qa, you can write tests in plain English (natural language). It is built upon battle-tested testing frameworks (Playwright for web and Appium for mobile). Playwright and Appium work as a kernel executing the planned actions, while AI runs in the harness doing observation -> planning -> executing planned actions (via kernel) -> self-healing (in case a planned action fails) -> verification.

The agent also evolves with every test run. It generates learning & product memories from each run, improving itself over time.

This is in an early stage, and I’m looking forward to your feedback.

Thanks!

GitHub - https://github.com/vostride/agent-qa Consider giving it a Good Day!

pranshuchittora··on Scrcpy v4.0
scrcpy new disappoints. Consistently stable, butter smooth & well maintained. Kudos to the maintainers.
pranshuchittora··on Ask HN: What Are You Working On? (March 2026)
Building Universal mobile devtool — control iOS Simulators, Android Emulators, and real devices from a single dashboard and CLI

GitHub: https://github.com/pranshuchittora/simvyn

Do give it a try, Thanks!

pranshuchittora··on Show HN: Simvyn – open-source Universal mobile devtool
All locales & device size can be configured on demand.
← PreviousPage 2 of 3Next →