HNHacker News
TopNewBestAskShowJobs

shahules

131 karma · joined October 6, 2021

submissionscomments
shahules··on Cloning Bench: Evaluating AI Agents on Visual Website Cloning
My team works on automatic environment generation for RL post-training. One of our projects is using coding agents to build web clones for BUAs/CUAs.

We tested Gemini, Claude Code, GLM, and Codex using our harness on their abilities to recreate a Slack workspace and benchmarked their performance.

Saw a variety of results:

- *Gemini 3 Pro:* Achieved the highest visual score (0.91 SSIM) but lacked interactive functionality. - *Claude Opus 4.6:* Developed the most complete application, balancing full interactivity with consistent self-correction. - *GLM-5:* Produced the best code architecture but reached a plateau in visual improvement. - *GPT-5.3 Codex:* Initialized quickly but entered a five-hour "scaling spiral" that failed to yield further progress.

Next, we’re planning:

- More web apps for cloning and benchmarking across the models - More functionality (the trajectory didn’t include full Slack features) - Better scoring for functionality (easier to catch Gemini’s mistake)

Repo: https://github.com/vibrantlabsai/cloning-bench

Blog post: https://vibrantlabs.com/blog/pa-bench

shahules··on PA bench: Evaluating web agents on real world personal assistant workflows
Nice, their training recipe seems unique.
shahules··on PA bench: Evaluating web agents on real world personal assistant workflows
After doing few experiments, I think that having Agents work on browser for all tasks wouldn't be best due to many factors like token cost, safety, etc. But browser/computer can be a tool that the agent can be alongside MCPs to complete tasks that requires interaction with such modalities.
shahules··on PA bench: Evaluating web agents on real world personal assistant workflows
There are few agents like browser-use, skyvern etc that may provide this capability.
shahules··on PA Bench: Evaluating Frontier Models on Multi-Tab Pa Tasks
Most current web agent benchmarks focus on single-tab tasks (e.g., 'go to Gmail and star this email'). We found that frontier models that score highly on those tasks (like in WebArena) often fall apart when they have to coordinate context across 2+ applications. We built a simulated environment with scenarios and deterministic verifiers to see why.
shahules··on Evals are not all you need
It's an interesting article and I agree with some points you brought up here. But here are some of them to which I don't agree to

1. Evals are used throughout the article in the sense of LLM benchmarking, but this is not the point. One could effectively evaluate any AI system by building custom evals.

2. The purpose of evals is to help devs systematically improve their AI systems (at least how we look at it) not any of the ones listed in your article. It's not a one-time thing, it's a practice like the scientific method.

shahules··on Too much efficiency makes everything worse (2022)
Can't agree with you more my friend. Another point on a philosophical level is efficiency or optimization in life, which always focuses on tangible aspects and ignores the greater intangible aspects of life.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Deepeval also uses Ragas underneath. They initially took a different approach by allowing uses to formulate test cases but we were focusing on RAGs only and creating metrics and features like synthetic test data generation for it. Now that we are doing good in the RAG category, we also want to expand to solve the greater challenge.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Hey, I would recommend checking out our PRs. There would be PRs that have modified some of the prompts to better suit Mixtral.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
I think it's true for any early-stage library/framework. The tradeoff is then you will have to keep maintaining it, add support to other LLMs if you change LLMs, etc. Then in the end OSS will be far ahead because by that time it will have smoothened its rough edges.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Hey, DeepEval is interesting. What do you mean by "evaluating any LLMs"?
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
1) Horizondal expansion and support are core to every framework/library. This won't make us a competitor to LC, we actually use langchain-core to support many of these like supporting different LLMs. 2) We operate in a layer underneath the stack of evals and testing because we want to solve the problem from the ground up rather than building a fancy tracing tool, which comes later in the stack.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Yes, this is an interesting challenge we are also excited about.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Hey, this is a highly requested feature. We will be implementing it soon. Something like a rough estimate is what we are planning to do.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Thanks for your feedback. We have tested Ragas on alternatives like Claude, Mixtral, Gemini, etc.

Although we support all LLMs supported by Langchain, sadly many of the OSS models out of the box aren't capable of generating JSON-type output which is important for us to ensure reproducibility.

shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Thank you.
shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Thank you for asking this question.

To answer this question, I will explain two directions of Ragas.

The first one is the horizontal expansion of the library which involves features like

- Giving you the ability to use any LLMs instantly without any hassle

- Asynchronous evaluations, integrations with tracing tools, etc

- Automatic support to adapt metrics to any language

The second is vertical expansion or adding more core features like metrics to Ragas which includes.

- Synthetic test data generation: this is something that is heavily loved by our community so we are continuously improving the quality of it. https://docs.ragas.io/en/stable/concepts/testset_generation....

Now, as we expand in both directions we aim to solve the problem of how to evaluate and test compound systems. Now, to solve this we will be innovating and working on features like feedback utilization, automatically synthesizing assertions, etc to solve this hard problem.

I hope I was able to answer your question. Would love to discuss more.

shahules··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Hey, thanks for trying out Ragas. As an open-source library, we are continuously improving from the feedback from the community which I see as our primary strength. I am sure that Ragas is not perfect yet, but I can assure you that it is 10x better than it was a few months ago.
shahules··on Show HN: Ragas – Open-source library for evals and testing RAG systems
Hey, The cost depends on how many samples you would want to generate. We are finetuning smaller OSS models to tackle the cost factor of synthetic data generation which we will open source this month.
shahules··on Show HN: Ragas – Open-source library for evals and testing RAG systems
Thank you. Yes. We support almost all opensource models. You can customise it easily by referring to https://docs.ragas.io/en/stable/howtos/customisations/bring-...
shahules··on Show HN: Ragas – Open-source library for evals and testing RAG systems
Yes,ofc. We support all LLMs supported by Langchain. https://docs.ragas.io/en/stable/howtos/customisations/bring-...
shahules··on Patterns for building LLM-based systems and products
Thank you. That's great. We are working on forming paradigms to evaluate Agents, I'll get in touch with you.
shahules··on Patterns for building LLM-based systems and products
Evals is not suitable for evaluating LLM applications such as RAG, etc because one has to evaluate on their own data where no golden test data exists, and techniqus used have poor correlation with human judgement. We have build RAGAS framework for this https://github.com/explodinggradients/ragas
shahules··on GPT-4 details leaked?
This guy doesn't have any idea what he is talking about. He consistently posts such bullshit on twitter. Mostly copy paste with added spice mix.
shahules··on Building LLM Applications for Production
Great, excellent read. This article I wrote some of the open-source finest foundation LLMs https://explodinggradients.com/the-rise-of-open-source-large...
shahules··on Enhance Speech from Adobe – Free AI filter for cleaning up spoken audio
Hi, most models performs best at 16KHz. Current architectures does not support real-time speech enhancement but I plan to add that in future.
shahules··on Enhance Speech from Adobe – Free AI filter for cleaning up spoken audio
Thanks for pointing out, I'm not a regular user of HN. I myself am the creator of Mayavoz.
shahules··on Enhance Speech from Adobe – Free AI filter for cleaning up spoken audio
There is an open-source tool that does the same https://github.com/shahules786/mayavoz
shahules··on Enhance Speech from Adobe – Free AI filter for cleaning up spoken audio
Hi there, I have made a free open-source tool that does better. Care you check that out? https://github.com/shahules786/mayavoz