HNHacker News
TopNewBestAskShowJobs

shardullavekar

221 karma · joined February 22, 2016

Founder at 100x.bot Twitter: x.com/shardullavekar LinkedIn:https://www.linkedin.com/in/shardul-lavekar-96049414/
submissionscomments
shardullavekar··on Why AI hasn't replaced software engineers, and won't
some x.com influencer claiming that they built a one man company with agents doing everything. Like, hey Claude - create a staffing automation company and a swarm of agents search, reach out, follow ups, and AI voice calls etc. Its bs marketing, I get it but "what if this comes true, not now but in future?" - drives the anxiety. The sandwich framework sounds great until the decision and accountability is passed on to AI with guardrails. The AI in its current state is great to take the feedback and decide the next course of actions.
shardullavekar··on Show HN: VidStudio, a browser based video editor that doesn't upload your files
looks great. any plans of abstracting these functions with an LLM integration?
shardullavekar··on Forcing an inversion of control on the SaaS stack
hence the sharability and subscription. other users need to explicitly subscribe to the page boosters. Else they continue with what they have.
shardullavekar··on Forcing an inversion of control on the SaaS stack
users didnt ask for slow apis either but there they are. I am speaking for the user here and sharing their frustration. Allowing UI modification to fit the user needs should be a default now. The APIs already act as a gaurdrail on what's possible
shardullavekar··on Forcing an inversion of control on the SaaS stack
I created that twitter responder after reading this post (https://news.ycombinator.com/item?id=47568028). That wasn't to call out what SaaS companies should prioratise but to show how easy it would be for a user to do it.
shardullavekar··on Forcing an Inversion of Control on the SaaS Stack
got our extension approved, post which we had no issues.
shardullavekar··on Forcing an inversion of control on the SaaS stack
these embeddable UI could be a direct ask on how users want a workflow, the SaaS vendors can distribute the embeddable UI and see if it clicks with a lot of users. Would push them to create a stable API
shardullavekar··on Forcing an inversion of control on the SaaS stack
> In the end, your obligation as a company, regardless of your product, is to generate profits.

No denying that. SaaS started with a user problem at the center of it and as they scaled, forgot about an individual user. This only presents the user frustration and a possible solution to it.

shardullavekar··on Forcing an inversion of control on the SaaS stack
I think the right step would be to somehow communicate to the vendor that this feature is needed (eliminating the PM backlog BS) and their coding Agents should pick it and build it. The real moat they have is SaaS vendors have everyone believe that trivial feature requests take time to implement.
shardullavekar··on LinkedIn is searching your browser extensions
wondering if all the new browsers in the market have the ability to block such scanning APIs explicitly.
shardullavekar··on Claude Code's source code has been leaked via a map file in their NPM registry
wondering how this fares for languages other than English.
shardullavekar··on Coding Agents Could Make Free Software Matter Again
is someone building an agent to manage self hosted infra? A lot of "convenience" issues around self hosting free software would go away.
shardullavekar··on A Visual Introduction to Machine Learning (2015)
has anyone come across an r2d3-style explainer for something as high-dimensional as a Transformer's attention mechanism?
shardullavekar··on Chaining FFmpeg with a Browser Agent
jack_pp made a point in the comments, worth noting.
shardullavekar··on Chaining FFmpeg with a Browser Agent
no 2nd thoughts about it, we are only making ffmpeg more accessible and embeddable.
shardullavekar··on Chaining FFmpeg with a Browser Agent
1. download a larger video from s3. 2. Use NLE and cut it into shorts. (crop, resize, subtitles etc.) 3. Upload shorts on YouTube, Instagram, Tiktok.

He does use davinci resolve but only for 2.

NLEs make ffmpeg a standalone yet easy to use tool.

Not denying that major heavy lifting is done by the NLE. We go a step ahead and make it embeddable in a larger workflow.

shardullavekar··on Chaining FFmpeg with a Browser Agent
people intimidated by a CLI tool but find tools like chatgpt easy to use and those who have video editing as a part of larger workflow.
shardullavekar··on Chaining FFmpeg with a Browser Agent
That's exactly how it is implemented!
shardullavekar··on Chaining FFmpeg with a Browser Agent
we took a basic example and described it. (will try adding a complex one)

we have our designer/intern in our minds who creates shorts, adds subtiles, crops them,and merges the audio generated. He is aware of ffmpeg and prefers using a SaaS UI on top of it.

However, we see him hanging out on chatgpt, or gemini all the time. He is literally the no coder we have in mind.

We just combined his type what you want + ffmpeg workflows.

shardullavekar··on Chaining FFmpeg with a Browser Agent
so creators on 100x will create well defined workflows that others can reuse. If a workflow is not found, llm creates one on the go and saves it.
shardullavekar··on Chaining FFmpeg with a Browser Agent
thanks for calling it out, I will correct the before vs after section. But you can describe any ffmpeg capability in plain English and the underlying ffmpeg tool call takes care of it.
shardullavekar··on Chaining FFmpeg with a Browser Agent
true, companies like Descript, Veed, or Kapwing exist because no coders find this syntax intimidating. Plus, a CLI tool stands out of a workflow. We wanted to change that.
shardullavekar··on A stateful browser agent using self-healing DOM maps
We launched Agent4 recently. You can install it from here: https://chromewebstore.google.com/detail/agent4/kipkglfnhnpb...

The one you refer will be taken down soon. Ping me on discord if you need help in trying it.

shardullavekar··on A stateful browser agent using self-healing DOM maps
a built-in mcp server that takes a look at what's broken and communicates with cursor is on our roadmap. Join discord and we will keep you posted there.
shardullavekar··on A stateful browser agent using self-healing DOM maps
It's a chrome extension. Works if you use chrome.
shardullavekar··on Your Browser Agent Is Thinking Too Hard
A workflow can have subjective parts too. For example, click on button A if it satisfies certain conditions I wrote in plain English, otherwise click on B.

These subjective elements can be defined with user inputs/prompts.

So a workflow is a literal script with embedded LLM calls for branching or even scraping details where literal script feels tedious.

shardullavekar··on Cloudflare Radar: AI Insights
would be interesting to see if linkedin (and the likes who don't want to be crawled) signs up for the pay-per-crawl that CF may come up with.
shardullavekar··on Claude for Chrome
I've been building a general browser agent myself, and I’ve found the biggest bottleneck in these systems isn’t capability demos but long-running reliability.

Tools like Manus / GPT Agent Mode / BrowserUse / Claude’s Chrome control typically make an LLM call per action/decision. That piles up latency, cost, and fragility as the DOM shifts, sessions expire, and sites rate-limit. Eventually you hit prompt-injection landmines or lose context and the run stalls.

I am approaching browser agents differently: record once, replay fast. We capture HTML snapshots + click targets + short voice notes to build a deterministic plan, then only use an LLM for rare ambiguities or recovery. That makes multi-hour jobs feasible. Concretely, users run things like:

Recruiter sourcing for hours at a stretch

SEO crawls: gather metadata → update internal dashboard → email a report

Bulk LinkedIn connection flows with lightweight personalization

Even long web-testing runs

A stress test I like (can share code/method): “Find 100+ GitHub profiles in Bangalore strong in Python + Java, extract links + metadata, and de-dupe.” Most per-step-LLM agents drift or stall after a few minutes due to DOM churn, pagination loops, or rate limits. A record-→-replay plan with checkpoints + idempotent steps tends to survive.

I’d benchmark on:

Throughput over time (actions/min sustained for 30–60+ mins)

End-to-end success rate on multi-page flows with infinite scroll/pagination

Resume semantics (crash → restart without duplicates)

Selector robustness (resilient to minor DOM changes)

Cost per 1,000 actions

Disclosure: I am the founder of 100x.bot (record-to-agent, long-run reliability focus). I’m putting together a public benchmark with the scenario above + a few gnarlier ones (auth walls, rate-limit backoff, content hashing for dedupe). If there’s interest, I can post the methodology and harness here so results are apples-to-apples.

shardullavekar··on Automated Unit Test Improvement Using Large Language Models at Meta
At unlogged.io, for some time - our primary focus was to auto-generate junit tests. The approach didn't take off for a few reasons: 1. A Lot of generated test code that no devs wanted to maintain. 2. The generated tests didn't simulate real-world scenarios. 3. Code coverage was a vanity metric. Devs worked around to reach their goals with scenarios that didn't matter.

We are currently working on offering no-code replay tests that simulate all unique production scenarios and developers can replay locally while mocking external dependencies.

Disclaimer: I am a founder at unlogged.io

shardullavekar··on Show HN: Unlogged – open-source record and replay for Java
Yes, we already offer it. In fact, Beepkart has just signed up for a paid support. Use the plugin and let us know if you need support with something immediately.
Page 1 of 2Next →