221 karma · joined February 22, 2016
No denying that. SaaS started with a user problem at the center of it and as they scaled, forgot about an individual user. This only presents the user frustration and a possible solution to it.
He does use davinci resolve but only for 2.
NLEs make ffmpeg a standalone yet easy to use tool.
Not denying that major heavy lifting is done by the NLE. We go a step ahead and make it embeddable in a larger workflow.
we have our designer/intern in our minds who creates shorts, adds subtiles, crops them,and merges the audio generated. He is aware of ffmpeg and prefers using a SaaS UI on top of it.
However, we see him hanging out on chatgpt, or gemini all the time. He is literally the no coder we have in mind.
We just combined his type what you want + ffmpeg workflows.
The one you refer will be taken down soon. Ping me on discord if you need help in trying it.
These subjective elements can be defined with user inputs/prompts.
So a workflow is a literal script with embedded LLM calls for branching or even scraping details where literal script feels tedious.
Tools like Manus / GPT Agent Mode / BrowserUse / Claude’s Chrome control typically make an LLM call per action/decision. That piles up latency, cost, and fragility as the DOM shifts, sessions expire, and sites rate-limit. Eventually you hit prompt-injection landmines or lose context and the run stalls.
I am approaching browser agents differently: record once, replay fast. We capture HTML snapshots + click targets + short voice notes to build a deterministic plan, then only use an LLM for rare ambiguities or recovery. That makes multi-hour jobs feasible. Concretely, users run things like:
Recruiter sourcing for hours at a stretch
SEO crawls: gather metadata → update internal dashboard → email a report
Bulk LinkedIn connection flows with lightweight personalization
Even long web-testing runs
A stress test I like (can share code/method): “Find 100+ GitHub profiles in Bangalore strong in Python + Java, extract links + metadata, and de-dupe.” Most per-step-LLM agents drift or stall after a few minutes due to DOM churn, pagination loops, or rate limits. A record-→-replay plan with checkpoints + idempotent steps tends to survive.
I’d benchmark on:
Throughput over time (actions/min sustained for 30–60+ mins)
End-to-end success rate on multi-page flows with infinite scroll/pagination
Resume semantics (crash → restart without duplicates)
Selector robustness (resilient to minor DOM changes)
Cost per 1,000 actions
Disclosure: I am the founder of 100x.bot (record-to-agent, long-run reliability focus). I’m putting together a public benchmark with the scenario above + a few gnarlier ones (auth walls, rate-limit backoff, content hashing for dedupe). If there’s interest, I can post the methodology and harness here so results are apples-to-apples.
We are currently working on offering no-code replay tests that simulate all unique production scenarios and developers can replay locally while mocking external dependencies.
Disclaimer: I am a founder at unlogged.io