HNHacker News
TopNewBestAskShowJobs

asim-shrestha

120 karma · joined April 8, 2023

submissionscomments
asim-shrestha··on Ask HN: Who is hiring? (November 2024)
Reworkd | Backend / Infrastructure | ONSITE San Francisco

At https://reworkd.ai/, we're building application layer LLM agents to extract web data at scale. We are foundational data infrastructure for startups today that are fine tuning models or building some web data constrained product. We're backed by YC, Paul Graham, AI grant, and many others.

We're looking for backend/infrastructure/full stack engineers to: 1. Help optimize our queue/browser infrastructure to enable ingesting millions of pages a day 2. Build Playwright systems to handle two way syncing between users and cloud browsers 3. Work on the UI/UX to enable user <> LLM interaction for code generation

If any of this sounds interesting, email me: asim (at) reworkd (dot) ai or apply directly at: https://www.ycombinator.com/companies/reworkd/jobs

asim-shrestha··on Ask HN: Who is hiring? (June 2024)
Reworkd (https://reworkd.ai/) | San Francisco (In-person) | Full-time | 2+ years of full-time experience

We're hiring a founding backend engineer to help us build infrastructure to run web agents at scale.

We're a super small scrappy team of four that's been working on the application layer of web agents since it's inception. Our projects have over 30k stars on GitHub, we're backed by PG himself + a bunch of great investors, and we have 3+ years of runway. Join us if you want to grind, want a lot of ownership, and are experienced. (More info about the role in our job posting)

You can either apply through bookface (https://www.ycombinator.com/companies/reworkd/jobs/4f6BHpT-f...) or directly email me the following (asim@reworkd.ai) the following:

1. Why you'd want to work for us. What makes you interested specifically in web agents or our open source projects (no generic responses)

2. Why you'd be a good fit

asim-shrestha··on Show HN: GPT-4 vision utilities to browse the web
Mind elaborating here? Happy to update the README if there are issues
asim-shrestha··on Ask HN: Would an easier way to scrape 100s of websites be useful to you?
No we haven't. We're building off existing scraping tools (eg. Selenium) and building the reasoning engine that will take actions on the page via these tools

Unfamiliar with blocking mechanisms, could you share some things you would do to block existing selenium scraping jobs?

asim-shrestha··on Ask HN: Would an easier way to scrape 100s of websites be useful to you?
Appreciate the input Damon, ethical concerns are definitely a consideration and we'll want to be respectful of mechanisms like robots.txt
asim-shrestha··on Ask HN: Would an easier way to scrape 100s of websites be useful to you?
Not specifically social media sites, getting through prevention would be difficult and there are already a lot of existing companies working on scraping popular social media sites.

Interesting idea, we're definitely looking into coupling OCR and LLMs today but not for that particular case. I think raw language models with a good workflow are typically good enough to extract structured data from things like books

ML training is definitely one area we can see this being useful. General data aggregation across a large industry (clothing, retail, etc) is something we want to look into. Also RPA style workflows involving multi-click actions across a variety of sites

asim-shrestha··on Ask HN: Would an easier way to scrape 100s of websites be useful to you?
Our take is scraping a website in isolation is typically quite easy for a somewhat technical person.

Scraping at larger scale is where it becomes challenging, the problems we want to tackle are: 1. You need at least a bit of technical expertise to do things like configure selectors properly 2. Websites typically have moderation in place to block scrapers 3. Scrapers are prone to changes in the site layout 4. Creating on the order of 100s of scrapers is difficult and time consuming. Creating this many will amplify the previous issue

asim-shrestha··on Do you need to scrape data from 100s to 1000s of websites?
Gotcha, will do

Appreciate the feedback, very new to posting here

asim-shrestha··on Do you need to scrape data from 100s to 1000s of websites?
Ah, none of those existing posts are relevant to this. Updated the messaging a bit here to sound less spammy - thank you

Additionally, we don't actually have something to show/launch. Was just interested in hearing perspectives of people that have had this problem or have tried to tackle this problem

asim-shrestha··on [dead]
A web based autonomous agent platform using lessons from AutoGPT and BabyAGI. A bit primitive in terms of functionality but serves as a demonstration of what is going to be capable

Website here: https://agentgpt.reworkd.ai/ GitHub: https://github.com/reworkd/AgentGPT Thread: https://twitter.com/asimdotshrestha/status/16448837277079592...

asim-shrestha··on [dead]
Recently helped build this using OpenAI, Langchain, and Next.js

Try it here: https://agentgpt.reworkd.ai/