HNHacker News
TopNewBestAskShowJobs

mfkhalil

148 karma · joined May 24, 2023

moe@webhound.ai
submissionscomments
mfkhalil··on Model Fatigue Is Real
It’s a mix. On benchmarks, yeah, we route a request, then run the same request on models within a range of capability and compare the results.

We also have our entire team using it for their day-to-day work, and they’re devs who are pumping out tens of PRs a day, so they’re a pretty unforgiving test group and are always giving feedback.

We also work with companies that have it either running on real traffic or running shadow evals in the background (which do a version of the +/-1 method you’re describing), and they’re a huge source of feedback.

At the end of the day, most of our focus is on how it performs on production workloads. As useful as benchmarks can be for initial testing, I think it’s dangerous to over-optimize for those kinds of tasks (and most frontier models have already learned how to solve those problems optimally during training, so results can be misleading there).

mfkhalil··on GPT-6 Astra
Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.

In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.

https://docs.litellm.ai/docs/proxy/auto_routing

mfkhalil··on How to become a 10x ramble-coder
Very salient point. My take on this (and I know it’s not very popular on HN) is that trying to preserve every pre-AI coding skill is probably a fruitless endeavor.

AI-assisted coding is here and it’s not going anywhere. In the same way most modern software engineers aren’t writing assembly anymore, I don’t think future engineers will need to understand exactly how every part of a codebase works under the hood. It’s just another layer of abstraction.

Having said that, I do think knowledge of architecture/performance/security remains a pretty large part of prompting correctly, although more conceptually than actual implementation. As long as you’re still using those concepts in your prompting, you can keep that judgment sharp, even if some of the implementation fluency fades.

For example, with the boxing analogy, yes, you lose some boxing skills, but you’re also becoming a way better boxing coach by consistently teaching it. I think that’s the tradeoff.

mfkhalil··on Skillhound: A live index of every SKILL.md
Built this internally for our coding agents but have been loving it so much that we decided to make it public.

The web UI is free to use and does not require sign up, and has every public SKILL.md on GitHub indexed, with the index refreshing every 48 hours.

Programmatic use is $20 a month for unlimited searches, and I can vouch for the fact that it's made my agents feel so much smarter and more knowledgeable. Only issue is right now you have to nudge it to use skillhound ("use skillhound first") otherwise it tends to try to figure out best practices on its own.

Hope it can be as useful to the community as it is to us, and would appreciate any feedback you have.

mfkhalil··on Ask HN: What are you working on? (May 2026)
We're working on Webhound - budget controlled long-running deep research. You set a budget and Webhound will use that much in compute/LLM tokens to research your prompt, with built in verification cycles and optional added verification budget. Every claim is cited with evidence and a direct link to the tool calls that produced the claim

The goal is to build a deep research product for actual researchers, since we believe that it is an extremely powerful product that is still nascent but has enormous potential - which we've already seen with some early users.

https://webhound.ai

mfkhalil··on Shooting down ideas is not a skill
The least productive teams I've been a part of are the ones where everyone is waiting for their turn to say why an idea is bad. Sometimes being "too smart" can hold you back from building something genuinely new.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Hey, appreciate the feedback. Will address all your points.

Regarding Reddit, we have our own custom handler for Reddit URLs which uses the Reddit API, which we are billed for when we exceed free limits.

For Terms of Service, you're right, that is definitely an oversight on our part. We just published both our Terms of Service and Privacy Policy on the website.

When it comes to comparing with GPT-5 and Claude, we do believe that our prompting, agent orchestration, and other core parts of the product such as parallel search results analysis and parallel agents are improvements on just GPT-5 and Claude, while also allowing it to run at much cheaper costs on significantly smaller models. Our v1 which we built months ago was essentially the same as what GPT-5 thinking with web search currently does, and we've since made the explicit choice to focus on data quality, user controllability, and cost efficiency over latency. So while yes, it might give faster results and work better for smaller datasets, both we and our users have found Webhound to work better for siloed sources and larger datasets.

Regarding account deletion, that is also a fair point. So far we've had people email us when they want their account deleted, but we will add account deletion ASAP.

Criticism like this helps us continue to hold ourselves to a high standard, so thanks for taking the time to write it up.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Could you share the session url via the feedback form if you still have access to it?

That's really strange, it sounds like Webhound for some reason deleted the schema after extraction ended, so although your data should still be tied to the session it just isn't being displayed. Definitely not the expected behavior.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Accuracy-wise we think it's almost there but probably still a few iterations away from being perfect. It's great at eliminating a lot of the collection time though.

Interestingly, we're working with B2B clients right now where we use Webhound to curate and then act as the "validation" layer ourselves. The agent lets us offer these datasets way cheaper with live updates, but still with human oversight.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks for testing it! That's definitely a miss, sounds like it got confused about what you were looking for and went after board member pages instead of the actual meeting/document sites.

We're working on better query interpretation, but in the meantime you could try being more specific like "find BoardDocs or meeting document websites for each district" to guide it better. Also, you can usually figure out how it interpreted your request by looking at the entity criteria, those are all the criteria a piece of data needs to meet to make it in the set.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Yes, translation layer would probably be better terminology.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Did this resolve itself? If not shoot us an email at team@webhound.ai and we can get it figured out.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Supabase
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks, fixed!
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Fair point, most of our users have come from referrals/word of mouth so it hasn't really been an issue for us, but you're probably right that we should have more information on the landing page
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
We have, 2.5 Flash is about as small as we've been able to go while still delivering consistent results.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Yep: NextJS frontend, NodeJS backend, Gemini 2.5 Flash LLM, Firecrawl for crawling, self-hosted SearXNG for web search, and fly.io for hosting. Beyond that everything else is built internally, we don’t use many frameworks.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Sorry about that. If you tell it to restructure the schema and search plan around MCP as model context protocol it should work. The agent can get stuck on its initial interpretation sometimes.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Hey, would be happy to chat. Shoot us an email at team@webhound.ai and we can set up a time.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Yeah, we've noticed it overthinks simple tasks that could be solved with a single table scrape. The agent architecture is built for complex, multi-source problems so it overengineers straightforward queries.

Working on better task classification upfront to route simple requests more directly.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks, we have noticed that it can tend to "give up" early on certain sources. Ideally the critic agent would guide it back to the correct path of continuing to go deeper, but if that doesn't work usually just adding something to the prompt or sending it a message later on telling it to go deep on these sources would work.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Good point. Our main differentiation is the shared workspace - users can step in and guide the agent mid-task, kind of like Cursor vs Claude (which can technically generate the same code that Cursor does). Firecrawl (or any crawler we may use) is only part of the process, we want to make the collaborative process for user <> agent as robust and user controllable as possible.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks for the feedback! From what we've seen it's actually the other way around - once it gets a sense of where this information lives the latter stages of data collection go quicker, especially since it's able to deploy search agents in parallel to get information and doesn't need to do the manual work as much anymore. Having said that, it does sometimes forget to do that, and although we've added the critic agent to remind it to do that it can be inconsistent but usually if you step in and ask it to deploy agents in parallel that fixes it.

We use Gemini 2.5 Flash which is already pretty cheap, so inference costs are actually not as high as they would seem given the number of steps. Our architecture allows for small models like that to operate well enough, and we think those kinds of models will only get cheaper.

Having said all that, we are working on improving latency and allowing for more parallelization wherever possible and hope to include that in future versions, especially for enrichment. We do think that one of the weaknesses of the product is for mass collection - it's better at finding medium sized datasets from siloed sources and less good at getting large comprehensive datasets, but we're also considering approaches that incorporate more traditional scraping tactics for finding these large datasets.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks a lot regarding UI and good point on the schema editing.

We've been having similar thoughts about pricing and offering unlimited, but since it is feasible for us in the short term due to credits we enjoy offering that option to early users, even if it may be a bit naive.

Having said that, we are currently working on a pilot with a company whom we are offering live updates, and they are paying per usage since they don't want to have to set it up themselves, so we can definitely see the demand there. We also offer an API for companies that want to reliably query the same thing at a preset cadence, which is also usage based.

For crawling we use Firecrawl. They handle most of the blocking issues and proxies.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Yep, you can paste the list as text, or we also accept file uploads. Then, you can prompt it to enrich with certain attributes and it will do that for you.
mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
We maintain a constant browser state that gets fed into the system prompt which shows the most recent results, current page, where you are in the content, what actions are available, etc. It's markdown by default but can switch to HTML if needed (for pagination or CSS selectors). The agent always has full context of its browsing session.

A few design decisions we made that turned out pretty interesting:

1. We gave it an analyze results function. When the agent is on a search results page, instead of visiting each page one by one, it can just ask "What are the pricing models?" and get answers from all search results in parallel.

2. Long web pages get broken into chunks with navigation hints so the agent always knows where it is and can jump around without overloading its context ("continue reading", "jump to middle", etc.).

3. For sites that are commonly visited but have messy layouts or spread out information, we built custom tool calls that let the agent request specific info that might be scattered on different pages and consolidates it all into one clean text response.

4. We're adding DOM interaction via text in the next couple of days, so the agent can click buttons, fill forms, enter keys, but everything still comes back as structured text instead of screenshots.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
We currently use Firecrawl for our crawling infrastructure. Looking at their documentation, they claim to respect robots.txt, but based on user reports in their GitHub issues, the implementation seems inconsistent - particularly for one-off scrapes vs full crawls.

This is definitely something we need to address on our end. Site owners should have clear ways to opt out, and crawlers should be identifiable. We're looking into either working with Firecrawl to improve this or potentially switching to a solution that gives us more control over respecting these standards.

Appreciate you bringing this up.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Right now it can do that via URL params if that is how the website handles pagination, although we are pushing a feature in the next couple of days which allows it to take action on the DOM.

If it isn't doing that in your session, you can usually just step in and tell it to and it will follow your instructions.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks! Unlike a lot of our competitors who use search-inspired UX, we went with an agentic approach inspired by tools like Cursor - basically iterative user control.

Instead of just search query → final result (though you can do that too), you can step in and guide it. Tell it exactly where to look, what sources to check, how to dig deeper, how to use its notepad.

We've found this gets you way better results that actually match what you're looking for, as well as being a more satisfying user experience for people who already know how they would do the job themselves. Plus it lets you tap into niche datasets that wouldn't show up with just generic search queries.

mfkhalil··on Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web
Thanks, glad to hear you had a good experience.

We were heavily inspired by tools like Cursor - basically tried to prioritize user control and visibility above everything else.

What we discovered during iteration was that our users are usually domain experts who know exactly what they want. The more we showed them what was happening under the hood and gave them control over the process, the better their results got.

Page 1 of 2Next →