HNHacker News
TopNewBestAskShowJobs

mdahardy

82 karma · joined November 1, 2022

submissionscomments
mdahardy··on AI capability isn't humanness
Our main argument is that outputs will become increasingly indistinguishable, but the processes won't. E.g. in 5 years if you watch an AI book a flight it will do it in a very non-human way, even if it gets the same flight you yourself would book.
mdahardy··on AI capability isn't humanness
This is a fair criticism we should've addressed. There's actually a nice study on this: Vong et al. (https://www.science.org/doi/10.1126/science.adi1374) hooked up a camera to a baby's head so it would get all the input data a baby gets. A model trained on this data learned some things babies do (eg word-object mappings), but not everything. However, this model couldn't actively manipulate the world in the way that a baby does and I think this is a big reason why humans can learn so quickly and efficiently.

That said, LLMs are still trained on significantly more data pretty much no matter how you look at it. E.g. a blind child might hear 10-15 million words by age 6 vs. trillions for LLMs.

mdahardy··on Show HN: ModelGuessr: Can you tell which AI you're chatting with?
Ah, nice idea. I hadn't considered locking it after you guess correctly.
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
The cross-tile challenges were quite robust - every model struggled with them, and we tried with several iterations of the prompt. I'm sure you could improve with specialized systems, but the models out-of-the-box definitely struggle with segmentation
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
That's a cool idea. I bet it would work better.
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
yes
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
After watching hundreds of these runs, Gemini was by far the least frustrating model to observe.
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
We have an example of a failed cross-tile result in the article - the models seem like they're much better at detecting whether something is in an image vs. identifying the boundaries of those items. This probably has to do with how they're trained - if you train on descriptions/image pairs, I'm not sure how well that does at learning boundaries.

Reload are challenging because of how the agent-action loop works. But the models were pretty good at identifying when a tile contained an item.

mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
Same! As we talk about in the article, the failures were less from raw model intelligence/ability than from challenges with timing and dynamic interfaces
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
While running this I looked at hundreds and hundreds of captchas. And I still get rejected on like 20% of them when I do them. I truly don't understand their algorithm lol
mdahardy··on Benchmarking leading AI agents against Google reCAPTCHA v2
You could definitely do better than we do here - this was just a test of how well these general-purpose systems are out-of-the-box
mdahardy··on Ask HN: Who is hiring? (September 2025)
Roundtable | https://roundtable.ai | On-site San Francisco, CA | Full-time

Roundtable is a research and deployment company building the proof-of-human layer in digital identity. Roundtable seeks to research and build real-world Turing Tests, and to productize these developments in bot detection, fraud prevention, and continuous authentication.

We're hiring a Member of Technical Staff. This is not a typical “Founding Engineer” position. In particular, this role will involve more research than similar positions. You should be prepared to spend around half of your time on research and half on engineering.

The ideal candidate has experience in both. However, we don’t strictly require research experience as long as you have a strong quantitative background and an eagerness to learn. We do require proficiency in both web development (Javascript, Node.js) and Python.

If you're interested shoot an email to matt@roundtable.ai with your resume and a short blurb on why you're interested (note: we are unable to offer visa sponsorship at this time).

mdahardy··on Bot or Human? Creating the Invisible Turing Test for the Internet
Co-founder of Roundtable here.

I agree that better authentication methods for AI agents are needed. But right now bots and malicious agents are a real problem for anyone running sites with significant traffic. In the long run I don’t think human traffic will go to zero even if its relative proportion is reduced.

mdahardy··on "The closer to the train station, the worse the kebab" – A "Study"
Seems like a lot of this could be explained by better food tending to be served in locations with lower commercial real estate prices (I believe Tyler Cowen has written about this).
mdahardy··on Show HN: RoundtableJS – Open-source programmatic survey library
Thanks for letting us know - we're planning to support TypeScript soon. A lot of stuff on our roadmap!