HNHacker News
TopNewBestAskShowJobs

goncharom

4 karma · joined November 9, 2025

submissionscomments
goncharom··on Ask HN: What Are You Working On? (December 2025)
A multi-purpose scrapper to turn any webpage into structured data: https://news.ycombinator.com/item?id=45870231

It uses LLMs to generate python code to scrap a webpage to fit any Pydantic model provided:

  from hikugen import HikuExtractor
  from pydantic import BaseModel
  from typing import List
  
  class Article(BaseModel):
      title: str
      author: str
      published_date: str
      content: str
  
  class ArticlePage(BaseModel):
      articles: List[Article]
  
  extractor = HikuExtractor(api_key="your-openrouter-api-key")
  
  result = extractor.extract(
      url="https://example.com/articles",
      schema=ArticlePage
  )
  
  for a in result.articles:
      print(a.title, a.author)
goncharom··on Schizophrenia sufferer mistakes smart fridge ad for psychotic episode
Relevant read (not my own): https://simone.org/advertising/
goncharom··on A new AI winter is coming?
it has a proper paper attached right at the beginning of the article
goncharom··on A new AI winter is coming?
As I said then, and probably echoing what other commenters are saying - what do you mean by understanding when you say computers understand nothing? do humans understand anything? if so, how?
goncharom··on A new AI winter is coming?
Every time I see comments like these I think about this research from anthropic: https://www.anthropic.com/research/mapping-mind-language-mod...

LLMs activate similar neurons for similar concepts not only across languages, but also across input types. I’d like to know if you’d consider that as a good representation of “understanding” and if not, how would you define it?

goncharom··on Greggit – Google but it's only the Reddit results
Yes this is literally just appending site: reddit.com to the query and redirecting to google. This page is a single HTML: https://github.com/goncharom/greggit

This was meant as a silly project! I caught myself adding site:reddit.com to a lot of google searches so I figured I’d just make a shortcut.

goncharom··on Show HN: Turn any webpage into structured data via LLM codegen
The regeneration loop was probably the most interesting part to work on: you need very strict constraints on what “good” content looks like and what the specific issue is when codegen fails. I found Pydantic annotations to be specifically useful for this.
goncharom··on Ask HN: What Are You Working On? (Nov 2025)
I've been working web scraping using LLMs, I just shared one of the libraries I created to get structured data from arbitrary pages: https://news.ycombinator.com/item?id=45870231

Instead of sending the page's HTML to an LLM, Hikugen asks it to generate python code to fetch the data and enforces the generated data conforms to a Pydantic schema defined by the user. I'm using this to power yomu (https://github.com/goncharom/yomu), a personal email newsletter built from arbitrary websites.