HNHacker News
TopNewBestAskShowJobs

KhoomeiK

1,299 karma · joined August 7, 2018

research @ openai

https://rpandey.tech

submissionscomments
KhoomeiK··on Diffusion models are real-time game engines
NVIDIA did something similar with GANs in 2020 [1], except users could actually play those games (unlike in this diffusion work which just plays back simulated video). Sentdex later adapted this to play GTA with a really cool demo [2].

[1] https://research.nvidia.com/labs/toronto-ai/gameGAN/

[2] https://www.youtube.com/watch?v=udPY5rQVoW0

KhoomeiK··on Introduction to Program Synthesis
Everything relevant in "program synthesis" moved to the new buzzword "codegen"
KhoomeiK··on Show HN: Revideo – Create Videos with Code
Interesting—LangChain seemed kinda like unnecessary abstractions in natural language (since everything is just string manipulations), but with AI video, there's so many different abstractions that I'd need to handle (images, puppeting, facegen, voicegen, etc).

Seems like there might be room for a "LangChain for Video" in this space...

KhoomeiK··on Atash Behram – Types of Fire
Just unrolled the thread for you here!

https://threadreaderapp.com/thread/1794082465398812770.html

KhoomeiK··on Atash Behram – Types of Fire
Thanks! I have no idea—unfortunately, very few Hindus maintain the Vedic fire rites. There are also no active central authorities on matters of Vedic ritual. The only plan of now is to use this interpretation in my own yajña practice.
KhoomeiK··on Atash Behram – Types of Fire
Vedic Hinduism had a similar concept of eternal fire. I recently wrote up a twitter thread [1] explaining how the modern interpretation of Vedic instructions on starting these sacred fires misunderstands the text.

Etymology tidbit: "Bhārata", India's Sanskrit name, refers to the forerunner clan that established India's first historically recorded political entity—the Kuru Kingdom—around 1200 BC near modern Delhi. The clan itself was named "Bhārata" due to their ardent bearing ("bhar-" in Sanskrit) of the sacred fire.

[1] https://x.com/khoomeik/status/1794082465398812770

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Maybe this project another commenter is working on?

https://news.ycombinator.com/item?id=40373310

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Awesome pics! We love tarsiers too
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Great question! See this thread:

https://news.ycombinator.com/item?id=40369713

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Yes it does work headless and we do grab a fullpage screenshot including scrolling (by resizing viewport to content height). We haven’t had to deal with infinite scrolling much but that’s an interesting feature we’d appreciate a PR for.

We haven’t tried Apple’s OCR but hopefully will integrate Azure OCR soon based on others’ advice.

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
They do show textboxes with labels. From our readme:

"Keep in mind that Tarsier tags different types of elements differently to help your LLM identify what actions are performable on each element. Specifically:

[#ID]: text-insertable fields (e.g. textarea, input with textual type)

[@ID]: hyperlinks (<a> tags)

[$ID]: other interactable elements (e.g. button, select)

[ID]: plain text (if you pass tag_text_elements=True)"

Do you see the search boxes labeled [#4] and [#5] at the top? And before you say that the tag is on a different line from the placeholder text—yes, and our agent is smart enough to handle that minor idiosyncrasy. Are you shocked? :)

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
We have a lot more powerful use-cases for Tarsier in web data extraction at the moment. Stay tuned for a broader launch soon!
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
We run OCR on the screenshot & convert it to whitespace-structured text, that is passed to the LLM. The images below might make it clearer for you:

[1] https://github.com/reworkd/tarsier/blob/main/.github/assets/...

[2] https://github.com/reworkd/tarsier/blob/main/.github/assets/...

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Yup, it could! There are a lot of players in the generalist personal web agent space but I personally think that use-case will be eaten by big players since fundamental foundation model improvements are required. That being said, Tarsier is a great place to start for building an open-source web agent for automating cool little tasks.

At Reworkd, we're focused on web agents for data extraction at scale, which isn't as hyped as the generalist agents but we find provides a lot of value and already works pretty well.

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
VimGPT couples the perception to a specific LLM/agent whereas Tarsier is solely a perception system that you can use for any uni/multi-modal web agent. So it's hard to compare, but you could say that VimGPT's performance probably lies somewhere in the middle of Tarsier's performance distribution (which varies as a function of your specific agent/prompt system).
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
We're hoping to release an evals paper about Bananalyzer this summer and compare Tarsier to a variety of other perception systems in it. The hard part with evaluating a perception/context system though is that it's very intertwined with the agent's architecture, and that's not something we're comfortable fully open-sourcing yet. We'll have to think of interesting ways to decouple the perception system and eval them with Bananalyzer.
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
More OCR providers are on the roadmap and we'd love for you to contribute any local OCR models you think could be useful! I wouldn't call it a wrapper though :)
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Thanks! We might put out a paper about it with some Carnegie Mellon collaborators this summer.
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Must have been the mods, I spent quite a bit of time on the content lol
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Thanks for pointing this out! Yeah, it's pretty strange. We thought including Show HN text was encouraged to engage with the community?
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
That's an interesting problem—Tarsier probably isn't the best solution here since it's focused on webpage perception rather than any kind of OCR. But one could try adapting the `format_text` function in tarsier/text_format.py to convert any set of OCR annotations to a whitespace-structured string. Curious to see if that works.
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Cool connection, hadn't seen this before but feels intuitively correct! I also formulate similar (but a bit more out-there) philosophical thoughts on word-meaning as being described by the topological structure of its corresponding images in embedding space, in Section 5.3 of my undergrad thesis [1].

[1] https://arxiv.org/abs/2305.16328

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Hm, not sure I follow why those situations would be especially difficult? Regarding website changes, the nice thing about using LLMs is that we can simply provide the previous scraper as context and have it regenerate the scraper to "self-heal" when significant website changes are detected.
KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Totally agreed—this is a design choice that basically comes from our agent architecture, and the codegen-based architecture that we think will likely proliferate for web agent tasks in the future. We provide Tarsier's text/screenshot to an LLM and have it write code with generically written selectors rather than the naive selectors that Tarsier assigns to each element.

It's sort of like when you (as a human) write a web scraper and visually click on individual elements to look at the surrounding HTML structure / their selectors, but then end up writing code with more general selectors—not copypasting the selectors of the elements you clicked.

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Yup, evals can definitely be tough. We basically have a suite of several hundred web data extraction evals in a tool we built called Bananalyzer [1]. It's made it pretty straightforward for us to benchmark how accurately our agent generates code when it uses Tarsier-text (+ GPT-4) for perception v.s. Tarsier-screenshot (+ GPT-4V/o).

Will have to look into supporting Azure OCR in Tarsier then—thanks for the tip!

[1] https://github.com/reworkd/bananalyzer

KhoomeiK··on Show HN: Tarsier – Vision utilities for web interaction agents
Thanks! Yeah, it seems like a lot can be done with just text while we wait for multimodal models to catch up. The recent Platonic Representation Hypothesis [1] also suggests that different models, regardless of modality, build the same internal representations of the world.

[1] https://arxiv.org/abs/2405.07987

KhoomeiK··on ScrapeGraphAI: Web scraping using LLM and direct graph logic
1 month away ;)
KhoomeiK··on xLSTM: Extended Long Short-Term Memory
For those who don't know, the senior author on this paper (Sepp Hochreiter) was the first author on the original paper with Schmidhuber introducing LSTMs in 1997.
KhoomeiK··on ScrapeGraphAI: Web scraping using LLM and direct graph logic
Yep, until you generate code—it's harder from a technical POV but you can get way higher performance & reliability.
KhoomeiK··on ScrapeGraphAI: Web scraping using LLM and direct graph logic
This is essentially what we're building at https://reworkd.ai (YC S23). We had thousands of users try using AgentGPT (our previous product) for scraping and we learned that using LLMs for web data extraction fundamentally does not work unless you generate code.
Page 1 of 9Next →