1,299 karma · joined August 7, 2018
https://rpandey.tech
Seems like there might be room for a "LangChain for Video" in this space...
Etymology tidbit: "Bhārata", India's Sanskrit name, refers to the forerunner clan that established India's first historically recorded political entity—the Kuru Kingdom—around 1200 BC near modern Delhi. The clan itself was named "Bhārata" due to their ardent bearing ("bhar-" in Sanskrit) of the sacred fire.
We haven’t tried Apple’s OCR but hopefully will integrate Azure OCR soon based on others’ advice.
"Keep in mind that Tarsier tags different types of elements differently to help your LLM identify what actions are performable on each element. Specifically:
[#ID]: text-insertable fields (e.g. textarea, input with textual type)
[@ID]: hyperlinks (<a> tags)
[$ID]: other interactable elements (e.g. button, select)
[ID]: plain text (if you pass tag_text_elements=True)"
Do you see the search boxes labeled [#4] and [#5] at the top? And before you say that the tag is on a different line from the placeholder text—yes, and our agent is smart enough to handle that minor idiosyncrasy. Are you shocked? :)
[1] https://github.com/reworkd/tarsier/blob/main/.github/assets/...
[2] https://github.com/reworkd/tarsier/blob/main/.github/assets/...
At Reworkd, we're focused on web agents for data extraction at scale, which isn't as hyped as the generalist agents but we find provides a lot of value and already works pretty well.
It's sort of like when you (as a human) write a web scraper and visually click on individual elements to look at the surrounding HTML structure / their selectors, but then end up writing code with more general selectors—not copypasting the selectors of the elements you clicked.
Will have to look into supporting Azure OCR in Tarsier then—thanks for the tip!