HNHacker News
TopNewBestAskShowJobs

hannah-pdx

35 karma · joined August 23, 2022

submissionscomments
hannah-pdx··on Archive your Twitter alt text
Twitter offers the ability to archive your tweets and other data they keep about you, but that archive doesn't include alt text. This tool will take an archive zip file or the extracted tweets.js and scrape twitter for all the alt text you've ever written. It will return everything in a JSON file for future work to incorporate with the downloaded archive. For now it also offers https://archive.alt-text.org/search, which will allow fuzzy searching of the alt text you've written.
hannah-pdx··on Non-Machine Learning Image Matching with a Vector DB
Oh, and did you explore using more lightness buckets than the 5 from the paper?
hannah-pdx··on Non-Machine Learning Image Matching with a Vector DB
Ooh, very cool. Were the quoted insertion rates including the vectorization times, and was that using the ES or Mongo backend?
hannah-pdx··on Non-Machine Learning Image Matching with a Vector DB
The full cropped results are in the "Correctness" tab of the Google sheet linked at the bottom, with more details in the "Scoring" tab, but TL;DR the intensity vector worked best, with Goldberg (the one I chose) a pretty close second. Goldberg correctly returned the correct result highest-scored in 79% of cases, with it present in the result in 90%.

I'm primarily interested in managed services. I've been an SRE and I hope to not be in that role again.

hannah-pdx··on Non-Machine Learning Image Matching with a Vector DB
Hmm, the smallest model I see is still 4.3MB, are there smaller?

My quick read of the stats offered says that the smaller model's accuracy suffers considerably. I could definitely see running similar tests on it though!

All that said, the matching tensorflow offers from my understanding is also not exactly what I'm after. I'm primarily concerned with matching identical-to-humans images, possibly with small modifications such as size changes. Think more "are these two images identical" vs "give me pictures of dogs"

hannah-pdx··on Non-Machine Learning Image Matching with a Vector DB
As part of the development of a system that requires searching by image, we needed to compute feature vectors for use with the Pinecone vector database. All the research we could find focused on either ML approaches, which were untenable due to hopes to perform vector generation in the browser, or hamming distance vector comparison, which are untenable for large scale search.

The README here contains my research into several algorithms' performance, and the repo contains the code that performed the data gathering.

The site alt-text.org is still alpha quality and under active development, and the library backing it is quite small so most searches will fail, but feel free to play around with it.

Twitter users can help build the library with the link in the upper right corner, though it does not yet work on mobile.