HNHacker News
TopNewBestAskShowJobs

taesiri

238 karma · joined April 24, 2015

submissionscomments
taesiri··on ChatGPT/Gemini can now draw on your screen to help you navigate complex software
would be sick with meta glasses; just look at broken things, draws what you mean, and get help fixing it. not just fixing but anything
taesiri··on Vision Language Models Are Biased
for overly represented concepts, like popular brands, it seems that the model “ignores” the details once it detects that the overall shapes or patterns are similar. Opening up the vision encoders to find out how these images cluster in the embedding space should provide better insights.
taesiri··on Vision Language Models Are Biased
State-of-the-art Vision Language Models achieve 100% accuracy counting on images of popular subjects (e.g. knowing that the Adidas logo has 3 stripes and a dog has 4 legs) but are only ~17% accurate in counting in counterfactual images (e.g. counting stripes in a 4-striped Adidas-like logo or counting legs in a 5-legged dog).
taesiri··on Understanding Generative AI Capabilities in Everyday Image Editing Tasks
tldr; We find that GenAI can satisfy 1/3 of everyday image editing requests, while 2/3 of the requests are better handled by human image editors.
taesiri··on Hot: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
Abstract:

An Achilles heel of Large Language Models (LLMs) is their tendency to hallucinate non-factual statements. A response mixed of factual and non-factual statements poses a challenge for humans to verify and accurately base their decisions on. To combat this problem, we propose Highlighted Chain-of-Thought Prompting (HoT), a technique for prompting LLMs to generate responses with XML tags that ground facts to those provided in the query. That is, given an input question, LLMs would first re-format the question to add XML tags highlighting key facts, and then, generate a response with highlights over the facts referenced from the input. Interestingly, in few-shot settings, HoT outperforms vanilla chain of thought prompting (CoT) on a wide range of 17 tasks from arithmetic, reading comprehension to logical reasoning. When asking humans to verify LLM responses, highlights help time-limited participants to more accurately and efficiently recognize when LLMs are correct. Yet, surprisingly, when LLMs are wrong, HoTs tend to make users believe that an answer is correct.

taesiri··on ZeroBench: An Impossible Visual Benchmark for Contemporary LMMs
Abstract:

Large Multimodal Models (LMMs) exhibit major shortfalls when interpreting images and, by some measures, have poorer spatial cognition than small children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by an ongoing surge of model progress. To address this, there is a pressing need for difficult benchmarks that remain relevant for longer. We take this idea to its limit by introducing ZeroBench-a lightweight visual reasoning benchmark that is entirely impossible for contemporary frontier LMMs. Our benchmark consists of 100 manually curated questions and 334 less difficult subquestions. We evaluate 20 LMMs on ZeroBench, all of which score 0.0%, and rigorously analyse the errors. To encourage progress in visual understanding, we publicly release ZeroBench.

taesiri··on ZeroBench: An Impossible Visual Benchmark for Contemporary LMMs
All frontier models, (o1, o1-pro, QVQ, gemini-flash-thinking) score exactly 0% on main questions of this benchmark.
taesiri··on Vision language models are blind
This paper examines the limitations of current vision-based language models, such as GPT-4 and Sonnet 3.5, in performing low-level vision tasks. Despite their high scores on numerous multimodal benchmarks, these models often fail on very basic cases. This raises a crucial question: are we evaluating these models accurately?
taesiri··on Is there any good place to track all updates related to LLMs?
Not one place, but there are some people tweeting about new papers daily (@arankomatsuzaki, @_akhaliq, @omarsar0) other people summarizing papers (@davisblalock, @rasbt). Latent Space podcast is also great and of course r/LocalLLaMA/ is an amazing place to share and learn.
taesiri··on Show HN: Fully isolated honeypot SSH server using thrussh
Coool! Would be nice to have an option to send commands to an LLM and show the results to the "user"! :D