HNHacker News
TopNewBestAskShowJobs

snats

289 karma · joined July 13, 2023

snats.xyz
submissionscomments
snats··on Show HN: Watch a neural net learn to play Snake
did a pretty similar thing last month for the text rendering library last month.

trained and made a viz for the model and then made it displace text.

should probably do a proper write-up:https://x.com/i/status/2038367016969724259

snats··on Show HN: Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training
you can also have removed layers of models and keep the same score in benchmarks [1].

i feel that sometimes a lot of the layers might just be redundant and are not fully needed once a model is trained.

[1] https://snats.xyz/pages/articles/pruningg.html

snats··on Ask HN: Share your personal website
personal website: snats.xyz weblog: weblog.snats.xyz
snats··on Just how resilient are large language models?
I also did a couple of experiments with pruning LLMs[1] using genetic algorithms and you can just keep removing a surprising amount of layers in big models before they start to have a stroke.

[1]https://snats.xyz/pages/articles/pruningg.html

snats··on I want an iPhone Mini-sized Android phone (2022)
i get it. i want one of those. the problem is that most cellphones are not actual cellphones, they are entertainment machines. they are a pocket tv / social media feed place. most usage for my normal friends is for that.
snats··on Ask HN: What are you working on? (April 2025)
I am working on the https://moviemovie.club/about, it's a tiny website about film review.

It works like a run club, where you have to make a review first to see other people's reviews.

I am currently implementing watchlists, comments and a mural to make it feel a bit less lonely. Right now I like the UI but it feels to lonely.

snats··on Ask HN: Any insider takes on Yann LeCun's push against current architectures?
Not an insider but imo the work on diffusion language models like LLaDA is really exciting. It's pretty obvious that LLMs are good but they are pretty slow. And in a world where people want agents you want a lot of the time something that might not be that smart but is capable of going really fast + searches fast. You only need to solve search in a specific domain for most agents. You don't need to solve the entire knowledge of human history in a single set of weights
snats··on 'A lot worse than expected': AI Pac-Man clones, reviewed
It's pretty funny to test in-distribution for AI models. But they fail horribly once you push them a bit[1].

I recently made LLMs play Minesweeper and ALL LLMs that I tested had a pretty bad win to loose ratio. Like the only model that won more than 3 times was R1 (mind you there were 50 games).

[1] https://snats.xyz/pages/articles/minesweeper_bench.html

snats··on Fewer students are enrolling in doctoral degrees
yup, if i went to do a PhD interpretability is the only interesting subject for academia IMO right now
snats··on OpenAI’s board, paraphrased: ‘All we need is unimaginable sums of money’
It's more of a distilled model, not a fair 1:1 comparison
snats··on Phishers Love New TLDs Like .shop, .top and .xyz
I use .XYZ because it was pretty cheap when I bought it
snats··on Ask HN: Bluesky Accounts Worth Following for HN Enthusiasts
I post about ML/DL and SWE at https://bsky.app/profile/snats.xyz
snats··on LLaVA-O1: Let Vision Language Models Reason Step-by-Step
This paper is not comparing against MOLMO or Qwen, so I would take it with a grain of salt
snats··on Low-poly image generation using evolutionary algorithms in Ruby (2023)
I made some image generation with CLIP and Evolutionary algorithms not so long ago[1] and the results had a bit more life than I was expecting. It is still a contested area of research and you have some cool stuff like CLIPDraw[2] where they use gradient decent to approximate a vector to the embeddings of CLIP.

[1] https://snats.xyz/pages/articles/optimizing_images.html [2] https://arxiv.org/pdf/2106.14843

snats··on Ask HN: Where After WordPress?
let me vouch for 11ty.dev, i love it. it is simple, easy to use, and slim. you don't have to worry about fighting the million dogs around and if that seems like to much. just go and use raw html, i think that ever since i've been using raw html i've actually been publishing more blogposts
snats··on I want to break some laws too
just changed! thanks for the heads up
snats··on Classifying all of the pdfs on the internet
hi! i used the average accuracy over the entire dataset made originally made by the llm
snats··on Classifying all of the pdfs on the internet
Hi! Author here, I wasn't expecting this to be at the top of HN, AMA
snats··on Show HN: LLM-aided OCR – Correcting Tesseract OCR errors with LLMs
Is there any goodmodel for OCR but on handwritten information? I feel like most models are currently kind of trash
snats··on Tech, Crunched: How the go-to site for startup news lost its way
Deeptech is going crazy, all of the AI boom is pretty much an obvious subject if you are in media, and finally you still have your usual gossip. You still have a lot of eyes
snats··on Crafting Interpreters
You can do the second part that is C!
snats··on “Meta spent almost as much as the Manhattan Project on GPUs in today's dollars”
Remember that Llama400B is still in training and could end up crushing a lot of the use cases for GPT-4
snats··on Infini-Gram: Scaling unbounded n-gram language models to a trillion tokens
I recently wrote a writeup on bigrams and the infinigram outputs[1]. I genuinely believe that ngrams are making a comeback. For query searching it's really good.

[1] https://snats.xyz/pages/articles/from_bigram_to_infinigram.h...

snats··on RAGFlow is an open-source RAG engine based on OCR and document parsing
A bit off topic but every time I see DTrOCR I remember that marketing is a good idea because DTrOCR[1] and Fuyu[2] are basically the same architecture[3].

[1] https://arxiv.org/pdf/2308.15996v1.pdf [2] https://www.adept.ai/blog/fuyu-8b [3] If you don't want to search for the figures I made a tiny post about it on my weblog: https://weblog.snats.xyz/posts/2024/02/16/

snats··on Can Demis Hassabis save Google?
It's been pretty interesting to see some research trying to solve this. For example, Stanford researchers recently published QuietSTaR[1]. This makes LLMs "think" before they speak. By generating a chain of thoughts and just choosing the best possible path from the generated thoughts.

This method is better than Chain of Thought and I think is a step in the right direction.

[1] https://arxiv.org/abs/2403.09629

snats··on Google's First Tensor Processing Unit: Architecture
to this day i am impressed that they have not figured out how to embed advertisments to bard outputs so that they can go free.
snats··on Mapping almost every law, regulation and case in Australia
No! I only tried PCA, but I still have the embeddings.

I'll try later and post results.

snats··on Mapping almost every law, regulation and case in Australia
I built a map of all the PDF urls on the internet recently.

I used a tiny embeddings model and PCA for dimensionality reduction.

https://weblog.snats.xyz/posts/2024/03/20/

snats··on Show HN: Pika – Simple blogging software
Love the website! Just beware that by making it completely free without any credit card required you are going to battle a bit of spam. I liked the solution that bearblog[1] implemented. I would love to hear how you solved this problem!

[1] https://herman.bearblog.dev/the-chatgpt-vs-bear-blog-spam-wa...

snats··on A decoder-only foundation model for time-series forecasting
A little bit of context. Basically, all of tabular deep learning has been stuck and SOTA has been tree based algorithms like Catboost and XGBoost. This seems like a big step forward towards getting a generalizable deep learning model besides this "tree models"
Page 1 of 2Next →