HNHacker News
TopNewBestAskShowJobs

LeonardoTolstoy

281 karma · joined May 3, 2021

submissionscomments
LeonardoTolstoy··on OCR It – pull text out of un-copyable documents for your LLM
I, at this point, use Qwen2.5-VL-3B-Instruct for most of the small OCR I want to do. It is much much better than my experience with Tesseract in general. The nice thing about it is that if you give it, say, a movie poster you can ask for the "title of the movie" and it will, to the best of its ability, do just that, no need for regex or filtering after. For smallish images after loading the 3B model runs in <1 second. 7B takes longer but is obviously more accurate.

I might be a bit behind, all of this is from early this year for the most part, but for something like "I have 3000 movie posters and I want to get the titles with like 90% accuracy" it is good (much better than Tesseract), and it'll do that in like an hour.

EDIT: I guess one thing is Tesseract will kind of give gibberish back when it fails. The main issue with the LLMs are that instead they take a stab at it (like for a movie poster it'll give part of a quote, or a actor name) back. Makes knowing when it fails a little harder. As long as you have some way to verify when it is likely failing they are very good though.

LeonardoTolstoy··on I indexed 669 GB of my GoPro videos using my M1 Max computer and local ML models
What models did you use for the stages? I see Qwen2.5-VL-7B-Instruct mentioned as an advanced option, so I assume maybe Qwen2.5-VL-3B-Instruct by default (which is what I also use for a lot of stuff, it is incredibly good at "clean" OCR, but as you maybe indicate not the best at "describing a scene").

EDITED: I didn't realize Whisper was a local model. I never tried transcription before, so I had always figured it was a pay model by OpenAI. I'll have to check it out (although the runtime listed here is a bit daunting).

For that project I'll say I don't see much degradation in embedding quality at much much worse quality than 720p (all the way down to 240p), which speeds things up considerably. Although I don't really do face or object detection, just scene embeddings. To me any process whereby it would take longer to process the video than watch it is probably a no go in general. Obviously a challenge for local-first analysis.

LeonardoTolstoy··on US special forces soldier arrested after allegedly winning $400k on Maduro raid
I feel like if you followed the NBA scandal involving Chauncey Billups the wire fraud charge for insider prediction market trading was inevitable.

Damon Jones didn't work for the NBA and basically just told some people the status of an injury to LeBron because he hangs out with him (in exchange for money). His crime I guess is gambling illegally? But wire fraud (I think they even say "creating a fraudulent market") was thrown in there.

Seemed inevitable they were going to start charging prediction market insiders the same way.

LeonardoTolstoy··on The threat is comfortable drift toward not understanding what you're doing
It is a spectrum. My advisor was very hands off. He didn't, ultimately, even really understand my PhD. He knew the problem, but he had no path in mind to solve it, that was up to me. I'm now working (as a software engineer) with a person who is very hands on with his students (and even postdocs) to the point of giving them specific tasks to do and then discussing the result every week. He defines the problems and structure of the solution, the students at least partially are an extension of himself, they are doing stuff he merely doesn't have time to do himself.

And there is everything in between.

LeonardoTolstoy··on Ask HN: Best Podcasts of 2025?
Most of my podcasts are movie related. If I had to purge them all and start with just 5 though I would go with.

Blank Check The Flophouse 99% Invisible Cautionary Tales The Rewatchables

I maintain The Flophouse is the funniest podcast around.

LeonardoTolstoy··on Ask HN: What did you read in 2025?
I read close to 40 books this year which I assume is up from the previous year, and way more than I used to read prior to starting to do the Rochester Library reading challenge (which I've done for the past two years). I won't get into a lot of it, but after reading Middlemarch three years ago, Napoleon: A Life two years ago, and The Power Broker last year I decided to read two "big" books.

In the first half of the year War and Peace, which obviously was excellent, although I liked Middlemarch more.

In the second half of the year David Copperfield which was very excellent. Just beautifully written. I still think I probably like Middlemarch more... but it might now require a re-read to know for sure.

This year is going to be Tom's Crossing to start as I just got that for Christmas.

LeonardoTolstoy··on Ask HN: Our AWS account got compromised after their outage
This is it. I had the same thing happen to me a year ago and there was a month between the original access to our system and the attack. And similarly they waited until a perceived lull in what might be org diligence (just prior to thanksgiving) to attack.
LeonardoTolstoy··on Ask HN: Our AWS account got compromised after their outage
Almost this exact thing happened to me about a year ago. Very old account login, SES access with request to raise the email limit. We were only quickly tipped off because they had to open a ticket to get the limit raised.

If you haven't check newly made Roles as well. We quashed the compromised users pretty quickly (including my own, the origin we figured out), but got a little lucky because I just started cruising the Roles and killing anything less than a month old or with admin access.

To play devil's advocate a bit. In our case we are pretty sure my key actually did get compromised although we aren't precisely sure how (probably a combination of me being dumb and my org being dumb and some guy putting two and two together). But we did trace the initial users being created to nearly a month prior to the actual SES request. It is entirely possible whomever did your thing had you compromised for a bit, and then once AWS went down they decided that was the perfect time to attack, when you might not notice just-another-AWS-thing happening.

LeonardoTolstoy··on A PhD in Snapshots
Does that count the time it takes to get a masters? I feel like I recall my coworkers in England doing a 1-2 year masters, and then the 3-4 year PhD after.
LeonardoTolstoy··on You’re a slow thinker. Now what?
Maybe a nonsequitur but in grad school I was in a study group which naturally split into two. In one group (mine) we'd read a problem and immediately charge in, sometimes have to backtrack, and meander around until the answer revealed itself. In the other they would plan everything out, and figure out what they needed to do, and from that the answer would reveal itself and they would write it all down.

The interesting part is neither group really finished the problem sets faster than the other. Individual problems my group could, if we knew or guessed the right path immediately, be faster. But over the span of a 10 question p-set it would mostly come out in the wash and both groups would finish in roughly the same amount of time.

I often think back on that when reflecting on how I still work that way years later.

LeonardoTolstoy··on Streaming services are driving viewers back to piracy
Library. I order DVDs to my local library all the time. Maybe your library system is terrible, but if it isn't you certainly can do this. There are hundreds a big wide release films with, effectively, are only available on DVD from a library (legally).

Bonus when I go I can still get that browsing the aisle experience like in an old video store (but in this case I am lucky, my local library has a large DVD / Blu-ray collection to browse)

LeonardoTolstoy··on What would you name a new programming language?
Tolstoy after my old dog. But also because I think a good language name should be able to be iterated on. E.g. C can become D or E or F. Java can be Mocha or Cappuccino. Etc.

With my new language Tolstoy you'd be able to have a little family of languages all named after classic authors. Tolstoy, Dickens, Melville etc. Plus my dog Tolstoy was the best dog ever so bonus for everyone as well.

LeonardoTolstoy··on U.S. senators introduce new pirate site blocking bill, "Block BEARD"
You have to go pretty deep though for the record. At least, using one of your examples, for Altman if you look at his top 25 films on Letterboxd, 20 of them are available to rent or stream online. And for me at least the other five I can get at the library. There are none that are totally unavailable of those 25.
LeonardoTolstoy··on U.S. senators introduce new pirate site blocking bill, "Block BEARD"
How's your local library system?
LeonardoTolstoy··on Have We Stopped Inventing Futures Worth Predicting?
My local library and the nearby highschool both have solar panels over the parking lot. They certainly exist. I suppose the ROI they've had would be interesting (although this is in Massachusetts so not exactly the solar capital of the world as far as potential is concerned).

I do like them though. Gives both shade and (some) rain cover.

LeonardoTolstoy··on Ask HN: What's your favorite book you've read?
Middlemarch by George Eliot. Well worth a read, possibly the greatest English novel ever written.
LeonardoTolstoy··on A non-anthropomorphized view of LLMs
And we are there. A boat sails, and a submarine sails. A model generates makes perfect sense to me. And saying chatgpt generated a poem feels correct personally. Indeed a model (e.g. a linear regression) generates predictions for the most part.
LeonardoTolstoy··on Fictional K-pop bands zoom to top of US music charts
Yes this is what I thought, that it was AI fake bands basically. This sounds a bit more like Sugar Sugar by The Archies, which is a made up cartoon band with a number one hit.
LeonardoTolstoy··on A non-anthropomorphized view of LLMs
What does a submarine do? Submarine? I suppose you "drive" a submarine which is getting to the idea: submarines don't swim because ultimately they are "driven"? I guess the issue is we don't make up a new word for what submarines do, we just don't use human words.

I think the above poster gets a little distracted by suggesting the models are creative which itself is disputed. Perhaps a better term, like above, would be to just use "model". They are models after all. We don't make up a new portmanteau for submarines. They float, or drive, or submarine around.

So maybe an LLM doesn't "write" a poem, but instead "models a poem" which maybe indeed take away a little of the sketchy magic and fake humanness they tend to be imbued with.

LeonardoTolstoy··on Ask HN: How to regain the ability to read with focus and learn
I faced a similar thing at one point. The thing that fixed it for me was going back to fiction. Started with a fantasy series. Did a little sci-fi. Did some easy classics (e.g. Dracula). Eventually that joy of reading, long focused reading, and effortless comprehension all came back. Just took practice.

Much like advice for writer's block being often "just write!". The same goes for reading. Start with something easy breezy and eventually it'll all start flowing again IMO.

LeonardoTolstoy··on Endometriosis is an interesting disease
It says that there is like a 10x risk of SIDS in the first four months of life with tummy sleeping.

I don't agree with her on everything, but Emily Oster's chapter on SIDS (in the second book I think, Cribsheet) I think does a good job outlining the data on it. And my brother just had a kid who also would absolutely not sleep on his back. Once he could roll he just sleeps on his tummy (but once they can roll SIDS is not really an issue)

LeonardoTolstoy··on Machine Learning: The Native Language of Biology
This person seems to work in a field (exercise / athletics) with an abundance of data, low stakes outcomes, reasonably well established biomarkers, etc. in other words, a field perfectly suited for a top down outcome driven analysis.

IMO the post is merely stating: "man, everyone should be doing this!" Without realizing that (1) everyone is doing this, and (2) it doesn't seem like it because many (most?) fields in biology don't work in the top down approach being suggested. Determining mechanism and function is vital in biology because in a lot of cases there just isn't the data to perform a fuzzy outcome driven analysis.

LeonardoTolstoy··on How to Read a Novel
The claim is not that far off from someone claiming they watch a movie a week. A book probably doesn't take much more than 10 hours to read on average. 50 movies are about 100 hours. Is one movie a week a mind-boggling amount of free time?
LeonardoTolstoy··on The Zach Attack Scratch 'N Solve Puzzle Pack
I used to crush Scratchees as a kid. I never knew they were made by Decipher (which IMO were more famous for their How to Host a Murder games than anything else). I'll definitely check these out.
LeonardoTolstoy··on Ask HN: Share your AI prompt that stumps every model
A bit of a non sequitur but I did ask a similar question to some models which provide links for the same small helicopter question. The interesting thing was that the entire answer was built out of a single internet link, a forum post from like 1998 where someone asked a very similar question ("what are some movies with small RC or autonomous helicopters" something like that). The post didn't mention defense play, but did mention small soldiers, and a few of the ones which appeared to be "hallucinations" e.g. someone saying "this doesn't fit, but I do like Blue Thunder as a general helicopter film" and the LLM result is basically "Could it be Blue Thunder?" Because it is associated with a similar associated question and films.

Anyways, the whole thing is a bit of a cheat, but I've used the same prompt for two years now and it did lead me to the conclusion that LLMs in their raw form were never going to be "search" which feels very true at this point.

LeonardoTolstoy··on Ask HN: Share your AI prompt that stumps every model
Something about an obscure movie.

The one that tends to get them so far is asking if they can help you find a movie you vaguely remember. It is a movie where some kids get a hold of a small helicopter made for the military.

The movie I'm concerned with is called Defense Play from 1988. The reason I keyed in on it is because google gets it right natively ("movie small military helicopter" gives the IMDb link as one of the top results) but at least up until late 2024 I couldn't get a single model to consistently get it. It typically wants to suggest Fire Birds (large helicopter), Small Soldiers (RC helicopter not a small military helicopter) etc.

Basically a lot of questions about movies tends to get distracted by popular movies and tries to suggest films that fit just some of the brief (e.g. this one has a helicopter could that be it?)

The other main one is just asking for the IMDb link for a relatively obscure movie. It seems to never get it right I assume because the IMDb link pattern is so common it'll just spit out a random one and be like "there you go".

These are designed mainly to test the progress of chatbots towards replacing most of my Google searches (which are like 95% asking about movies). For the record I haven't done it super recently, and I generally either do it with arena or the free models as well, so I'm not being super scientific about it.

LeonardoTolstoy··on How to Make Superbabies
With today's knowledge of genetics it ain't happening. We have barely started gene editing for single point mutation diseases like sickle cell. The stuff described in this blog post (which is mostly gobbledygook by the way) are traits that are (1) hugely polygenic and (2) not known well enough from a genetic perspective anyways. It will be decades at best before they'd attempt genetic editing in that scale on a healthy person.

But if we want to solve the second one I welcome it as a person who works in genetics research, always good to have more money.

LeonardoTolstoy··on Introduction to Stochastic Calculus
It has been a while since I studied along these lines (stochastic chemical reaction simulations in my case) but I think the answer is often yes, but not always (I don't think). A random walk for example will be a normal distribution (and you know the mean, and you know the variance is going to infinity), so I do think in that case you end up with an elegant analytical solution if I'm understanding correctly as the inputs can determine the function the variance follows through time.

But often no, you need to run a stochastic algorithm (e.g. Gillespie's algorithm in the case of simple stochastic chemical kinetics) as there will be no analytical solution.

Again it has been a while though.

LeonardoTolstoy··on Tennis great Steffi Graf got bitten by pickleball's 'fun factor'
Padel is great, and I do wish it would make it over to the northeast US somewhat. Paddle tennis has some presence, but Padel is basically nonexistent still
LeonardoTolstoy··on Robot Jailbreak: Researchers Trick Bots into Dangerous Tasks
To actually implement it we would have to completely understand how the underlying model works and how to manually manipulate the structure. It might be impossible with LLMs. Not to take Asimov as gospel truth, he was just writing stories afterall not writing a treatise about how robots have to work, but in his stories at least the three laws were encoding explicitly in the structure of the robot's brain. They couldn't be circumvented (in most stories).

And in those stories it was enforced in the following way: the earth banned robots. In response the three laws were created and it was proved that robots couldn't disobey them.

So I guess the first step is to ban LLMs until they can prove they are safe ... Something tells be that ain't happening.

Page 1 of 3Next →