HNHacker News
TopNewBestAskShowJobs

tomhazledine

182 karma · joined March 22, 2017

I’m a frontend developer, specializing in complex UI for web apps. Check out my RSS reader app rssisawesome.com and my ESM-first static site generator jsssg.org
submissionscomments
tomhazledine··on Force a better group decision by making a bad choice first
I'm thinking of adopting a similar policy in our house. Already getting a lot of utility from the "how much do you care?" system https://adam.whittingham.dev/posts/how-much-do-you-care.html Hoping they'll work well in combination
tomhazledine··on Was an ancient bacterium awakened by an industrial accident?
If anyone else out there was put off by the potential for the title to conform to Betteridge’s Law of Headlines (“Any headline that ends in a question mark can be answered by the word no”), you can rest easy. This one is a “maybe”

https://en.m.wikipedia.org/wiki/Betteridge%27s_law_of_headli...

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
Some excellent points in there, thanks!

Really drawn to the idea of storing the final recommendations in a separate file. Didn't write it that way initially because most SSGs handle "data" files in their own unique way, and one of my goals was to be as "light touch" as possible. But I guess that could be solved with some config options (set `path/to/data.json` or similar).

I agree that "79% match" doesn't mean anything on it's own (and is, yes, totally arbitrary in a lot of ways), but it does provide some context when browsing across the whole site. It's a way to indicate that a "90% match" is more similar than a "60% match". Felt like useful info to me, so that's why I included it.

As for "why only top two?" - I'm constantly paranoid about adding too many bells and whistles to my blog and overloading the patience of whoever's taken the time to read. If I had official rules, they'd read "no carousels, and as few Calls To Action on a page as possible". I'm not super strict about it, but 1 recommendation felt too stingy and 3 felt like too many.

BM25 is new to me, and I'm mostly a n00b when it comes to "proper" search. But I'll definitely do some more reading. Currently setting up a head-to-head with an embedding index and a fuzzy-search library, but don't have any scientific way to measure the results. Sounds like you may have pointed me in the direction of a missing piece of the puzzle. Thanks!

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
Thanks for the heads up, although the legacy shut-off won't affect this script.

Not sure if it's clear from the post or not, but this script uses the newer, recommended, `v1/chat/completions` endpoint rather than the deprecated `v1/completions` endpoint

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
Very true! The real limiting factor on the success of this feature is the lack of posts on my blog :D

Works much more effectively for posts about topics I've covered a lot, like SVG or web audio. Happy with the outcomes in general, though, as I can now apply to more content-rich projects.

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
Even that simple explanation makes my brain itch a little :D I never did master trig

I'm curious if there ARE alternative methods to cosine similarity. A lot of the things I've read mention that cosine similarity is "one of the ways to compute distance..." or "a simple way...". But I've not seen any real suggestions for alternatives. Guess everyone's thinking "if it ain't broke, don't fix it" as cosine similarity works pretty darn well

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
To be honest I've sidestepped the issue now that GPT4 context size has increased to ~8000 tokens. None of my content goes over that limit.

I did build in a "chunking" mechanism to break down the article into sections if it was over the limit, but I'm not entirely sure how effective the summarisation would be for those... Summarising the first part of an article and then doing a separate summary for the n-th part would probably make confusing results when recombined.

Might be some "prompt magic" that can make it possible, but for the 90%-case I bet you'd get perfectly useable results just by using the first under-limit part of the content. Not tested that idea in the real world yet, though.

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
Yeah - it's a great idea. The size of the embeddings is the big restricting factor IMO. Even with my approach of embedding the entire article, my embeddings index was about the same size as my "regular" search index.

Once you start increasing the granularity of what you're embedding (either by paragraph or sentence) then the old-fashioned search index has a big advantage.

Might be worth it in some scenarios because of the quality of the results. I bet there are places where an embedding search would be more effective by orders of magnitude.

tomhazledine··on Show HN: My related-posts finder script (with LLM and GPT4 enhancement)
My default approach is to use as few dependencies as possible for Proof of Concept experiments like this one, but it DOES favour convenience over long-term efficiency and performance.

I think a single-author blog is about the limit for what can be handled by my flat-file + working memory approach. Any dataset that's much larger would almost certainly need a vector-friendly DB. Typesense looks like it might be a good fit, and so does the pgvector extension of postgres

tomhazledine··on Show HN: Fuse.js: The opinionated framework for creating typesafe data layers
Looks interesting. But instantly makes me think of the JavaScript fuzzy-search library Fuse.js https://www.fusejs.io/

Possibly a confusing name collision here.

tomhazledine··on Acoustic Location and Sound Mirrors (2004)
Such an effective trick when you experience it in person.

I remember visiting Goonhilly Earth Station as a boy (a collection of huge satellite dishes at the very tip of Cornwall, UK) and they had two small (~2 or 3m diameter) dishes facing each other on either side of the visitor centre. Completely blew my mind to hear someone whispering from all the way across the room!

tomhazledine··on What is a decibel?
That's a great point. The piece was written from a very web-centric point of view. Even though I used analogue mixing desks as my jumping off point, I only really looked at decibels in a digital context.

In fact, all the digital VU meters I've looked at do account for positive values, so I definitely shouldn't have breezed over that aspect so quickly.