PostgreSQL built-in FTS provides a score for each row based just on the data for that row.
Relevance algorithms like BM25 take overall corpus statistics into account. If you search for a bunch of words and some of them are less common than others in the overall set of documents, documents that match THOSE words will score higher than matches for other words in your search.
That's what all of these additional extensions are providing.
Thanks for pointing out a real difference.
- superior performance
- superior operational overhead
- no second copy of data in tsvector form
- BM25 scoring support with optimized top-k output
- runtime configurable scoring knobs
- expression-attached score boosting
- sophisticated span query support -- this is proximity search on steroids (https://github.com/planetscale/lead/tree/main/tinql/docs)
- lossless term positions
- index-answerable negative expressions (find all docs that don't contain a word)
- full document hit highlighting
- optimized exact `count(\*)`
- term expansion via any of fuzzy matching, wildcards, regular expressions, and dictionary ranges
- intentionally smaller user-facing SQL API surface
There's a lot we didn't cover in the announcement blog. I'm sure we'll do more as time goes on.As an aside, something I personally think is cool, and I suppose you can do this with Postgres' built-in `@@` too, is that you can use TIN's full query language (linked above) against any text datum. This is a valid query:
SELECT pid, query
FROM pg_stat_activity
WHERE query ==> 'select OR copy'
in other words, you don't need an index at all to use TIN's full query language against any text field in any query.Incredible? No.
You want to use this thing instead?