584 karma · joined January 22, 2011
Their prediction was spot on: “We expect that advertising-funded search engines will be inherently biased toward the advertisers and away from the needs of the consumers.”
I wonder how it compares to an vector indexing approach.
[0] https://twelvetables.blog/comparing-claude-fable-5s-system-p...
> we found that employees worked at a faster pace, took on a broader scope of tasks, and extended work into more hours of the day, often without being asked to do so.
> On their own initiative workers did more because AI made “doing more” feel possible, accessible, and in many cases intrinsically rewarding.
I tried OneCommander and they're super fast, so it's not something slowing down disk IO, it's purely File Explorer.
Now I'm still struggling with closing chrome tabs being super slow sometimes.
[0] https://en.wikipedia.org/wiki/Static_single-assignment_form
Since they're using Arrow they might look into Flight RPC [1] which is made for this use case.
Haven't yet had the same issue with Wikipedia.
If you needed to look up say the 100 most recent documents, that would require ~100+ disk seeks at random locations just to look up the index due to the random nature of UUIDv4. If they were sequential or even just semi-sequential that would reduce the number of lookups to just a few, and they would be more likely to be cached since most hits would be to more recent rows. Having it roughly ordered by time would also help with e.g. partitioning. With no partitioning, as the table grows, it'll still have to traverse the B-Tree that has lots of entries from 5 years ago. With partitioning by year or year-month it only has to look at a small subset of that, which could fit easily in memory.